Commit Graph

11 Commits

Author SHA1 Message Date
clockwork-labs-bot 59b5d49ba2 Fix Internal Tests paths filter checkout (#5295)
## What changed

Adds an explicit checkout step before `dorny/paths-filter` in the
Internal Tests workflow.

## Why

`dorny/paths-filter@v3` needs a git working tree for `push` events. The
Internal Tests workflow ran it before any checkout, so every `push` run
on `master` failed immediately in `Detect non-docs changes` with:

```text
fatal: not a git repository (or any of the parent directories): .git
```

This only showed up consistently on `master` because those runs are
`push` events. On `pull_request` events, `dorny/paths-filter` can use
the GitHub pull request files API with the PR number, so it does not
need a local checkout for the same file detection path.

Adding checkout gives the action a repository when it handles `push`
events, while leaving PR behavior unchanged.

## Testing

- `git diff --check`
- PR #5295 `Internal Tests` job completed `Checkout` and `Detect
non-docs changes` successfully, then moved on to private dispatch/wait.

---------

Signed-off-by: Zeke Foppa <196249+bfops@users.noreply.github.com>
Co-authored-by: clockwork-labs-bot <clockwork-labs-bot@users.noreply.github.com>
Co-authored-by: Zeke Foppa <196249+bfops@users.noreply.github.com>
2026-06-16 20:14:16 +00:00
Zeke Foppa d31301a8f6 Move Internal Tests to its own workflow (#5147)
# Description of Changes

Moving `Internal Tests` to its own workflow so it can be canceled and
re-run independently from the rest of CI.

# API and ABI breaking changes

<!-- If this is an API or ABI breaking change, please apply the
corresponding GitHub label. -->

# Expected complexity level and risk

1

# Testing
- [x] `Internal Tests` succeed on this PR and appear to have run
properly

---------

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2026-05-29 20:49:12 +00:00
bradleyshep b75bf6decf LLM Benchmarking (#3486)
# Description of Changes

Introduce a new **LLM benchmarking app** and supporting code.

* **CLI:** `llm` with subcommands `run`, `routes list`, `diff`,
`ci-check`.
* **Runner:** executes globally numbered tasks; filters by `--lang`,
`--categories`, `--tasks`, `--providers`, `--models`.
* **Providers/clients:** route layer (`provider:model`) with HTTP LLM
Vendor clients; env-driven keys/base URLs.
* **Evaluation:** deterministic scorers (hash/equality, JSON
shape/count, light schema/reducer parity) with clear failure messages.
* **Results:** stable JSON schema; single-file HTML viewer to
inspect/filter/export CSV.
* **Build & guards:** build script for compile-time setup;
* **Docs:** `DEVELOP.md` includes `cargo llm …` usage.

This PR is the initial addition of the app and its modules (runner,
config, routes, prompt/segmentation, scorers, schema/types,
defaults/constants/paths/hashing/combine, publishers, spacetime guard,
HTML stats viewer).

### How it works
1. **Pick what to run**

* Choose tasks (`--tasks 0,7,12`), or a language (`--lang rust|csharp`),
or categories (`--categories basics,schema`).
   * Optionally limit vendors/models (`--providers …`, `--models …`).

2. **Resolve routes**

* Read env (API keys + base URLs) and build the active set (e.g.,
`openai:gpt-5`).

3. **Build context**

   * Start Spacetime
   * Publish golden answer modules
   * Prepare prompts and send to LLM model
   * Attempt to publish LLM module

4. **Execute calls**

* Run the selected tasks within each test against selected models and
languages.

5. **Score outputs**

* Apply deterministic scorers (hash/equality, JSON shape/count, simple
schema/reducer checks).
   * Record the score and any short failure reason.

6. **Update results file**

* Write/update the single results JSON with task/route outcomes,
timings, and summaries.


# API and ABI breaking changes

None. New application and modules; no existing public APIs/ABIs altered.

# Expected complexity level and risk

**4/5.** New CLI, routing, evaluation, and artifact format.

* External model APIs may rate-limit/timeout; concurrency tunable via
`LLM_BENCH_CONCURRENCY` / `LLM_BENCH_ROUTE_CONCURRENCY`.

# Testing

I ran the full test matrix and generated results for every task against
every vendor, model, and language (rust + C#). I also tested the CI
check locally using [act](https://github.com/nektos/act).

**Please verify**

* [ ] `llm run --tasks 0,1,2` (explicit `run`)
* [ ] `llm run --lang rust --categories basics` (filters)
* [ ] `llm run --categories basics,schema` (multiple categories)
* [ ] `llm run --lang csharp` (language switch)
* [ ] `llm run --providers openai,anthropic --models "openai:gpt-5
anthropic:claude-sonnet-4-5"` (provider/model limits)
* [ ] `llm run --hash-only` (dry integrity)
* [ ] `llm run --goldens-only` (test goldens only)
* [ ] `llm run --force` (skip hash check)
* [ ] `llm ci-check`
* [ ] Stats viewer loads the JSON; filtering and CSV export work
* [ ] CI works as intended

---------

Signed-off-by: bradleyshep <148254416+bradleyshep@users.noreply.github.com>
Signed-off-by: Tyler Cloutier <cloutiertyler@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: Tyler Cloutier <cloutiertyler@aol.com>
Co-authored-by: Tyler Cloutier <cloutiertyler@users.noreply.github.com>
Co-authored-by: spacetimedb-bot <spacetimedb-bot@users.noreply.github.com>
Co-authored-by: John Detter <4099508+jdetter@users.noreply.github.com>
2026-01-06 22:22:57 +00:00
Zeke Foppa ee69be7073 CI - Cancel internal tests if cancelled (#3674)
# Description of Changes

`Internal Tests` invoke another workflow. If the calling workflow is
cancelled, now we also cancel the invoked workflow.

# API and ABI breaking changes

None.

# Expected complexity level and risk

2

# Testing

- [x] Push multiple commits to this PR, see that the invoked job is
cancelled
- [x] Manually cancel the `Internal Tests` job, see that the invoked job
is cancelled

---------

Co-authored-by: John Detter <4099508+jdetter@users.noreply.github.com>
Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-11-17 22:14:33 +00:00
Zeke Foppa 703fae2ed6 Switch Internal Tests to reuse CI job (#3625)
# Description of Changes

Switch the `Internal Tests` job to re-use the existing CI workflow
instead of a separate one.

# API and ABI breaking changes

None. CI only.

# Expected complexity level and risk

1

# Testing

- [x] Internal Tests passes on this PR

---------

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-11-12 18:21:08 +00:00
Zeke Foppa 6bf3efc006 CI - Fix format strings (#3627)
# Description of Changes

I allowed chatgpt to mislead me about how to do these format strings.
Apparently this is the wrong syntax. I've now verified the correct
syntax:
https://docs.github.com/en/actions/reference/workflows-and-actions/expressions#format

# API and ABI breaking changes

CI only.

# Expected complexity level and risk

1

# Testing

I honestly don't know how to check what concurrency group something is
running in..

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-11-11 00:12:53 +00:00
Zeke Foppa 09bca44c56 CI - Make Internal Tests less brittle (#3536)
# Description of Changes

For some reason, the Internal Tests have trouble fetching the commit
sha, especially when the job is re-run. This PR switches it to using ref
names rather than commit sha, because the ref names are much more
durable than GitHub's ephemeral commits that it generates during
workflows.

# API and ABI breaking changes

None. CI only.

# Expected complexity level and risk

1

# Testing

- [x] CI still passes
- [x] Ref still gets checked out successfully on re-run.

---------

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-11-07 22:44:52 +00:00
Zeke Foppa 7c4c3ddeea CI - Fix the merge queue (#3571)
# Description of Changes

The merge queue was (partly) getting borked because we were putting all
non-PR CI events into the same concurrency group, which meant they all
non-PR CI jobs would run sequentially instead of running in parallel.
This sometimes caused _painfully_ long delays in the merge queue.

This was due to my misunderstanding in
https://github.com/clockworklabs/SpacetimeDB/pull/3501#discussion_r2466570395,
where I didn't realize that `cancel-in-progress: false` would cause
everything to queue up.

Now, for non-PR events, we append the commit SHA to the concurrency
group. For merge queue events, this should be the SHA of the ephemeral
merge commit that GH creates, so it will never conflict. For push events
or manual workflow dispatch events, the SHA should be a sane way to
recognize/cancel redundant events.

# API and ABI breaking changes

None. CI-only change.

# Expected complexity level and risk

1

# Testing

- [x] PR CI passes on this PR
- [x] PR CI is still canceled on this PR if a new commit is pushed

Unfortunately it's hard to test the behavior for non-PR events without
merging and seeing if it works.

---------

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-11-05 20:25:00 +00:00
Zeke Foppa 2516357c8d CI - Skip Internal Tests and Unreal Tests on external PRs (#3522)
# Description of Changes

These tests fail on external PRs, but not for any real reasons - just
because GH secrets are missing. "Skipped" is more informative than
"failed".

# API and ABI breaking changes

None.

# Expected complexity level and risk

1

# Testing

None, but I just copied the logic from the unity testsuite.

---------

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-10-28 19:28:12 +00:00
Zeke Foppa 9ad5e7038a CI - Cancel runs on new pushes (#3501)
# Description of Changes

Add `cancel-in-progress` to our GitHub workflows.

# API and ABI breaking changes

None

# Expected complexity level and risk

1

# Testing

- [x] Pushing new commits to this PR causes cancels of previous CI runs

---------

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-10-27 19:15:21 +00:00
Zeke Foppa 0f0cb47d03 CI - Move Internal Tests to GitHub (#3436)
# Description of Changes

Move a version of the Jenkins logic into a GitHub workflow. See the
linked issue for more context.

# API and ABI breaking changes

None. CI only.

# Expected complexity level and risk

2

# Testing

- [x] Tests currently passing on this PR
- [x] Tests fail with appropriate error if made to fail, e.g. version
bump
(https://github.com/clockworklabs/SpacetimeDB/actions/runs/18666569784/job/53219039721)
- [x] If the timeout is sharply reduced, we get an appropriate timeout
message
(https://github.com/clockworklabs/SpacetimeDB/actions/runs/18693760959/job/53305759272?pr=3436)

---------

Co-authored-by: Zeke Foppa <bfops@users.noreply.github.com>
2025-10-21 19:33:17 +00:00