Repository navigation
CI: get development green and run the tests in ~7 minutes (#132: CI-0, CI-12, CI-1/2/3) - #137
Merged
Blumster merged 2 commits intoOct 5, 2026
Conversation
- LoginPacket.Read rejects trailing bytes (EnsureFullyConsumed), which AuthLoginRejectsTrailingBytes has expected since InfiniteRasa#96. - MissionReviewFindingTests builds its paths with Path.Combine, so it finds the files on the Linux runner. - Tests read each navmesh once per process (TestNavMeshes) and give every harness its own NavMeshQuery over the shared, read-only mesh, instead of reading a fresh 15 MB Wilderness mesh per harness. The whole suite in one process now peaks at ~1.4 GB, down from 12 GB+ and climbing. - Two tests are quarantined, each with a comment on why: ForeanEscortsCanFollowFromTheCaveExitToTheReclaimedBase depends on wall-clock time (creature timers use Environment.TickCount64), and ConcurrentDeadlineEvaluationCommitsOneCompletion catches a real race (SceneDueQueue is read outside the dispatch lock, so two overlapping deadline evaluations can both drop the deadline). Both failed intermittently on loaded CI runners. Refs InfiniteRasa#132. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012rNg9gkomRvRWZ23oM7NVX
…/2/3) The suite takes ~56 minutes in one process on a hosted runner. It now runs in 16 test processes, 2 on each of 8 runners, in ~7 minutes once the runners start. - build: restores (NuGet cache restored on every run, saved only from development: CI-3), builds once, plans the lanes and hands the test build to the runners as one artifact. - .github/scripts/TestShards.cs finds the tests in the built assembly and deals them out slowest first, on durations averaged over the last 3 green runs (test case counts when there are none). Classes bigger than half a lane are split by method, so MissionDialogueTests no longer sets the run's length. The last lane runs everything the others don't list, so a test the planner misses still runs once. - test (runner N): 2 lanes side by side, each its own test host, with a 10-minute hang dump, a GC heap cap so an overgrown host fails with a stack trace instead of losing the runner, and TRX results uploaded. - tests complete: the one check to require. It totals every lane's results and fails unless the build and every runner passed; a PR that changes only docs skips the build and tests and still passes it (CI-2; job conditions, not paths-ignore). - ubuntu-24.04, a timeout on every job, checkout v7, setup-dotnet v6, a read-only token by default, and a new push to a PR cancels its older run (CI-0, CI-1). Why 8 x 2: a 4-vCPU runner behaves like ~2 cores. Measured, 2 lanes per runner slow each test ~1.1x, 3 lanes ~1.4x, 4 lanes ~1.8x. 8 x 2 is the fastest, keeps clock-sensitive tests near their unloaded speed, and its 11 jobs fit the Free plan's 20 concurrent jobs. Refs InfiniteRasa#132. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012rNg9gkomRvRWZ23oM7NVX
This was referenced Oct 5, 2026
Open
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Gets
developmentgreen and fast. This is the first PR for #132 and covers CI-0 and CI-12, plus the parts of CI-1, CI-2 and CI-3 that live in the same workflow. CI-4 through CI-11 aren't included.Result: the full suite (2,613 tests) passes in about 7 minutes of wall time. In one process it takes about 56 minutes on a hosted runner, and today it never finishes because the runner is lost. This was measured on a fork of this repo with these exact commits (run).
Two commits
1. Get the test suite passing on Linux in bounded memory (CI-0)
LoginPacket.Readrejects trailing bytes withEnsureFullyConsumed.AuthLoginRejectsTrailingByteshas expected this since Complete mission runtime and Bootcamp workflows #96.MissionReviewFindingTestsbuilds its paths withPath.Combine. The@"src\…"literals were file names with backslashes in them on Linux.TestSupport/TestNavMeshes.cshands every harness its ownNavMeshQueryover a shared mesh. Nothing in the game edits tiles after a load. Before, every Wilderness harness read its own 15 MB mesh, and the whole suite in one process grew past 12 GB. Now it peaks at about 1.4 GB. Production code is unchanged.[Ignore], a[GitHubWorkItem]and a comment on why. Both failed intermittently on loaded CI runners.ForeanEscortsCanFollowFromTheCaveExitToTheReclaimedBasedepends on real time. Creature timers (buffs, bombs, habits) run onEnvironment.TickCount64, so how far the escort gets in 1,200 simulated ticks depends on how fast the machine runs them. It also fails every time in class order on Windows. It needs an injectable clock, which is a bigger change than this PR.ConcurrentDeadlineEvaluationCommitsOneCompletioncatches a real race. It failed 2 times in about 9 full runs, withexpected: 1, actual: 0.SceneApplication.Ticktakes due timers fromSceneDueQueue(aPriorityQueueplus aDictionary, with no locking) outside_dispatchGate. So two overlapping deadline evaluations can both see the same due item, and each then skips it after the other's removal: the deadline is dropped and neither commits. It passed 48 of 48 times locally, because the threads rarely overlap on an idle machine. The fix belongs in the scene host's locking, not in a CI PR. It deserves its own issue.2. Run the tests in parallel lanes across runners
dotnet.ymlnow has four jobs:detect code changes*.md,docs/,.github/agents/,.claude/), the build and tests are skipped. Anything in doubt runs everything.builddevelopment, because a PR's cache can only be read by that PR and the 10 GB limit evicts the oldest first.test (0)…test (7)tests completepaths-ignore, so this check always reports.Also:
ubuntu-24.04, a timeout on every job,checkout@v7,setup-dotnet@v6, a read-only token by default, and a new push to a PR cancels that PR's older run. Runs ondevelopmentare never cancelled, since a cancelled run saves no cache and leaves no timings.How the lanes are planned:
.github/scripts/TestShards.csruns withdotnet runand needs no packages.Rasa.Test.dllwithSystem.Reflection.Metadataand lists every[TestMethod]on every[TestClass].development. A PR whose base has none uses its own green runs, and with neither it balances on test-case counts.MissionDialogueTestsalone (about 4.8 minutes) would set the run's length.Where this differs from #132
Balanced on measured time, not hashed by class name. With hashing, the slowest shard came out about 50% over the ideal on this suite. Balancing on measured time hits the ideal.
No "listed total equals TRX total" check.
dotnet test --list-testsreports 2,492 tests while the TRX files contain 2,613, because data rows are counted differently. So that comparison would fail on a correct run. The "last lane runs the rest" rule makes completeness hold by construction instead.Lanes per runner. A hosted runner has 4 vCPUs, and the suite runs one test at a time. These are measured contention numbers from fork runs:
A runner behaves like about 2 cores. 8 runners × 2 lanes is the fastest configuration and keeps clock-sensitive tests close to their unloaded speed. It uses 11 jobs per run, inside the Free plan's 20 concurrent jobs. Both numbers are
envsettings at the top of the workflow.Debug is kept, as today. That's open question 4 in CI modernization: get
developmentgreen, then make it stay green #132, and the build configuration is a one-word change.masterstays in the push triggers. That's CI-11, a maintainer decision.Verified on the fork
All of these ran with this workflow on SandboxServers/Rasa.NET:
development: green, 2,613 results (the 2 quarantined tests skipped), 7.4 minutes (run).development: "Cache saved", and its timings used by later runs (run).docs/setup.md: build and tests skipped,tests completepassed (run).Worth knowing
developmentgreen, then make it stay green #132's memory section has thegcrootstep to find it.tests completeondevelopmentonce it has been green for a week (CI-10).Refs #132.
🤖 Generated with Claude Code
https://claude.ai/code/session_012rNg9gkomRvRWZ23oM7NVX