Repository navigation
fix(ci): run the SDK examples gate 2-wide, not 4-wide - #13
Merged
Merged
Conversation
The gate failed 2 of 3 runs on ubuntu-latest, always on the same example
and always the same way:
[basic] page.screenshot: Protocol error (Page.captureScreenshot):
Unable to capture screenshot
- waiting for fonts to load...
- fonts loaded
The page had loaded and Chromium then failed to capture it, so this is a
capture-time resource failure rather than a test failure. 7 of 8 examples
passed every time.
The width was 4 on the 8-vCPU self-hosted runner, and I kept it at 4 when
moving to ubuntu-latest, reasoning that these block on live-AI calls rather
than CPU. That reasoning was incomplete: four concurrent headed Chromium
instances also contend for memory and display surfaces, neither of which is
a core count. The review on #12 flagged the width as unverified on this
runner class, and the one dry run that passed gave false confidence.
Halving the width trades wall time on that step for a gate that holds. What
this does not do is identify the exact constraint — there was no OOM in the
logs, so memory, /dev/shm and display surfaces are all still candidates. If
2-wide proves stable, the diagnosis is contention; if it does not, the
explanation is wrong and worth reopening.
The guard in publish-sdk-workflow.test.ts pinned -P 4 and carried the
disproven claim in a comment; both are corrected, and it fails against the
4-wide workflow.
There was a problem hiding this comment.
Review
Clean, minimal fix. Two files changed, both correctly updated.
What changes
.github/workflows/publish-sdk.yml:xargs -P 4toxargs -P 2in the SDK examples gate, with an updated comment explaining thePage.captureScreenshotfailure mode.scripts/__tests__/publish-sdk-workflow.test.ts: guard regex updated from-P 4to-P 2; stale "blocks on live-AI calls, not on cores" comment removed.
Findings
No CRITICAL / HIGH / MEDIUM issues.
LOW — none actionable: The remaining comment in the test (github-actions-runner-labels.test.ts holds this for every workflow) still reads cleanly after the stale sentence was removed; no correction needed.
Checklist
- No literal
::error::/::warning::/::notice::annotation commands introduced. - No untrusted
github.event.*interpolation inrun:blocks. - No secret handling or permission changes.
- Guard test correctly pins the new value and the PR description confirms it fails against the 4-wide workflow and passes at 2-wide.
- No application code touched; no i18n or publishing surface affected.
The tradeoff (slower wall time for gate stability) is correctly acknowledged in the PR description. Approved.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The
E2E — public SDK exampleshard gate failed 2 of 3 runs onubuntu-latest, always on the same example and always the same way:The page had loaded and Chromium then failed to capture it, so this is a capture-time resource failure rather than a test failure. 7 of 8 examples passed in every run.
basicCode was effectively identical across all three — the only
mainmovement between them was the MCP version bump, which does not touch the SDK.Why it regressed
The width was 4 on the 8-vCPU self-hosted runner. I kept it at 4 when moving to
ubuntu-latestin #12, reasoning that these examples block on live-AI calls rather than CPU. That reasoning was incomplete: four concurrent headed Chromium instances also contend for memory and display surfaces, neither of which is a core count. The review on #12 flagged the width as unverified on this runner class, and the single dry run that passed gave false confidence.Neither publish attempt shipped anything — the gate runs before
npm publish, so@shiplightai/sdkstayed at 0.1.11 andmainstayed clean.What this does not establish
Halving the width trades wall time on that step for a gate that holds. It does not identify the exact constraint: there was no OOM line in any log, so memory,
/dev/shmand display surfaces all remain candidates. If 2-wide proves stable, contention is the explanation; if it does not, the explanation is wrong and worth reopening rather than halving again.Test plan
scripts/__tests__/publish-sdk-workflow.test.tspinnedxargs -P 4and carried the disproven "blocks on live-AI calls, not on cores" claim in a comment. Both corrected.publish-sdk.yml: 4 pass, 1 fail) and passes at 2-wide.🤖 Generated with Claude Code