Skip to content

fix: keep runtime alive during channel recovery - #25

Draft
tylerslaton wants to merge 1 commit into
mainfrom
fix/channel-startup-recovery
Draft

fix: keep runtime alive during channel recovery#25
tylerslaton wants to merge 1 commit into
mainfrom
fix/channel-startup-recovery

Conversation

@tylerslaton

Copy link
Copy Markdown
Contributor

What changed

  • If the initial 15-second Channel readiness wait expires while the runtime is still connecting or reconnecting, start the HTTP server so Railway health checks keep the process alive.
  • Continue waiting for Channel readiness in the background.
  • Keep permanent initial failures terminal, and shut down if background recovery later becomes terminal.

Why

A runtime that starts during a realtime gateway outage currently exits before a retry can connect after the gateway recovers. Keeping the HTTP health route alive lets the CopilotKit retry loop eventually restore the managed Channel without a process restart.

Dependency

Depends on CopilotKit/CopilotKit#6347 and a CopilotKit release that includes it. That change retries transient initial gateway activation with exponential backoff.

Validation

  • pnpm exec vitest run app/server-recovery.test.ts app/server.test.ts --reporter=verbose — 6 tests passed
  • git diff --check

Existing repository failures

  • pnpm check-types fails unchanged on the clean base checkout because @copilotkit/channels 0.6.1 and the runtime transitive channels-core 0.6.0 expose incompatible types.
  • The full test suite also has existing failures from the same package skew. The two server test files changed or exercised here pass.

AlemTuzlak added a commit to CopilotKit/CopilotKit that referenced this pull request Aug 4, 2026
## What changed

- Mark initial gateway HTTP 5xx and transient transport failures as
retryable.
- Retry initial managed Channel activation with exponential backoff from
1 second to a 30-second cap until it connects or the manager stops.
- Preserve retry hints from `gateway_draining` join replies and retry
initial join timeouts.
- Keep HTTP 4xx and NXDOMAIN failures terminal.
- Back off established-session outage reminders from 30 seconds to a
15-minute cap while Phoenix continues reconnecting.

## Why

The OpenTag Railway runtime saw the gateway host return HTTP 502 during
an outage. Established Phoenix sessions keep retrying, but a runtime
that starts during the outage stops after its one initial connect
window. It cannot recover when the gateway comes back unless the process
restarts. Fixed 30-second reminder logs also flood long outages.

The gateway drain work now rejects new joins with a structured retryable
response. The client must preserve that response so the runtime can
retry instead of leaving the Channel in a terminal error state.

## Companion change

CopilotKit/OpenTag#25 keeps the Railway HTTP server alive while an
initial Channel retry is pending. OpenTag must consume a CopilotKit
release containing this PR before that companion change can recover by
itself.

## Validation

- `pnpm nx run-many -t test,check-types,build -p
@copilotkit/runtime,@copilotkit/channels-intelligence`
- `pnpm nx run-many -t publint,attw -p
@copilotkit/runtime,@copilotkit/channels-intelligence`
- pre-commit tests and package checks for all affected projects
- `pnpm exec oxfmt --check` on all five changed files
- `pnpm exec oxlint` on all five changed files
- `git diff --check`
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant