Skip to content

test(e2e): prove a SIGKILLed hop loses nothing - #195

Merged
catinspace-au merged 1 commit into
mainfrom
fix/outage-kill-mode
Oct 1, 2026
Merged

catinspace-au merged 1 commit into
mainfrom
fix/outage-kill-mode

Conversation

@catinspace-au

Copy link
Copy Markdown
Contributor

The outage tests only ever stopped a service gracefully, so a hop always got to drain what it held before it went. A crash gets no drain. This adds a SIGKILL mode and kill tests on both data paths, so "a 2xx means it lands" is checked against a crash, not just a polite stop.

  • outage.method is stop (the default, unchanged) or kill (docker compose kill --signal SIGKILL, then start). It is validated with the other outage fields, and the test header prints it.
  • Four new tests under make test-resilience:
    • loader-kill-kafka -- SIGKILL dfe-loader while it consumes, so it resumes from its committed offset
    • receiver-kill-kafka -- SIGKILL dfe-receiver while it produces
    • loader-kill-grpc -- SIGKILL dfe-loader on the direct receiver -> loader gRPC path
    • transform-vrl-kill-grpc -- SIGKILL dfe-transform-vrl on the gRPC hop
  • A kill that catches nothing in flight proves nothing, so a kill test FAILS on it.
    • For the receiver or a gRPC hop, that means no request was out at the instant the container exited, read off State.FinishedAt.
    • For a Kafka consumer, it means its group had nothing uncommitted just before the kill.
  • outage.workers and outage.interval set the load. The kill tests run 16 workers 0.01s apart so every kill lands with records in flight. The default is still 4 workers 0.5s apart.
  • The landing verdict is unchanged: every record answered 2xx must land, duplicates allowed. It now counts rows per record, so a run reports sent, accepted, landed, lost, duplicate rows, and rows for records that were never accepted.
  • When the killed service is the receiver itself, a request overlapping its own outage may go unanswered. Silence outside that window still fails.
  • On a Kafka profile, every outage now gets the "the load reached Kafka" gate and the consumer catch-up check. Before, only a broker outage got them.
  • Outage-field validation moved into _outage.config_problems, so it is unit-tested in scripts/tests/test_outage.py.

NOT wired into CI. Each test needs a full stack (ClickHouse, the engine, a broker, the apps) and takes about 3 minutes, so they stay opt-in under make test-resilience, like the other outage tests.

Run locally on 2026-10-02 against the 2.2.0-rc.14 pins, one test per invocation:

test hop killed mode caught at the kill sent accepted landed lost duplicate rows landed, never accepted
loader-kill-kafka dfe-loader (Kafka) kill 400 uncommitted records 40670 40670 40670 0 0 0
receiver-kill-kafka dfe-receiver (Kafka) kill 7 requests 59637 25010 25010 0 0 0
loader-kill-grpc dfe-loader (gRPC) kill 16 requests 6826 2996 2996 0 0 0
transform-vrl-kill-grpc dfe-transform-vrl (gRPC) kill 16 requests 5175 564 564 0 0 9
loader-outage-grpc (control) dfe-loader (gRPC) stop -- 880 368 368 0 0 0

dfe.main is a plain MergeTree with non_replicated_deduplication_window at 0, so neither merges nor insert dedup can hide a duplicate row. The duplicate counts are real counts. A direct count() against uniqExact(marker) per run matched every row above.

The 9 rows in transform-vrl-kill-grpc are records the kill caught after they had passed transform-vrl. They landed, but the receiver answered 503, which is the at-least-once trade-off: a client that retries on 503 duplicates them.

The outage tests only ever stopped a service gracefully, so every hop got to drain before it went. A crash gets no drain. outage.method now takes kill (docker compose kill --signal SIGKILL, then start) beside the default stop, and four kill tests cover both data paths: dfe-loader and dfe-receiver on Kafka, dfe-loader and dfe-transform-vrl on direct gRPC.

A kill that catches nothing in flight proves nothing, so a kill test fails on it. For the receiver or a gRPC hop that means no request was out at the instant the container exited (State.FinishedAt). For a Kafka consumer it means its group had nothing uncommitted just before the kill. outage.workers and outage.interval set the load, and the kill tests run 16 workers 0.01s apart so the kill lands with records in flight.

The landing verdict is unchanged, every 2xx must land, but it now counts rows per record, so a run reports sent, accepted, landed, lost, duplicate rows and rows that were never accepted. On a Kafka profile every outage now gets the reached-Kafka gate and the consumer catch-up check, not only a broker outage. Outage-field validation moved into _outage so it is unit-tested.

Not wired into CI: each test needs a full stack. They run under make test-resilience.
@catinspace-au
catinspace-au merged commit c7fd415 into main Oct 1, 2026
7 checks passed
@catinspace-au
catinspace-au deleted the fix/outage-kill-mode branch October 1, 2026 20:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant