test(e2e): prove a SIGKILLed hop loses nothing - #195
Merged
Merged
Conversation
The outage tests only ever stopped a service gracefully, so every hop got to drain before it went. A crash gets no drain. outage.method now takes kill (docker compose kill --signal SIGKILL, then start) beside the default stop, and four kill tests cover both data paths: dfe-loader and dfe-receiver on Kafka, dfe-loader and dfe-transform-vrl on direct gRPC. A kill that catches nothing in flight proves nothing, so a kill test fails on it. For the receiver or a gRPC hop that means no request was out at the instant the container exited (State.FinishedAt). For a Kafka consumer it means its group had nothing uncommitted just before the kill. outage.workers and outage.interval set the load, and the kill tests run 16 workers 0.01s apart so the kill lands with records in flight. The landing verdict is unchanged, every 2xx must land, but it now counts rows per record, so a run reports sent, accepted, landed, lost, duplicate rows and rows that were never accepted. On a Kafka profile every outage now gets the reached-Kafka gate and the consumer catch-up check, not only a broker outage. Outage-field validation moved into _outage so it is unit-tested. Not wired into CI: each test needs a full stack. They run under make test-resilience.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The outage tests only ever stopped a service gracefully, so a hop always got to drain what it held before it went. A crash gets no drain. This adds a SIGKILL mode and kill tests on both data paths, so "a 2xx means it lands" is checked against a crash, not just a polite stop.
outage.methodisstop(the default, unchanged) orkill(docker compose kill --signal SIGKILL, then start). It is validated with the other outage fields, and the test header prints it.make test-resilience:loader-kill-kafka-- SIGKILL dfe-loader while it consumes, so it resumes from its committed offsetreceiver-kill-kafka-- SIGKILL dfe-receiver while it producesloader-kill-grpc-- SIGKILL dfe-loader on the direct receiver -> loader gRPC pathtransform-vrl-kill-grpc-- SIGKILL dfe-transform-vrl on the gRPC hopState.FinishedAt.outage.workersandoutage.intervalset the load. The kill tests run 16 workers 0.01s apart so every kill lands with records in flight. The default is still 4 workers 0.5s apart._outage.config_problems, so it is unit-tested inscripts/tests/test_outage.py.NOT wired into CI. Each test needs a full stack (ClickHouse, the engine, a broker, the apps) and takes about 3 minutes, so they stay opt-in under
make test-resilience, like the other outage tests.Run locally on 2026-10-02 against the 2.2.0-rc.14 pins, one test per invocation:
dfe.mainis a plainMergeTreewithnon_replicated_deduplication_windowat 0, so neither merges nor insert dedup can hide a duplicate row. The duplicate counts are real counts. A directcount()againstuniqExact(marker)per run matched every row above.The 9 rows in
transform-vrl-kill-grpcare records the kill caught after they had passed transform-vrl. They landed, but the receiver answered 503, which is the at-least-once trade-off: a client that retries on 503 duplicates them.