Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
135 changes: 116 additions & 19 deletions .github/workflows/rust-benchmark.yml
Original file line number Diff line number Diff line change
Expand Up @@ -3,14 +3,18 @@ name: Rust Benchmark
on:
push:
branches: [main, master]
paths: ['rust/**', 'spacetime-module/**', '.github/workflows/rust-benchmark.yml']
paths: ['rust/**', '.github/workflows/rust-benchmark.yml']
pull_request:
branches: [main, master]
paths: ['rust/**', 'spacetime-module/**', '.github/workflows/rust-benchmark.yml']
paths: ['rust/**', '.github/workflows/rust-benchmark.yml']

env:
CARGO_TERM_COLOR: always
RUST_BACKTRACE: 1
# Pinned nightly, kept in sync with rust/rust-toolchain.toml. The patched
# doublets crates rely on unstable APIs, so a rolling nightly breaks the build
# without any change in this repository.
toolchain: nightly-2026-04-14

defaults:
run:
Expand All @@ -28,10 +32,10 @@ jobs:
steps:
- uses: actions/checkout@v4

- name: Setup Rust (nightly)
- name: Setup Rust (pinned nightly)
uses: dtolnay/rust-toolchain@master
with:
toolchain: nightly
toolchain: ${{ env.toolchain }}
components: rustfmt, clippy
targets: wasm32-unknown-unknown

Expand Down Expand Up @@ -89,6 +93,27 @@ jobs:
# Run tests sequentially to avoid parallel interference with shared SpacetimeDB state.
run: cargo test -- --test-threads=1

# Unit tests for the reporting pipeline (rust/out.py). They are cheap and must
# pass before a benchmark is run, otherwise a 40 minute benchmark could finish
# only to fail while publishing its results.
results-pipeline:
name: Results pipeline tests
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4

- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.11'

- name: Install Python dependencies
run: pip install matplotlib numpy

- name: Run results pipeline tests
run: python3 -m unittest test_out -v

# Quick benchmark validation for pull requests.
# Runs benchmarks with reduced scale to verify they work and produce results
# in well under 10 minutes. Results are not committed but uploaded as artifacts.
Expand All @@ -103,16 +128,21 @@ jobs:
benchmark-pr:
name: Benchmark (PR validation)
runs-on: ubuntu-latest
needs: [test]
needs: [test, results-pipeline]
if: github.event_name == 'pull_request'
timeout-minutes: 20
steps:
- uses: actions/checkout@v4

- name: Setup Rust (nightly)
- name: Setup Rust (pinned nightly)
uses: dtolnay/rust-toolchain@master
with:
toolchain: nightly
toolchain: ${{ env.toolchain }}
# rust/rust-toolchain.toml requires these components. Installing them
# with the toolchain avoids rustup adding them lazily on the first
# cargo invocation, which fails on the runner image with
# "detected conflict: 'bin/cargo-fmt'".
components: rustfmt, clippy
targets: wasm32-unknown-unknown

- name: Setup Python
Expand Down Expand Up @@ -173,6 +203,12 @@ jobs:
SPACETIMEDB_URI: http://localhost:3000
SPACETIMEDB_DB: benchmark-links
run: |
set -o pipefail
# Criterion compares against target/criterion/<id>/<size>/base and prints
# the failure to *stdout* when that directory was restored incompletely by
# the cache, which splices "Criterion.rs ERROR: ..." into the middle of the
# bencher records (see CI run 35028280108). Start from a clean state.
rm -rf target/criterion
cargo bench --bench bench -- \
--output-format bencher \
--sample-size 10 \
Expand All @@ -181,15 +217,33 @@ jobs:
--nresamples 1000 \
| tee out.txt

- name: Generate charts
run: python3 out.py
- name: Generate results table and charts
env:
BENCHMARK_LINK_COUNT: 10
BACKGROUND_LINK_COUNT: 30
run: python3 out.py out.txt --results results.md

- name: Publish results to the job summary
run: |
{
echo "## Benchmark results (PR validation, reduced scale)"
echo
cat results.md
echo
echo "_Reduced scale: these numbers only prove the benchmark runs;"
echo "the published results come from the full run on \`main\`._"
} >> "$GITHUB_STEP_SUMMARY"

- name: Upload PR benchmark artifacts
# Keep the raw output even when the results step fails,
# so a broken run can be diagnosed from the artifact.
if: always()
uses: actions/upload-artifact@v4
with:
name: benchmark-results-pr
path: |
rust/out.txt
rust/results.md
rust/bench_rust.png
rust/bench_rust_log_scale.png

Expand All @@ -205,19 +259,27 @@ jobs:
benchmark:
name: Benchmark (full)
runs-on: ubuntu-latest
needs: [test]
needs: [test, results-pipeline]
if: github.event_name == 'push' && (github.ref == 'refs/heads/main' || github.ref == 'refs/heads/master')
timeout-minutes: 180
# Required to push the regenerated README table and docs/benchmarks/ charts back.
permissions:
contents: write
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
token: ${{ secrets.GITHUB_TOKEN }}

- name: Setup Rust (nightly)
- name: Setup Rust (pinned nightly)
uses: dtolnay/rust-toolchain@master
with:
toolchain: nightly
toolchain: ${{ env.toolchain }}
# rust/rust-toolchain.toml requires these components. Installing them
# with the toolchain avoids rustup adding them lazily on the first
# cargo invocation, which fails on the runner image with
# "detected conflict: 'bin/cargo-fmt'".
components: rustfmt, clippy
targets: wasm32-unknown-unknown

- name: Setup Python
Expand Down Expand Up @@ -278,31 +340,66 @@ jobs:
SPACETIMEDB_URI: http://localhost:3000
SPACETIMEDB_DB: benchmark-links
run: |
set -o pipefail
# Criterion compares against target/criterion/<id>/<size>/base and prints
# the failure to *stdout* when that directory was restored incompletely by
# the cache, which splices "Criterion.rs ERROR: ..." into the middle of the
# bencher records (see CI run 35028280108). Start from a clean state.
rm -rf target/criterion
cargo bench --bench bench -- \
--output-format bencher \
--sample-size 20 \
--nresamples 10000 \
| tee out.txt

- name: Generate charts
run: python3 out.py
# Regenerates results.md, both charts, copies the charts into docs/benchmarks/ and
# replaces the results section of README.md, so the numbers are readable
# in the repository without running the benchmark locally.
- name: Generate results table and charts
env:
BENCHMARK_LINK_COUNT: 1000
BACKGROUND_LINK_COUNT: 3000
run: |
python3 out.py out.txt \
--results results.md \
--readme ../README.md \
--docs-dir ../docs/benchmarks

- name: Publish results to the job summary
run: |
{
echo "## Benchmark results (full scale)"
echo
cat results.md
} >> "$GITHUB_STEP_SUMMARY"

- name: Configure git
run: |
git config user.name "github-actions[bot]"
git config user.email "github-actions[bot]@users.noreply.github.com"
git config user.email "linksplatform@gmail.com"
git config user.name "LinksPlatformBencher"

- name: Commit benchmark results
working-directory: .
run: |
git add -f out.txt bench_rust.png bench_rust_log_scale.png 2>/dev/null || true
git diff --staged --quiet || git commit -m "chore: update benchmark results [skip ci]"
git push
# The charts live in docs/benchmarks/ only; the copies in rust/ are build
# output and are kept as workflow artifacts instead of being committed twice.
git add docs/benchmarks README.md rust/results.md rust/out.txt
if git diff --staged --quiet; then
echo "No changes to commit"
else
git commit -m "Update benchmark results [skip ci]"
git push origin HEAD:${GITHUB_REF_NAME}
fi

- name: Upload benchmark artifacts
# Keep the raw output even when the results step fails,
# so a broken run can be diagnosed from the artifact.
if: always()
uses: actions/upload-artifact@v4
with:
name: benchmark-results
path: |
rust/out.txt
rust/results.md
rust/bench_rust.png
rust/bench_rust_log_scale.png
73 changes: 60 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,9 +40,42 @@ Each benchmark iteration pre-populates the database with background links to sim

## Results

> _Benchmark results will be automatically generated and committed here by CI when changes are merged to main._
The numbers below represent the amount of time (ns) a single benchmark iteration takes.

<!--RESULTS_TABLE_PLACEHOLDER-->
- The first chart shows time in a pixel (linear) scale. Doublets bars are drawn with a
minimum visible width, otherwise they would not be visible next to SpacetimeDB.
- The second chart shows time in a logarithmic scale, to see the difference clearly,
because it is around 3-5 orders of magnitude.

Charts and the table are recalculated by the
[Rust Benchmark workflow](.github/workflows/rust-benchmark.yml) on every push to `main`
and committed back to this repository, so the results are visible here without running
the benchmark locally.

### Rust

![Image of Rust benchmark (pixel scale)](https://github.com/linksplatform/Comparisons.SpacetimeDBVSDoublets/blob/main/docs/benchmarks/bench_rust.png?raw=true)
![Image of Rust benchmark (log scale)](https://github.com/linksplatform/Comparisons.SpacetimeDBVSDoublets/blob/main/docs/benchmarks/bench_rust_log_scale.png?raw=true)

### Raw benchmark results (all numbers are in nanoseconds)

<!--BENCHMARK_RESULTS_START-->
_No benchmark results have been published yet. They are generated by the first
[Rust Benchmark](.github/workflows/rust-benchmark.yml) run on `main`._
<!--BENCHMARK_RESULTS_END-->

Each Doublets cell is annotated with how many times faster (or slower) it is than
SpacetimeDB for the same operation.

## Conclusion

Doublets is an embedded store: an operation is a few pointer dereferences and tree
rotations in memory (or in a memory-mapped file), while every SpacetimeDB operation is a
reducer call over a WebSocket connection to a separate process, and every query is served
from the client-side subscription cache. The measured difference is dominated by that
architectural difference rather than by the data structures themselves.

To get fresh numbers, please fork the repository and rerun the benchmark in GitHub Actions.

## Operation Complexity

Expand All @@ -68,7 +101,7 @@ The algorithmic complexity is the same for volatile and non-volatile Doublets va

### Prerequisites

- Rust nightly (see `rust/rust-toolchain.toml`)
- Rust nightly, pinned in `rust/rust-toolchain.toml` (`rustup` installs it automatically)
- SpacetimeDB CLI: `curl -sSf https://install.spacetimedb.com | sh`

### Start SpacetimeDB server and publish module
Expand All @@ -78,8 +111,8 @@ The algorithmic complexity is the same for volatile and non-volatile Doublets va
spacetime start &

# Build and publish the links module
spacetime build --project-path spacetime-module
spacetime publish --project-path spacetime-module benchmark-links
spacetime build --project-path rust/spacetime-module
spacetime publish --project-path rust/spacetime-module --yes benchmark-links
```

### Run benchmarks
Expand All @@ -96,8 +129,13 @@ BENCHMARK_LINK_COUNT=10 BACKGROUND_LINK_COUNT=100 \
SPACETIMEDB_URI=http://localhost:3000 SPACETIMEDB_DB=benchmark-links \
cargo bench --bench bench

# Generate charts from results
python3 out.py
# Generate the results table and charts from out.txt
python3 out.py out.txt --results results.md

# Regenerate everything the CI publishes: results.md, docs/benchmarks/ charts
# and the results section of README.md
python3 out.py out.txt --results results.md --readme ../README.md \
--docs-dir ../docs/benchmarks
```

### Run tests
Expand All @@ -113,23 +151,32 @@ SPACETIMEDB_URI=http://localhost:3000 SPACETIMEDB_DB=benchmark-links cargo test
cd rust
cargo fmt --all
cargo clippy --all-targets

# Unit tests for the results reporting pipeline (no benchmark run required)
python3 -m unittest test_out -v
```

## Project Structure

```
.
├── spacetime-module/ # SpacetimeDB WASM module (links table + reducers)
│ ├── Cargo.toml
│ └── src/
│ └── lib.rs # Table definition and reducers using `spacetimedb` crate
├── docs/
│ └── benchmarks/ # Benchmark charts published by CI and shown above
│ ├── bench_rust.png
│ └── bench_rust_log_scale.png
├── rust/
│ ├── spacetime-module/ # SpacetimeDB WASM module (links table + reducers)
│ │ ├── Cargo.toml
│ │ └── src/
│ │ └── lib.rs # Table definition and reducers using `spacetimedb` crate
│ ├── Cargo.toml # Package manifest with pinned dependencies
│ ├── doublets-patched/ # Local patches to doublets-rs for modern nightly compatibility
│ │ └── PATCHES.md # Documents why patches are needed and what was changed
│ ├── rust-toolchain.toml # Pinned Rust nightly toolchain
│ ├── rustfmt.toml # Rust formatting config
│ ├── out.py # Chart generation script (matplotlib)
│ ├── out.py # Results table, charts and README update
│ ├── test_out.py # Unit tests for out.py
│ ├── results.md # Generated results table (committed by CI)
│ ├── src/
│ │ ├── lib.rs # Links trait, constants (BENCHMARK_LINK_COUNT, BACKGROUND_LINK_COUNT)
│ │ ├── module_bindings/ # spacetimedb-sdk client bindings for the links module
Expand All @@ -145,7 +192,7 @@ cargo clippy --all-targets
│ └── bench.rs # Criterion benchmark suite (7 operations x 5 backends)
└── .github/
└── workflows/
└── rust-benchmark.yml # CI: test on 3 OS + benchmark + chart generation
└── rust-benchmark.yml # CI: test on Linux/macOS, benchmark, publish results
```

## License
Expand Down
21 changes: 21 additions & 0 deletions changelog.d/20260915_223000_benchmark_results_publication.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
---
bump: patch
---

### Fixed
- Pinned the Rust nightly toolchain (`nightly-2026-04-14`), unbreaking the Rust Benchmark workflow that had been failing on every branch since April with `E0512` in `ethnum` and `E0277` in the patched `doublets` dev dependency
- Added `set -o pipefail` around `cargo bench ... | tee out.txt`, so a failing benchmark is no longer masked by `tee`
- The results parser now tolerates Criterion's own error messages, which it prints to stdout in the middle of a bencher record; a stale `target/criterion` baseline used to split every record over two lines and fail the run with `No benchmark data found in out.txt`
- The benchmark steps reset `target/criterion` before running, so a partially restored cache cannot pollute the benchmark output

### Added
- Benchmark results are now published automatically: `rust/out.py` writes `rust/results.md`, copies both charts into `docs/benchmarks/` and replaces the results section of `README.md`, which CI commits back to `main`
- `rust/test_out.py` — unit tests for the results reporting pipeline, run by a dedicated `results-pipeline` CI job that gates the benchmark jobs
- Benchmark results are written to the GitHub Actions job summary for both the pull request and the full run

### Changed
- `rust/out.py` now reports all five benchmarked backends (the two NonVolatile Doublets variants were previously missing) and annotates every Doublets result relative to the SpacetimeDB baseline
- `README.md` documents the results in the same style as the sibling Neo4j and PostgreSQL comparisons

### Removed
- `rust/rust_out` — a 4.2 MB compiled binary committed by accident
Loading
Loading