A rigorous, reproducible, open-source benchmark framework for evaluating event sourcing databases.
This project exists to define a credible performance standard for event stores — one that measures real-world behavior under realistic workloads, not synthetic best-case scenarios.
This project is implemented with Rust and Python.
Clone the project repository from GitHub.
git clone https://github.com/pyeventsourcing/event-store-benchmark.gitInstall the Rust toolchain, the protobuf compiler, and Python 3.11+.
Then, create a Python virtual environment (for report generation) and build the benchmark tool.
For convenience, a Makefile is provided to simplify common tasks.
- Make a Python virtual environment:
make venv - Build the benchmark tool:
make build - Run the 'smoke test' workload:
make run-smoke-test - Run the 'scaling readers' workload:
make run-scaling-readers - Run the 'scaling writers' workload:
make run-scaling-writers - Generate HTML reports:
make report - Read HTML reports: Open
results/published/index.htmlin your brower - Print available Makefile targets:
make help
Most existing benchmarks for event stores:
- Measure only peak append throughput
- Ignore latency percentiles
- Skip recovery and crash behavior
- Do not model realistic workload shapes
- Are difficult to reproduce
- Favor a specific implementation
This project aims to correct that.
We treat benchmarking as an engineering discipline — not a marketing exercise.
This benchmark suite is built around the following principles:
Benchmarks must model real event-sourced applications:
- Many small streams
- Some hot streams
- Heavy-tailed (Zipf-like) distributions
- Tag/category filtering
- Concurrent writers
- Catch-up subscribers
- Mixed read/write workloads
Synthetic “write 1 million events to one stream” tests are insufficient.
We measure latency percentiles using the HDR (high dynamic range) Histogram.
Average throughput alone is misleading.
Latency distribution under contention is what matters.
All benchmarks must be:
- Deterministic (fixed random seeds)
- Configurable via versioned YAML definitions
- Hardware documented
- OS and fsync mode documented
- Repeatable across environments
Raw results must be published alongside summarized results.
The benchmark must not favor a specific implementation.
Adapters are used to interface with different systems, but workloads are defined independently of implementation details.
Benchmark runs capture:
- Throughput: Events per second
- Latency percentiles: p50, p95, p99, p999
- Container metrics: CPU, memory, startup time
- Raw samples: Per-operation timing data
- Environment: Hardware, OS, disk, runtime info
- Reproducibility: Git commit hash, seed, exact config
Each published result must document:
- CPU model
- Core count
- RAM
- Disk type (NVMe, SSD, HDD)
- Filesystem
- OS version
- Fsync configuration
- Kernel tuning (if any)
- Store configuration
Responsible for:
- Event store adaption
- Workload execution
- Raw metrics output
No analysis logic lives in Rust — only measurement.
Event stores are adapted using common Rust traits:
trait StoreManager {
/// Start the container and return success status
async fn start(&mut self) -> anyhow::Result<()>;
/// Stop and cleanup the container
async fn stop(&mut self) -> anyhow::Result<()>;
/// Get the container ID for stats collection (if applicable)
fn container_id(&self) -> Option<String>;
/// Store name (adapter name)
fn name(&self) -> &'static str;
/// Create a new adapter instance (client)
fn create_adapter(&self) -> anyhow::Result<Arc<dyn EventStoreAdapter>>;
}
trait EventStoreAdapter {
/// Append an event
async fn append(&self, events: Vec<EventData>) -> anyhow::Result<()>;
/// Read events
async fn read(&self, req: ReadRequest) -> anyhow::Result<Vec<ReadEvent>>;
}This allows the same workload to run across different systems.
In alphabetical order:
- Axon Server
- EventsourcingDB
- KurrentDB
- UmaDB
To ensure fair and reproducible comparisons between KurrentDB, Axon Server, and UmaDB, the benchmark suite establishes a "level playing ground" by standardizing low-level transport settings and client instantiation strategies.
All three adapters utilize the tonic gRPC library in Rust, and they are configured with identical network and flow control parameters:
- TCP NoDelay: Set to
truefor all clients. This disables Nagle's algorithm, ensuring that small packets (like event append requests) are sent immediately, which is critical for accurate latency measurement. - HTTP/2 Keep-Alive: Configured with a 5-second interval and a 10-second timeout.
- Window Sizes (Flow Control):
- Initial Stream Window Size: 4 MB (
4 * 1024 * 1024). - Initial Connection Window Size: 8 MB (
8 * 1024 * 1024). - These enlarged window sizes prevent the benchmark from being throttled by default small gRPC flow-control limits, allowing higher throughput over single connections.
- Initial Stream Window Size: 4 MB (
The benchmark follows a consistent "One Client Per Worker" model:
- Independent Connections: Each worker task (reader or writer) establishes its own independent gRPC connection during initialization.
- Adapter Instances: When the benchmark starts worker tasks, it calls
create_adapter()for each worker.- For KurrentDB, Axon Server, and UmaDB, this creates a new instance of the respective minimalist gRPC client, which establishes a dedicated connection to the store.
- Concurrency: The number of these connections is strictly controlled by the
concurrencysettings in the benchmark configuration (e.g.,writers: [1, 4]), ensuring that all databases are tested with the same number of active client connections.
- Minimalist Clients: The suite uses "minimal" gRPC client implementations for KurrentDB. This implementation strips away high-level background state machines or complex coordination logic found in the official client SDK, ensuring the benchmark measures the database's performance rather than the client library's overhead. A Rust gRPC client for Axon Server has also been implemented with the same design, because no official SDK exists.
- Standardized Payload: All adapters transform the internal
EventData(binary payload + type + tags) into their respective proto formats just before the gRPC call, keeping the transformation overhead comparable across all tests.
The benchmark supports four workload categories:
Generic event store usage patterns with configurable concurrency and operations:
- Write mode: Concurrent writers appending events
- Read mode: Concurrent readers consuming events
- Mixed mode: Combined read/write operations
Testing persistence guarantees:
- Crash recovery testing
- fsync timing analysis
- WAL replay verification
Testing correctness guarantees:
- Optimistic concurrency conflict detection
- Read-after-write verification
- Event ordering validation
Testing operational characteristics:
- Startup/shutdown performance
- Backup/restore speed
- Storage growth measurement
Each workload is defined by a named YAML file. The name field identifies the workload, and workload_type specifies which implementation to use.
# configs/smoke-test.yaml
name: smoke-test
workload_type: performance
mode: write
duration_seconds: 10
concurrency:
writers: [1, 4]
operations:
write:
event_size_bytes: 256
stores: [umadb, dummy]# configs/scaling/writers.yaml
name: scaling-writers
workload_type: performance
mode: write
duration_seconds: 120
concurrency:
writers: [1, 2, 4, 8, 16, 32]
operations:
write:
event_size_bytes: 256
stores: [umadb, kurrentdb, axonserver, eventsourcingdb]# configs/concurrent_readers.yaml
name: scaling-readers
workload_type: performance
mode: read
duration_seconds: 6
concurrency:
readers: [1, 2, 4, 8, 16, 32]
operations:
write:
event_size_bytes: 256
read:
batch_size: 100
setup:
prepopulate_events: 50000
prepopulate_streams: 5000
stores: [umadb, kurrentdb, axonserver, eventsourcingdb]Responsible for:
- Aggregating benchmark runs
- Computing statistical comparisons
- Plotting latency distributions
- Generating tables for publication
- Producing PDF/HTML reports
- Detecting regressions between runs
Published benchmark reports must include:
- Workload definition
- Raw metrics
- Summary tables
- Latency distribution graphs
- Environment specification
- Exact commit hash of benchmark suite
- Exact version of target system
Transparency is mandatory.
This benchmark suite does not:
- Optimize systems for artificial workloads
- Hide durability settings
- Benchmark in-memory-only configurations
- Publish results without reproducibility metadata
- Declare “winners”
The goal is measurement, not marketing.
Contributions are welcome for:
- New workload definitions
- New system adapters
- Improved statistical analysis
- Improved reporting templates
- Environment automation scripts
All contributions must preserve:
- Determinism
- Reproducibility
- Neutrality
This project aims to become:
- A reference benchmark for event sourcing systems
- A research-grade measurement framework
- A regression detection tool for event store developers
- A shared standard for comparing durability and performance trade-offs
If adopted broadly, this could meaningfully improve the quality of performance claims in the event sourcing ecosystem.
Open source under MIT.
