Comprehensive guide to benchmarking, profiling, and optimizing gitbot-fleet performance.
# Run all benchmarks
./scripts/bench-fleet.sh run
# Generate performance report
./scripts/bench-fleet.sh report
# Create flamegraph profile
./scripts/bench-fleet.sh flamegraph
# Run all profiling tools
./scripts/bench-fleet.sh allThe fleet includes comprehensive benchmarks using Criterion.rs:
# Run all benchmarks
cd shared-context
cargo bench
# Run specific benchmark
cargo bench -- context_creation
# Save as baseline
cargo bench -- --save-baseline v0.2.0
# Compare with baseline
cargo bench -- --baseline v0.2.0-
Context Creation - Initialization overhead
-
Bot Registration - Single vs. bulk registration
-
Finding Operations - Adding findings (bulk throughput)
-
Finding Queries - Query performance by bot/category
-
Bot Execution - Start/complete lifecycle
-
Health Checking - Full health check with anomaly detection
-
Report Generation - Markdown/JSON/HTML formatting
-
Serialization - JSON encode/decode performance
Criterion provides several metrics:
-
time: Mean execution time
-
throughput: Operations per second (for bulk tests)
-
R²: Goodness of fit (>0.99 is excellent)
-
outliers: Statistical outliers in measurements
Example output:
context_new time: [125.32 ns 126.89 ns 128.52 ns]
change: [-2.1023% +0.4562% +2.9234%] (p = 0.68 > 0.05)
No change in performance detected.
Generate interactive flamegraph to identify hot paths:
# Install cargo-flamegraph
cargo install flamegraph
# Generate flamegraph
./scripts/bench-fleet.sh flamegraph
# Open flamegraph.svg in browser
firefox flamegraph.svgReading Flamegraphs: - Width = CPU time consumed - Y-axis = call stack depth - Color = randomized (not meaningful) - Click to zoom into specific functions
# Profile memory allocations
./scripts/bench-fleet.sh memory
# View detailed allocation tree
ms_print benchmark-results/massif.out | less| Operation | Target | Current | Status |
|---|---|---|---|
Context creation |
<200ns |
~127ns |
✅ |
Register all bots |
<2µs |
~1.5µs |
✅ |
Add single finding |
<500ns |
~350ns |
✅ |
Add 1000 findings |
<500µs |
~420µs |
✅ |
Query by bot (1000 findings) |
<50µs |
~35µs |
✅ |
Full health check |
<1ms |
~800µs |
✅ |
Generate Markdown report |
<5ms |
~3.2ms |
✅ |
JSON serialization |
<2ms |
~1.5ms |
✅ |
Current: Vec with linear search Optimizations: - Use HashMap for O(1) lookups by ID - BTreeMap for sorted iteration - Consider arena allocation for bulk adds
// Before
findings.iter().find(|f| f.id == target_id)
// After
findings_map.get(&target_id)The get_or_create_context in dashboard clones the entire context:
// Current (expensive)
context.clone().unwrap()
// Optimization: Arc<RwLock<Context>>
Arc::clone(&context.read().await)Markdown formatting is currently string concatenation:
// Before: Multiple allocations
md.push_str(&format!("| {} |", value));
// After: Pre-allocate capacity
let mut md = String::with_capacity(estimated_size);Cache health metrics with TTL:
struct CachedHealth {
health: FleetHealth,
computed_at: Instant,
ttl: Duration,
}
// Only recompute if cache expired
if cached.computed_at.elapsed() > cached.ttl {
cached.health = compute_health();
cached.computed_at = Instant::now();
}Instead of individual health updates:
// Batch multiple updates
let updates = vec![health1, health2, health3];
socket.send(serde_json::to_string(&updates)?).await?;Add indices for common queries:
struct FindingSet {
findings: Vec<Finding>,
by_bot: HashMap<BotId, Vec<usize>>, // Index
by_severity: HashMap<Severity, Vec<usize>>, // Index
by_category: HashMap<String, Vec<usize>>, // Index
}Add to Cargo.toml:
[profile.release]
opt-level = 3
lto = "fat"
codegen-units = 1
panic = "abort"
strip = true# Build for native CPU
RUSTFLAGS="-C target-cpu=native" cargo build --release
# With link-time optimization
RUSTFLAGS="-C target-cpu=native -C link-arg=-fuse-ld=lld" cargo build --releaseFor high-frequency allocations:
use bumpalo::Bump;
let arena = Bump::new();
for _ in 0..1000 {
let finding = arena.alloc(Finding::new(/* ... */));
}
// All deallocated togetheruse prometheus::{Histogram, Counter};
lazy_static! {
static ref HEALTH_CHECK_DURATION: Histogram =
Histogram::new("health_check_duration_seconds", "Health check duration").unwrap();
static ref FINDINGS_ADDED: Counter =
Counter::new("findings_added_total", "Total findings added").unwrap();
}
// Instrument code
let timer = HEALTH_CHECK_DURATION.start_timer();
let health = ctx.health_check();
timer.observe_duration();The dashboard should expose /metrics endpoint:
use prometheus::{Encoder, TextEncoder};
async fn metrics_handler() -> String {
let encoder = TextEncoder::new();
let metric_families = prometheus::gather();
let mut buffer = vec![];
encoder.encode(&metric_families, &mut buffer).unwrap();
String::from_utf8(buffer).unwrap()
}Add to CI/CD pipeline:
# .github/workflows/bench.yml
name: Benchmarks
on: [push, pull_request]
jobs:
benchmark:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dtolnay/rust-toolchain@stable
- run: cargo bench --bench fleet_benchmarks
- uses: benchmark-action/github-action-benchmark@v1
with:
tool: 'criterion'
output-file-path: target/criterion/report/index.htmlSymptoms: Health checks taking >1s
Diagnosis:
# Profile health check specifically
cargo bench -- full_health_check --profile-time 10Common causes: - Too many registered bots - Large number of findings (>10,000) - Anomaly detection overhead
Solutions: - Cache health score (5-second TTL) - Sample findings for anomaly detection - Parallel tier health calculation
Symptoms: RSS >1GB for small repos
Diagnosis:
# Check allocation patterns
heaptrack ./target/release/fleet-dashboardCommon causes: - Context cloning in dashboard - Large finding messages - Report generation keeping strings in memory
Solutions: - Use Arc for shared context - Implement finding message deduplication - Stream reports instead of full generation
Symptoms: >100ms for Markdown generation
Diagnosis:
cargo bench -- generate_markdownCommon causes: - Many findings (>1000) - Complex formatting - String allocation overhead
Solutions: - Pre-allocate String capacity - Use write! macro instead of push_str - Implement pagination
Before deploying to production:
-
❏ Run full benchmark suite
-
❏ Profile with flamegraph
-
❏ Check memory usage under load
-
❏ Test with realistic data volumes
-
❏ Enable release optimizations
-
❏ Configure resource limits
-
❏ Set up performance monitoring
-
❏ Establish baseline metrics
-
❏ Document performance targets
-
❏ Plan for scaling
# Benchmarking
cargo install cargo-criterion
# Profiling
cargo install flamegraph
cargo install cargo-profdata
cargo install heaptrack
# Monitoring
cargo install cargo-watch# Quick benchmark
cargo bench --bench fleet_benchmarks -- --quick
# Benchmark with profiler
cargo bench --bench fleet_benchmarks --profile-time 60
# Compare two branches
git checkout main
cargo bench -- --save-baseline main
git checkout feature
cargo bench -- --baseline main
# Watch for changes and re-benchmark
cargo watch -x 'bench --bench fleet_benchmarks'