Reproducible analysis of reasoning-effort saturation on DeepSWE v1.1
-
Updated
Aug 22, 2026 - Python
Reproducible analysis of reasoning-effort saturation on DeepSWE v1.1
DeepSWE v1.1: Perfect score (113/113) — all tasks solved with reward=1.0
DeepSWE Launcher ⚡ lightweight DeepSWE benchmark launcher: batching, resume, multi-model profiles, live dashboard (terminal + WebUI), Excel reports
Daily mirror of the DeepSWE live leaderboard with TrueIQ benchmarking analysis.
To associate your repository with the deepswe topic, visit your repo's landing page and select "manage topics."