#
swebench
Here are 5 public repositories matching this topic...
Toolkit for measuring Claude Code and Codex performance over time against a baseline using SWEbench-lite dataset **No API key required for Max or Pro subscribers**
-
Updated
Nov 22, 2025 - Python
Wrapper of common LLM evaluation frameworks
evaluation artificial-intelligence llm lm-evaluation-harness vllm lighteval openai-compatible swebench
-
Updated
Apr 2, 2026 - Python
Autonomous coding loop engine — solo worker or multi-agent team. GLM-5.2[1m], MCP search, mechanical gate, human-in-the-loop.
-
Updated
Jun 26, 2026 - Shell
DeepSWE v1.1: Perfect score (113/113) — all tasks solved with reward=1.0
-
Updated
Aug 13, 2026 - Python
Add this topic to your repo
To associate your repository with the swebench topic, visit your repo's landing page and select "manage topics."