Interesting benchmark, but I think the main conclusion (agents write low-quality code) is not well supported by the setup. A few points:
-
Between sessions, there is no handoff doc or harness-level memory, so even for human it is hard for them to understand and improve a fresh codebase. The baseline prompt tells the agent "Implement a program that 100% solves the specification. That is all you need to do." At C1 there is no signal that later checkpoints exist. This also does not imply to the model a note/handoff is needed.
-
The only pressure for the model is correctness (even pushed to 100% in the prompt). That combination naturally pushes toward defensive, edge-case-heavy code. Without tests, model also cannot know what is enough and what other cases there is. Coding without testing also does not mimic real world coding agent usage. Even with anti-slop prompt, the model is still being pushed to finish the spec sheet and it only shows how the model acts at somewhat conflicting prompt.
-
The quality metrics have two separate problems. First, I doubt that these hard-coded rules can actually capture verbosity and erosion. Secondly, even if they do capture these things, the verbosity might not be a bad thing, they are argubly just a style choice (e.g., list comprehension vs filter).
-
And lastly, a lot of these make the code taste worse, but not really harming the accuracy. There are many SKILLs out there that can help you refactor your code, just like human code needs refactoring as well. Humans do that by themselves on the way; maybe agents don't: they just want to finish the current task asap (as the prompt stated). The sampled commits from human repo where spanning long time frame, which makes it almost ceterinly a refactor or rewrite happened somewhere.
I think the benchmark is interesting and worth noticing, but I think it tells us more about how to use these tools, instead of how good are these tools.
Also,
problems that did not meaningfully test design decisions, or that frontier agents could solve in a single shot, were removed from the pool.
This is a bit weird, it is like you filtered out the easy questions, and then claiming that they cannot solve them. It is like asking people who did not catch their flight if they failed to board.
This is a bit weird, it is like you filtered out the easy questions, and then claiming that they cannot solve them. If you filter out the questions agent cannot solve in one shot, it means they will keep refining it and make it bolt.
Interesting benchmark, but I think the main conclusion (agents write low-quality code) is not well supported by the setup. A few points:
Between sessions, there is no handoff doc or harness-level memory, so even for human it is hard for them to understand and improve a fresh codebase. The baseline prompt tells the agent "Implement a program that 100% solves the specification. That is all you need to do." At C1 there is no signal that later checkpoints exist. This also does not imply to the model a note/handoff is needed.
The only pressure for the model is correctness (even pushed to 100% in the prompt). That combination naturally pushes toward defensive, edge-case-heavy code. Without tests, model also cannot know what is enough and what other cases there is. Coding without testing also does not mimic real world coding agent usage. Even with anti-slop prompt, the model is still being pushed to finish the spec sheet and it only shows how the model acts at somewhat conflicting prompt.
The quality metrics have two separate problems. First, I doubt that these hard-coded rules can actually capture verbosity and erosion. Secondly, even if they do capture these things, the verbosity might not be a bad thing, they are argubly just a style choice (e.g., list comprehension vs filter).
And lastly, a lot of these make the code taste worse, but not really harming the accuracy. There are many SKILLs out there that can help you refactor your code, just like human code needs refactoring as well. Humans do that by themselves on the way; maybe agents don't: they just want to finish the current task asap (as the prompt stated). The sampled commits from human repo where spanning long time frame, which makes it almost ceterinly a refactor or rewrite happened somewhere.
I think the benchmark is interesting and worth noticing, but I think it tells us more about how to use these tools, instead of how good are these tools.
Also,
This is a bit weird, it is like you filtered out the easy questions, and then claiming that they cannot solve them. It is like asking people who did not catch their flight if they failed to board.
This is a bit weird, it is like you filtered out the easy questions, and then claiming that they cannot solve them. If you filter out the questions agent cannot solve in one shot, it means they will keep refining it and make it bolt.