thorough work, good stuff.. it even runs a selection-bias analysis against their own benchmark and reports that some tasks that were disproportionately hard for a model. Rare to see a benchmark paper attack itself like that.
The comparison between harnesses is very nice. Interesting to see that using a different harness can bump the performance of the model as much as a new version (e.g., GPT 5.5+Codex ~= GPT 5.6+Terminus, at lower cost)
That’s a nice discussion. Some people say that with current model capabilities, the real differentiator is the harness. What are the best harnesses you guys are using?
This approach of not only producing the benchmark tasks, but also focusing on creating a data engine that will improve over time and produce up-to-date tasks that challenge the cutting-edge models is very interesting and valuable.
Interesting the idea of treating the benchmark as an evolving system rather than a static dataset.
thorough work, good stuff.. it even runs a selection-bias analysis against their own benchmark and reports that some tasks that were disproportionately hard for a model. Rare to see a benchmark paper attack itself like that.
The comparison between harnesses is very nice. Interesting to see that using a different harness can bump the performance of the model as much as a new version (e.g., GPT 5.5+Codex ~= GPT 5.6+Terminus, at lower cost)
That’s a nice discussion. Some people say that with current model capabilities, the real differentiator is the harness. What are the best harnesses you guys are using?
Using semantic perturbation to test whether difficulty survives rewording is really smart. Great work!
The methodology was the most interesting part for me. The paper spends as much time explaining how the benchmark was built as the benchmark itself.
Refreshing to see something practical instead of another leaderboard battle. Also, props to the team for being so meticulous.
This approach of not only producing the benchmark tasks, but also focusing on creating a data engine that will improve over time and produce up-to-date tasks that challenge the cutting-edge models is very interesting and valuable.
good work!
nice!
[flagged]
[dead]
[flagged]
[flagged]
[flagged]
[dead]