I just released ab-lab. It is a Python tool that lets you run benches with a specific model/harness/setting configuration.
For example, you can do:
python3 -m lab run ./benchmarks/swe-milestone-scikit-learn-light.json \
--out runs/bench33 \
--harness codex \
--model gpt-6-luna \
--effort xhigh \
--preset all --off C08,C16,C17 \
--review-priorities P0,P1 \
--seconds 14400 \
--max-raw 50000000 \
--max-turns 100 \
--max-review-loops 3 \
--parallel 4
The above command will run the scikitlight bench using codex with gpt-6-luna xhigh. It
will allow a max of 3 review loops, and only P0 and P1 issues will be fixed. It adds a safety net, so
after either 4 hours of running, 50M total tokens used, or 100 turns, the runner will stop. The run
result will be saved in ./runs/bench33. The last option tells the runner that we want to run 4
instances of the same bench.
The above command also instructs the runner to run the bench with specific conditions, in
particular all conditions but C08,C16,C17.
For now, I support codex and pi (1.0). I plan to add support for other harnesses. Benches are defined in https://github.com/Fi3/ab-lab/tree/master/benchmarks and are composed of a repo URL, a commit, a set of prompts for the agent, and a set of tests to evaluate the output.
This tool can be used as a benchmark for any repo that you own. You need to define a specific set of tests over a specific commit of that repo, then you can run this tool every time you change the harness or you want to try new SKILLS or a different model, etc.
I’m randomly testing different configurations to see if I can find any big token difference that can’t be explained only by variance. For now, the results are not very interesting. The first bench that I added was slope bench, but it was too easy. Then I tried swe-milestone-scikit-learn, and it seemed to be a little bit more challenging for the agents, but it was using too many tokens, so I ended up running most of the experiments on a reduced version of swe-milestone-scikit-learn that I called swe-milestone-scikit-learn-light. You can find the results of all these experiments in results.md. The output, other than with the specific bench evaluator, is also rated with the scb-check.
For now, from the bench data, we can’t see any improvement in solving a bench using an “author/reviewer”
loop vs using only one agent. (We don’t let the agent spawn sub-agents unless in native mode, so if you set
--max-review-loops=0, this is exactly what you get.)
This is in disagreement with my experience working with agents on real-world issues. Most of the
time, if the issue is not super-simple, doing a review with a new agent that does not have the issue
description / requirements in the context improves the output a lot. Right now the tool send
feature requested to the reviewer but I will fix it soon removing that or making it optional.
The next step would be to find a bench where I can observe it.