About the project
The benchmark becomes a stable lab where techniques can change and be compared without rewriting the runner or binding the system to one backend.
A modular capability harness
for pushing agents beyond
scaling-law limits
# practical pipeline
GAIA Benchmark
-> InspectAI GAIA runner
-> Test-Time Scaling Proxy
-> OpenAI-compatible model backend
-> Local / Colab / Mac / GPU inference
Start with the big picture, then check the results, and move to the technical section when you want to reproduce or extend the experiment.
The benchmark becomes a stable lab where techniques can change and be compared without rewriting the runner or binding the system to one backend.
The main signal is already visible because some test-time techniques improve quality, but the gain has to be read together with time, tokens, and run stability.
The architecture separates research logic from run infrastructure, so a new technique can be plugged in as a module and reproduced across heterogeneous compute.
These two charts give a quick read: what the baseline looks like across the full task set, and which techniques sit on the accuracy-versus-cost Pareto front.