Overview

CAPA

A modular capability harness
for pushing agents beyond
scaling-law limits

# practical pipeline
GAIA Benchmark
  -> InspectAI GAIA runner
  -> Test-Time Scaling Proxy
  -> OpenAI-compatible model backend
  -> Local / Colab / Mac / GPU inference

Project Map

Start with the big picture, then check the results, and move to the technical section when you want to reproduce or extend the experiment.

About

About the project

The benchmark becomes a stable lab where techniques can change and be compared without rewriting the runner or binding the system to one backend.

Project Results

Experiment Results

The main signal is already visible because some test-time techniques improve quality, but the gain has to be read together with time, tokens, and run stability.

Results

These two charts give a quick read: what the baseline looks like across the full task set, and which techniques sit on the accuracy-versus-cost Pareto front.

Baseline on 53 tasks: accuracy, time, and token usage per task
Baseline over all 53 GAIA Level 1 tasks with correctness, time, and token usage. Dashed lines show mean values.
Feature efficiency quadrant: accuracy gain vs token and time overhead
Pareto front of accuracy gain versus token/time overhead. Techniques toward the upper left give more accuracy for less cost.