#6266·deepagents

harness matrix evals

Author: sydney-runkleCreated Sep 11, 2026Updated Sep 16, 2026
Labelsorg:internalpriority:mediumpackage:evalstype:chore

support running our benchmark suite across deepagents + other harnesses ideally these would be regularly run so we can identify areas we need to improve and learn from other open harnesses

accordingly, we should have an up to date benchmarks page on our docs