harness matrix evals
Author: sydney-runkleCreated Sep 11, 2026Updated Sep 16, 2026
Labelsorg:internalpriority:mediumpackage:evalstype:chore
support running our benchmark suite across deepagents + other harnesses ideally these would be regularly run so we can identify areas we need to improve and learn from other open harnesses
accordingly, we should have an up to date benchmarks page on our docs
Source: langchain-ai/deepagents