Evaluating Custom Reports with run_evaluate.py

Author: guijuzhejiangCreated Jan 23, 2026Updated Jan 23, 2026

I noticed that the python tests/run_evaluate.py script can be used to evaluate the "Deep Research Bench" dataset. However, I would like to evaluate a custom report I generated. Can run_evaluate.py be used for this purpose as well? If so, how should I modify the dataset_name to accommodate my custom report?

Source: langchain-ai/open_deep_research