Clarification on Appendix E benchmark: issue-recall ground truth and compute matching
Context — I'm scoping a small pilot of the Appendix E controlled benchmark with a local paper-reading community. Not a reproduction: roughly six drafts from a single ML subfield, four of the five conditions (single-model self-critique, same-model two-agent, cross-model, cross-model reversed), three blinded raters. The aim is to produce a workable ground-truth procedure, a rating rubric with a measured inter-rater agreement figure, and a realistic cost estimate, so whoever runs the full study starts with those in hand. Noting up front that the repo has clearly developed past the v0.4 snapshot in the paper, so I'd welcome knowing if your thinking here has moved too.
Issue recall — Appendix E names issue recall as a metric, which implies a reference set of issues per draft. How do you envision that set being constructed? Expert-authored or adjudicated issue sets, seeded defects, the union of findings across conditions, or something else? Asking because a union-based reference seems like it could reward conditions that simply produce more candidates, while seeded defects may not resemble the failures that show up in real drafts.
Compute matching — What quantity is intended to be held roughly constant across conditions? Tokens, inference cost, reviewer call count, wall clock, or something else. The effort contract is explicit that reviewer tier never moves with effort, and that pipeline workload and reviewer reasoning depth are separate axes. That separation seems like it would make the choice of matching unit consequential for the comparison.
The forensic review traces in .aris/traces/ look like they'd give a rater panel exactly the artifacts to score — would you see those as a reasonable basis? Happy to share whatever the pilot produces.
Source: wanshuiyin/Auto-claude-code-research-in-sleep