Reviewer Round-1 paired controls: presentation-only rewrite invariance, soundness-counterfactual sensitivity, bounded-feedback rounds (paper-derived: arXiv:2609.07713)

Author: Imbad0202Created Sep 15, 2026Updated Sep 15, 2026
Labelsenhancementresearchpaper-derivedstatus/needs-design

Motivation (paper-derived)

Wang, Li et al. (2026), The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing (arXiv:2609.07713v1, survey of 230 sources) collects three findings that bear directly on ARS's simulated review panel:

  • §4.5 (Dycke & Gurevych 2026, experimentally reproduced). 931 variants of 133 accepted AI/NLP papers: 391 edits that break a critical scientific support relation and 540 soundness-neutral controls. Across the tested automated reviewers, soundness-critical edits did not produce statistically significant differences in review aspects, sentiment, or scores compared with surface-level control edits, while the models were sensitive to irrelevant wording changes.
  • §5.2 (Yang et al. 2026c; Baumann et al. 2026, experimentally reproduced). Holding methods, data, figures, equations, and numerical results fixed while iterating presentation against reviewer feedback achieved a 75.1% success rate and +1.21 points on a ten-point scale across three reviewer models; criticisms of unchanged limitations sometimes weakened or disappeared. A content-preserving rewrite alone raised AI-review scores by +0.45 on average.
  • §5.3 / §9.2 (synthesized across the literature). A repeatable evaluator turns its regularities into an optimization signal; static evaluations of AI reviewers overestimate real-world reliability once authors can observe and adapt to the evaluator.

Figures are quoted as the survey reports them; the primary studies were not re-verified for this issue.

What ARS already has (checked on main at b06ceaf)

  • evals/heldout/re_review_persuasion_invariance/ measures Round-2 invariance to Response-Letter rhetoric (P-1) and sensitivity to manuscript substance under an identical letter (P-2). It does not cover Round 1, and v0.1 is archived against re-review contract 1.0.
  • evals/heldout/reviewer_seeded_defects/ measures Round-1 recall on planted defects plus clean-control false findings. It has no paired variant of the same manuscript.
  • evals/heldout/reviewer_calibration/ (#653, blocked by #828) measures a static gold corpus; #648 covers severity-band anchoring. Neither asks whether a judgement moves when only presentation moves.
  • Claim-strength ladder (#569) and token conservation (#570) constrain what a revision may silently change; they say nothing about whether the panel's judgement of unchanged science is stable.

So the gap is a measurement gap, not a missing guard.

Proposed shape: one held-out paired-control seed, no runtime change

A seed under evals/heldout/ reusing the existing synthetic fixtures (reviewer_seeded_defects/manuscripts/ms00_clean_control.md, ms01_quant_defective.md, ms02_qual_defective.md) with three arm families per manuscript:

  1. Presentation-only rewrite (APRES-style): related-work positioning, discussion framing, and limitation wording rewritten; every number, table, method sentence, and citation held byte-identical. Expected: per-seat criterion judgements, severity, and decision invariant; each planted defect still found.
  2. Soundness-breaking edit with surface control (Dycke-style): one support relation broken (e.g., a claim now outruns the reported statistic) paired with a same-length neutral edit elsewhere. Expected: judgement moves in the direction the evidence dictates, and only for the broken relation.
  3. Bounded-feedback rounds (the adaptive case): the rewrite for round k+1 is chosen using the panel's round-k output, with a fixed query budget (e.g., 2 rounds) and no change to the scientific content. Expected: the planted defects and the original limitations remain in the round-k+1 findings.

Each arm carries a human-adjudicated equivalence attestation (that the science did not change, or changed only where stated), per-item expected outcomes anchored to the review criteria, and a heldout-measurement report per evals/heldout/MEASUREMENT_CONTRACT.md. Measure the current model first; any prompt change afterwards is measured against that baseline (the #569/#570 and #574 E4 discipline).

Non-goals

  • No runtime "manipulation detector" and no score correction factor. The paper's own boundary (§5.1) is that a score change with science held fixed shows evaluator sensitivity; it does not by itself distinguish clearer presentation from manipulation.
  • No claim about ARS's reviewer before the seed is run; the result is a measured susceptibility, NOT_RUN until executed.
  • Not a calibration set and not a substitute for #653.

Relations

  • Reuses the #574 E4 harness machinery (record status fields, replicate discipline, raw-output preservation).
  • Not blocked by #828 (no public corpus needed; fixtures are synthetic).
  • Companion measurement: author-identity cue paired controls (opened alongside this issue).
  • Source: dual read of the paper on 2026-09-15 (Claude Fable 5.1 and Codex gpt-6-astra, independent reads, then compared).

Source: Imbad0202/academic-research-skills