[META] Gartenberg et al. 2026 (Org. Sci. 37(3) editorial): "more versus better", journal-side evidence for the failure mode ARS is built to avoid
Paper
Gartenberg, C., Hasan, S., Murray, A., & Pierce, L. (2026). More versus better: Artificial intelligence, incentives, and the emerging crisis in peer review. Organization Science, 37(3), 795-812. https://doi.org/10.1287/orsc.2026.ed.v37.n3
An AI Task Force editorial reporting on every first submission (6,957) and text-format review (10,389) at one management journal from January 2021 to February 2026, scored with a commercial AI-writing classifier (Pangram v3.1) and standard readability indices, with a two-year pre-ChatGPT placebo window. It is observational, single-journal, and aggregate-only by design. The authors state they cannot make normative assessments about appropriate AI use.
Why it matters for ARS
ARS positions itself as a quality-first, human-in-the-loop design, and the README already carries three literature anchors for that stance (The AI Scientist, Zhao et al., Ren et al.). This editorial supplies the first journal-side outcome evidence on the failure mode ARS is built to avoid: AI use that produces more submissions rather than better ones. Four findings map onto ARS mechanisms that either exist without measurement or do not exist:
| Editorial finding (section) | What it reports | ARS mechanism today | Gap |
|---|---|---|---|
| §3.3, Fig. 8: heavily AI-written manuscripts read worse | Flesch Reading Ease 1.28 SD lower by Jan 2026 vs Jan 2021; higher grade level, FOG, SMOG, nominalization, jargon; but less passive voice, less hedging, more specificity | academic-paper/references/writing_quality_check.md (prose rules, no measurement) |
No deterministic readability or nominalization measurement anywhere in the suite → #830 |
| §3.4, Tables 1-3: worse writing and higher AI score each predict desk rejection | Above ~30% AI share, desk rejection is ~30 points higher; readability and AI score are separately predictive; AI use does not help all-non-native-English author teams (interaction ≈ 0) | none | Same as A; this is the outcome evidence that makes A worth doing |
| §4.2, Table 5 and Fig. 12: AI-written reviews are narrower | With manuscript and reviewer fixed effects, AI share shifts review emphasis toward theory (+0.25 SD) and away from data (-0.28 SD); PCA shows a narrower evaluative range | Five-seat panel with role-specific criteria (academic-paper-reviewer) is designed for breadth but breadth is not measured |
No coverage profile on panel outputs → #831 |
| §5.1, §5.4.1: production mode matters and editors want to see it | "Cognitive surrender" (Shaw & Nave 2026) vs. human-first use; detection data should inform triage, not gatekeep; the authors disclose their own production process and classifier score | disclosure mode renders venue statements from use records; #390 block manifest tracks block hashes |
No block-level record of what was model-drafted vs. human-written or human-edited, so a disclosure cannot be evidence-backed → #832 |
| §5.2: "the same dynamic extends to empirical analysis"; craft erosion (Bechky & Davis 2025) | Delegating the struggle erodes the capacity to recognize puzzles | Collaboration Depth Observer (Wang & Zhang 2026 rubric), claim-strength ladder (#569) | Docs anchor only → #833 |
Sub-issues
- A (#830): deterministic writing-surface measurement (readability, nominalization, jargon, hedging) as an advisory on drafts and on reviewer outputs
- B (#831): review topical-coverage profile on panel outputs, and as a measured dimension in the #653 calibration corpus
- C (#832): design: block-level prose provenance so a disclosure statement can be evidence-backed
- D (#833): docs: add the editorial as the fourth human-in-the-loop anchor and record the "more versus better" non-goal in POSITIONING
Boundaries
- Nothing here is an AI-detection feature. ARS does not score its own output for "AI-ness" and does not aim to fool classifiers;
writing_quality_check.mdalready states that boundary and it stays. - The editorial is single-journal and correlational. It motivates measurement inside ARS; it is not evidence that any ARS mechanism works.
- No ranking, gating, or auto-rejection of a user's draft on any of these measures. Advisory rows only, consistent with the Kong L2 advisory-not-generation lesson (
docs/design/2026-06-08-kong-255-l2-advisory-not-generation.md).
Convention follows #255 (Kong et al.).
Source: Imbad0202/academic-research-skills