[META] Gartenberg et al. 2026 (Org. Sci. 37(3) editorial): "more versus better", journal-side evidence for the failure mode ARS is built to avoid

Author: Imbad0202Created Sep 7, 2026Updated Sep 7, 2026
Labelspaper-derivedepic

Paper

Gartenberg, C., Hasan, S., Murray, A., & Pierce, L. (2026). More versus better: Artificial intelligence, incentives, and the emerging crisis in peer review. Organization Science, 37(3), 795-812. https://doi.org/10.1287/orsc.2026.ed.v37.n3

An AI Task Force editorial reporting on every first submission (6,957) and text-format review (10,389) at one management journal from January 2021 to February 2026, scored with a commercial AI-writing classifier (Pangram v3.1) and standard readability indices, with a two-year pre-ChatGPT placebo window. It is observational, single-journal, and aggregate-only by design. The authors state they cannot make normative assessments about appropriate AI use.

Why it matters for ARS

ARS positions itself as a quality-first, human-in-the-loop design, and the README already carries three literature anchors for that stance (The AI Scientist, Zhao et al., Ren et al.). This editorial supplies the first journal-side outcome evidence on the failure mode ARS is built to avoid: AI use that produces more submissions rather than better ones. Four findings map onto ARS mechanisms that either exist without measurement or do not exist:

Editorial finding (section) What it reports ARS mechanism today Gap
§3.3, Fig. 8: heavily AI-written manuscripts read worse Flesch Reading Ease 1.28 SD lower by Jan 2026 vs Jan 2021; higher grade level, FOG, SMOG, nominalization, jargon; but less passive voice, less hedging, more specificity academic-paper/references/writing_quality_check.md (prose rules, no measurement) No deterministic readability or nominalization measurement anywhere in the suite → #830
§3.4, Tables 1-3: worse writing and higher AI score each predict desk rejection Above ~30% AI share, desk rejection is ~30 points higher; readability and AI score are separately predictive; AI use does not help all-non-native-English author teams (interaction ≈ 0) none Same as A; this is the outcome evidence that makes A worth doing
§4.2, Table 5 and Fig. 12: AI-written reviews are narrower With manuscript and reviewer fixed effects, AI share shifts review emphasis toward theory (+0.25 SD) and away from data (-0.28 SD); PCA shows a narrower evaluative range Five-seat panel with role-specific criteria (academic-paper-reviewer) is designed for breadth but breadth is not measured No coverage profile on panel outputs → #831
§5.1, §5.4.1: production mode matters and editors want to see it "Cognitive surrender" (Shaw & Nave 2026) vs. human-first use; detection data should inform triage, not gatekeep; the authors disclose their own production process and classifier score disclosure mode renders venue statements from use records; #390 block manifest tracks block hashes No block-level record of what was model-drafted vs. human-written or human-edited, so a disclosure cannot be evidence-backed → #832
§5.2: "the same dynamic extends to empirical analysis"; craft erosion (Bechky & Davis 2025) Delegating the struggle erodes the capacity to recognize puzzles Collaboration Depth Observer (Wang & Zhang 2026 rubric), claim-strength ladder (#569) Docs anchor only → #833

Sub-issues

  • A (#830): deterministic writing-surface measurement (readability, nominalization, jargon, hedging) as an advisory on drafts and on reviewer outputs
  • B (#831): review topical-coverage profile on panel outputs, and as a measured dimension in the #653 calibration corpus
  • C (#832): design: block-level prose provenance so a disclosure statement can be evidence-backed
  • D (#833): docs: add the editorial as the fourth human-in-the-loop anchor and record the "more versus better" non-goal in POSITIONING

Boundaries

  • Nothing here is an AI-detection feature. ARS does not score its own output for "AI-ness" and does not aim to fool classifiers; writing_quality_check.md already states that boundary and it stays.
  • The editorial is single-journal and correlational. It motivates measurement inside ARS; it is not evidence that any ARS mechanism works.
  • No ranking, gating, or auto-rejection of a user's draft on any of these measures. Advisory rows only, consistent with the Kong L2 advisory-not-generation lesson (docs/design/2026-06-08-kong-255-l2-advisory-not-generation.md).

Convention follows #255 (Kong et al.).

Source: Imbad0202/academic-research-skills