#155·SkillOpt

Proposal: AgenticReplayBackend — replay mined tasks with a real agentic CLI in a throwaway worktree, gate on the diff

Author: AlphaxalchemyCreated Jul 19, 2026Updated Sep 18, 2026

Motivation

Replay currently scores a single text completion: the candidate skill text + memory + task description go in as one prompt, one block of text comes out, and it is checked programmatically. (The only tool access is the small sanctioned attempt_with_tools loop for tool_called checks.)

That is honest for text-expressible behaviors — formatting, structure, phrasing, conventions — but it structurally cannot evaluate agentic behaviors: "run the tests, read the failure, fix the config" needs an environment that reacts. Much of what users actually want improved in coding agents lives in that multi-step space, and it also connects to #154: agentic tasks are exactly the ones whose success can't be expressed as a regex.

Proposal: AgenticReplayBackend

A replay backend that drives a real open agentic CLI, following the existing backend/attempt pattern:

  1. Give the CLI the mined task in a throwaway git worktree/branch (sandboxed, disposable).
  2. Let it actually run tools and edit files.
  3. Gate on the diff rather than on output text: does it build? do tests pass? does the diff resemble the change the user eventually accepted in the original session?
  4. Surface the diff in the report/dashboard for human review.

This gives real ground truth with no simulator-fidelity questions. (The heavier alternative — a language world model like Qwen-AgentWorld simulating environment feedback — is interesting for offline multi-step replay but adds a large unvalidated link to the evidence chain; probably a research direction rather than a v-next feature.)

Cost alignment

An obvious objection: SkillOpt's premise is cheap, token-efficient improvement of frozen agents, and a real agentic replay is orders of magnitude heavier than a single text completion (full CLI session, tool calls, builds, test runs). To be clear, this proposal is not "make every replay agentic." It's a tiered gate:

  1. Cheap text replay stays the default for everything text-expressible — no change for the vast majority of candidates.
  2. Agentic replay is invoked only for the candidate-rule category that cheap replay structurally cannot validate (the agentic behaviors from #154), and only for candidates that have already survived the cheap filters.
  3. It runs under an explicit per-night budget (e.g., at most N agentic replays per night, token-capped), opt-in per user.

Two further points on alignment with the original premise:

  • SkillOpt's cheapness claim is fundamentally about the deployed artifact: a small text file improving a frozen model at zero inference-time cost. This proposal doesn't touch that — the skill that comes out is exactly as cheap to deploy. The added cost is one-time validation expense, amortized over every future session, which is consistent with the existing design (the optimizer itself already uses a frontier model).
  • The pitch is to spend expensive validation only where cheap validation is worthless. A wrong-but-accepted agentic rule costs more — in degraded future sessions — than the validation would have. An unvalidated rule that passes a regex gate but breaks real workflows is the truly expensive outcome.

Bigger picture

This is one of the key steps toward SkillOpt functioning as on-the-job training for a specific job role: learning from real work, evaluated on real work.

Happy to help spec this out if there's maintainer interest.