☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE
☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE
Most coding benchmarks ask:
“Can the agent solve this task?”
I wanted to ask something nastier:
“Can the agent survive?”
So I built Agent Killer.
Not another leaderboard where agents solve clean LeetCode-style problems.
Agent Killer throws coding agents into hostile trials designed around:
adversarial reasoning cascading failures ☠️ deceptive repositories mutation attacks agent behavior fingerprints network isolation cryptographic result verification hidden evaluation ⚔️ gauntlets & kill chains black-box challenges Human vs AI trials
And the interesting part:
The benchmark itself has to prove that its results can be trusted.
That means:
hidden evaluators signed receipts trusted result chains replay protection resource governance adversarial challenge generation hardened execution boundaries I spent way too much time trying to break my own benchmark.
And every time I found something ugly…
I patched it.
v0.7.15 is now public.
GitHub: https://github.com/Rosellines/Agent-Killers
The question isn't:
“Which AI is smartest?”
The question is:
“Which coding agent can actually survive the trial?”
Try your favorite agent.
Try to break the benchmark.
Try to prove me wrong.
I genuinely want you to. ☠️
#AgentKiller #CodingAgents #AI #Benchmark #LLM #OpenSource #GitHub #AIAgents
Source: openai/skills