#536·skills

☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE

Author: RosellinesCreated Sep 8, 2026Updated Sep 13, 2026

☠️ I BUILT A BENCHMARK WHERE CODING AGENTS COME TO DIE

Most coding benchmarks ask:

“Can the agent solve this task?”

I wanted to ask something nastier:

“Can the agent survive?”

So I built Agent Killer.

Not another leaderboard where agents solve clean LeetCode-style problems.

Agent Killer throws coding agents into hostile trials designed around:

adversarial reasoning cascading failures ☠️ deceptive repositories mutation attacks agent behavior fingerprints network isolation cryptographic result verification hidden evaluation ⚔️ gauntlets & kill chains black-box challenges Human vs AI trials

And the interesting part:

The benchmark itself has to prove that its results can be trusted.

That means:

hidden evaluators signed receipts trusted result chains replay protection resource governance adversarial challenge generation hardened execution boundaries I spent way too much time trying to break my own benchmark.

And every time I found something ugly…

I patched it.

v0.7.15 is now public.

GitHub: https://github.com/Rosellines/Agent-Killers

The question isn't:

“Which AI is smartest?”

The question is:

“Which coding agent can actually survive the trial?”

Try your favorite agent.

Try to break the benchmark.

Try to prove me wrong.

I genuinely want you to. ☠️

#AgentKiller #CodingAgents #AI #Benchmark #LLM #OpenSource #GitHub #AIAgents