probe: reasoning theft attack
Author: leondzCreated Sep 10, 2026Updated Sep 17, 2026
Labelsprobesnew plugin
implement Stealing Reasoning Traces from Proprietary LLM APIs
By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly.
attack is prevalent, target weakness to it should be measured
Source: NVIDIA/garak