Consider whether the "adversarial review" framing can be softened
I had a Fable thread where I essentially looped /codex:adversarial-review -> ask Claude to fix -> repeat. The work was not cybersecurity related at all, it only involved simple networking. And yet, on about the 6th cycle (yeah, it's cooked, don't ask), the "Verify and address Codex's findings" prompt tripped Fable's cyber classifier.
Now, I obviously can't know for sure whether this is what tripped it. But I wouldn't be surprised at all if having "adversarial review" repeated 6 times in the context (and likely double that if you include the CoT), would trip the classifier. The term "adversarial" has a cybersecurity/red teaming connotation, even though it isn't being used that way there. So I wonder if maybe it's worth replacing it with a different term to reduce the risk of false flags with frontier models.
Source: openai/codex-plugin-cc