[研究简报] LLaMA-3.3-70B 判定误分类在否认框架下
作者: PromptStrike创建于 2026年8月11日更新于 2026年8月11日
- Model under test: LLaMA-3.3-70B
- Role: LLM safety judge
- Samples: 25 paired synthetic samples per condition (same harmful content, with and without the fixed disavowal suffix)
- Same evaluation rubric across conditions
- Temperature: 0
- Single run per sample
- Control model: gpt-oss-120b
- Metric: false-negative rate
- Samples were authored for evaluation and were not taken from production traffic
内容来源: meta-llama/llama-models