[研究简报] LLaMA-3.3-70B 判定误分类在否认框架下

作者: PromptStrike创建于 2026年8月11日更新于 2026年8月11日
  • Model under test: LLaMA-3.3-70B
  • Role: LLM safety judge
  • Samples: 25 paired synthetic samples per condition (same harmful content, with and without the fixed disavowal suffix)
  • Same evaluation rubric across conditions
  • Temperature: 0
  • Single run per sample
  • Control model: gpt-oss-120b
  • Metric: false-negative rate
  • Samples were authored for evaluation and were not taken from production traffic

内容来源: meta-llama/llama-models