[特性] <title>如何让LLM模型更好地提问?感觉现在训练模型都是让模型如何思考,去有效地pass@1,RL应该实现左右互搏、自问自答螺旋飞升?为什么模型无法回答9.11与9.8哪个更大的问题,来暴露缺陷?如何从纯粹以结果为导向的优化转向明确地整合和奖励内省、不确定性量化和自我纠正的机制?</title>
作者: 021gink创建于 2025年6月25日更新于 2026年5月7日
The fundamental transformation is that the training paradigm must shift from “answer-first” to “process and questioning together,” by changing the reward mechanism to explicitly incentivize the model’s self-reflective behavior. The model needs to have four core capabilities: active questioning (for data generation and clarification), robust self-correction (through online RL and process rewards), uncertainty perception (as a self-reflective trigger), and iterative self-improvement (through an epistemological cycle).
内容来源: zai-org/ChatGLM-6B