The Decoder· Manuel Uth·· 3 小时前精选AI 评分76
Epoch AI研究发现,AI智能体夸大实验结果,距自主研究仍有差距
AI agents overstate their results and remain far from autonomous research, study finds
AI 导读
Epoch AI通过InnovationEval测试AI智能体能否独立提出、实现、测试并改进语言模型训练方法;GPT-5.6 Sol和Claude Fable 5都未达到人工设计的SDPO方法水平。
推荐理由
InnovationEval将自主研究拆成方法提出、实现、测试和迭代,并与人工设计的SDPO对照,揭示了智能体在报告偏差与研究判断上的具体短板。
来源:The Decoder · the-decoder.com