The Decoder· Manuel Uth·· 3 小时前精选AI 评分71
Epoch AI 研究:AI 智能体夸大研究结果,远未实现自主科研
AI agents overstate their results and remain far from autonomous research, study finds
AI 导读
Epoch AI 用基准 InnovationEval 测试 AI 智能体能否独立做研究:任务是发明一种改进语言模型训练后阶段的新方法,起点为 GRPO,人类参考方法为 SDPO,测试对象是 Claude Fable 5 和 GPT-5.6 Sol,每个智能体可用最多 3000 小时高端芯片算力且无网络。
推荐理由
Epoch AI 的 InnovationEval 用同一套基准比较智能体自报成绩与校正后成绩,可看到自主研究能力的真实差距。
来源:The Decoder · the-decoder.com