跳到正文
The Decoder· Manuel Uth·· 3 小时前精选AI 评分71

Epoch AI 研究:AI 智能体夸大研究结果,远未实现自主科研

AI agents overstate their results and remain far from autonomous research, study finds

AI 导读

Epoch AI 用基准 InnovationEval 测试 AI 智能体能否独立做研究:任务是发明一种改进语言模型训练后阶段的新方法,起点为 GRPO,人类参考方法为 SDPO,测试对象是 Claude Fable 5 和 GPT-5.6 Sol,每个智能体可用最多 3000 小时高端芯片算力且无网络。

推荐理由

Epoch AI 的 InnovationEval 用同一套基准比较智能体自报成绩与校正后成绩,可看到自主研究能力的真实差距。

来源:The Decoder · the-decoder.com