arXiv cs.SE· Aleksi Huotala, Miikka Kuutila, Mika M\"antyl\"a·· 2 小时前AI 评分25
LLM 在软件工程系统综述中的筛选性能是否已停滞?
Has LLM Screening Performance Stalled in Software Engineering Systematic Reviews?
AI 导读
研究基于 SESR-Eval 基准及新构建的 SESR-Eval-Mini 数据集,评估了八款新 LLM 在软件工程系统综述筛选任务上的表现,平均 MCC 仅从旧模型的 0.347 微升至 0.365。
来源:arXiv cs.SE · arxiv.org