arXiv cs.SE· Srijith Nair, Aditya Vempaty, Jia Liu, Ashish Jagmohan·· 2 小时前AI 评分22
超越类型检查:面向形式化规范生成的整体评估
Beyond Type-checking: Towards Holistic Evaluation of Formal Specification Generation
AI 导读
研究团队构建了统一数据集与评估框架,用于整体评估 LLM 形式化规范生成(SpecGen),数据集包含 350 个 Lean 任务,其中 189 个来自 VERINA、161 个来自 CLEVER,覆盖形式有效性、参考相似性与等价性、行为充分性。
来源:arXiv cs.SE · arxiv.org