شورای راهبری میز هوش مصنوعی استان اصفهان رصد، حکمرانی و توسعه اکوسیستم AI

arXiv یادگیری ماشین دیروز

SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?

LLM AI 63/100

خلاصه خبر واقعیت از منبع

arXiv:2609.21190v1 Announce Type: new Abstract: Ensuring the correctness of LLM-generated code is a core challenge for modern software engineering. Benchmarks for agentic code generation check correctness with held-out test suites, which are inherently incomplete and increasingly susceptible to memorization. Formal verification avoids both problems, but existing work covers only standalone tasks whose specifications are given as input, not real issues, which touch large repo…

وضعیت تحلیل

تحلیل هوشمند برای این خبر هنوز با ارائه‌دهنده واقعی تولید نشده است. خلاصه بالا از متادیتای منبع یا ترجمه عنوان/خلاصه است و ادعای تحلیل ساختگی ندارد.

موجودیت‌های مرتبط

عامل هوشمند · technology مدل زبانی بزرگ · technology هوش عامل‌محور · technology تراشه AI · technology

منبع اصلی

مشاهده در منبع اصلی

اخبار مرتبط

بازگشت به مرکز اخبار

ثبت نیاز آموزشی ورود به پلتفرم