شورای راهبری میز هوش مصنوعی استان اصفهان رصد، حکمرانی و توسعه اکوسیستم AI

arXiv هوش مصنوعی 22 ساعت پیش

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

LLM AI 73/100

خلاصه خبر واقعیت از منبع

arXiv:2609.21267v1 Announce Type: new Abstract: Production LLM agents are evaluated repeatedly as they evolve, but full agent benchmarks are costly to rerun. We study efficient recurring evaluation for a production analytics agent serving tens of thousands of monthly active users and report first-hand deployment experience. Using 574 historical runs of the production benchmark, split chronologically into calibration and held-out periods, we compare random sampling, historica…

وضعیت تحلیل

تحلیل هوشمند برای این خبر هنوز با ارائه‌دهنده واقعی تولید نشده است. خلاصه بالا از متادیتای منبع یا ترجمه عنوان/خلاصه است و ادعای تحلیل ساختگی ندارد.

موجودیت‌های مرتبط

عامل هوشمند · technology مدل زبانی بزرگ · technology آمریکا · country

منبع اصلی

مشاهده در منبع اصلی

اخبار مرتبط

بازگشت به مرکز اخبار

ثبت نیاز آموزشی ورود به پلتفرم