شورای راهبری میز هوش مصنوعی استان اصفهان رصد، حکمرانی و توسعه اکوسیستم AI

arXiv هوش مصنوعی ۱۴۰۵/۰۶/۱۲

HeadWiseKV: Budgeted Per-Head Cache Residency for Hybrid Long-Context Language Models

LLM AI 60/100

خلاصه خبر واقعیت از منبع

arXiv:2609.02029v1 Announce Type: new Abstract: Long-context inference retains a growing key--value (KV) cache during decoding, which consumes substantial GPU memory and can reduce generation throughput. This bottleneck remains in hybrid language models because their residual global-attention layers can dominate context-dependent cache demand. We study how to allocate this state under an aggregate KV-residency budget.

وضعیت تحلیل

تحلیل هوشمند برای این خبر هنوز با ارائه‌دهنده واقعی تولید نشده است. خلاصه بالا از متادیتای منبع یا ترجمه عنوان/خلاصه است و ادعای تحلیل ساختگی ندارد.

موجودیت‌های مرتبط

پردازنده گرافیکی · technology مدل زبانی بزرگ · technology

منبع اصلی

مشاهده در منبع اصلی

اخبار مرتبط

بازگشت به مرکز اخبار

ثبت نیاز آموزشی ورود به پلتفرم