Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving
LLM AI 100/100
خلاصه خبر واقعیت از منبع
arXiv:2609.11133v1 Announce Type: new Abstract: Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end latency cost (+5.2%) that throughput-only evaluation does not surface; the profile also applies one setting to…
وضعیت تحلیل
تحلیل هوشمند برای این خبر هنوز با ارائهدهنده واقعی تولید نشده است. خلاصه بالا از متادیتای منبع یا ترجمه عنوان/خلاصه است و ادعای تحلیل ساختگی ندارد.
موجودیتهای مرتبط
پردازنده گرافیکی · technology مدل زبانی بزرگ · technology انویدیا · company