文章摘要
胡存琛* ** ***,孙凝晖* **,王卅* **.大语言模型推理优化研究综述[J].高技术通讯(中文),2026,36(7):663~676
大语言模型推理优化研究综述
A survey on optimization research for large language model inference
  
DOI:10. 3772 / j. issn. 1002 - 0470. 2026. 07. 001
中文关键词: 大模型服务; 请求调度; 缓存加速; 显存管理; 算子优化
英文关键词: large language model serving, request scheduling, caching accelerate, memory management, operator optimization
基金项目:
作者单位
胡存琛* ** *** (*中国科学院计算技术研究所处理器芯片全国重点实验室北京 100190) (**中国科学院大学北京100049) (***中国电信云计算研究院北京 100053) 
孙凝晖* **  
王卅* **  
摘要点击次数: 89
全文下载次数: 65
中文摘要:
      大语言模型(large language model,LLM)快速增长的参数规模以及自回归特性使得推理的高服务成本与低资源使用效率之间出现巨大的矛盾。推理系统的服务性能依赖于系统架构、算法、模型等各方面。本文总结了推理系统面临的问题和挑战,并且从请求处理流程和模型服务架构出发,重点介绍了请求调度、资源管理、令牌生成、分布式推理以及推理缓存方面的研究进展,分析了提升推理系统效率所采用的关键技术。最后,本文探讨了异构硬件资源与文生图场景下降低推理服务成本的挑战,认为软硬协同、硬件能力感知的设计有助于进一步提升资源效率以及降低推理成本。
英文摘要:
      The rapid growth in parameter size of large language model (LLM) and their autoregressive nature has led to a significant contradiction between high service costs for inference and low resource utilization efficiency. The performance of inference systems depends on various factors, including system architecture, algorithms, and models. The paper summarizes the challenges faced by inference systems from the request processing flow and model service architecture. It focuses on research advancements in request scheduling, resource management, token generation, distributed inference, and inference caching. The paper also analyzes key technologies for improving inference system efficiency. Finally, this work discusses the challenges of reducing inference service costs in heterogeneous hardware resources and text-to-image scenarios, suggesting that software-hardware co-design and hardware-aware strategies can further enhance resource efficiency and reduce inference costs.
查看全文   查看/发表评论  下载PDF阅读器
关闭