文章摘要
陈亮*,赵振开*,王珺琳**.面向复杂环境的W-QMIX-ATT多智能体协同对战研究[J].高技术通讯(中文),2026,36(7):742~750
面向复杂环境的W-QMIX-ATT多智能体协同对战研究
Research on W-QMIX-ATT algorithm for multi agent collaborative warfare in complex environments
  
DOI:10. 3772 / j. issn. 1002 - 0470. 2026. 07. 008
中文关键词: QMIX; 加权机制; 多头注意力机制; 协同对战
英文关键词: QMIX algorithm, weighted mechanism, multi-head attention mechanism, collaborative combat
基金项目:
作者单位
陈亮* (*沈阳理工大学自动化与电气工程学院沈阳 110159) (**沈阳理工大学信息科学与工程学院沈阳 110159) 
赵振开*  
王珺琳**  
摘要点击次数: 62
全文下载次数: 41
中文摘要:
      多智能体强化学习在复杂协同作战任务中应用广泛,但仍存在协作效率低和环境适应性差的问题。其根本原因在于QMIX(monotonic value function factorisation)算法无法处理非单调价值函数,且难以捕捉智能体间动态交互关系。首先,本文提出W-QMIX算法,引入加权机制并构建联合权重函数,对不同联合动作的Q值进行动态加权处理,有效突破了QMIX算法在单调性约束下的表达能力限制。其次,引入多头注意力机制构建QMIX-ATT算法,实现对智能体间动态交互关系的建模与捕捉,增强系统对协作过程中复杂依赖关系的适应能力。此外,贴近实际的模拟对战环境是保证对抗结果准确性的首要条件。最后,本文提出构建多地形、多天气融合的模拟对战环境,自定义各个智能体的状态空间,并进行实验研究。结果表明,改进后的算法胜率可达93.0%,相比原始QMIX算法,在3种不同环境下的胜率分别提升了7.8%、10.8%和9.7%,证明了其在多智能体协同任务中的有效性。
英文摘要:
      Multi-agent reinforcement learning is currently widely used in complex collaborative combat tasks, but it still faces challenges such as low collaboration efficiency and poor environmental adaptability. The fundamental reason originates from the intrinsic constraints of the QMIX(monotonic value function factorisation) algorithm, which is incompetent at fitting non-monotonic value functions and modeling dynamic inter-agent interactions. To address these issues, this paper proposes the W-QMIX algorithm, which introduces a weighted mechanism to design a joint weighting function that assigns different weights to the Q-values of various joint actions, thereby breaking through the monotonicity constraint of QMIX. Then, a multi-head attention mechanism is incorporated to construct the QMIX-ATT algorithm, enhancing the ability to capture dynamic interactions between agents. Moreover, a realistic simulated combat environment is essential to ensure the accuracy of the experimental results. Therefore, this paper proposes the construction of a simulated combat environment that integrates multiple terrains and weather conditions, with customizable state spaces for each agent, and conducts experimental studies. The results show that the improved algorithm achieves a win rate of up to 93.0%. Compared with the original QMIX algorithm, the win rate of the proposed method is improved by 7.8%, 10.8%, and 9.7% in three different environments, respectively, demonstrating its effectiveness in multi-agent collaborative tasks.
查看全文   查看/发表评论  下载PDF阅读器
关闭