VG-ADP:基于价值引导与扩散对抗的模仿学习方法
集成电路应用
唐坤1,王东升1,韩兴豪2,薄其乐1,王黎铭1,潘舒1
1.江苏科技大学计算机学院;2.江苏自动化所
摘要: 针对多智能体模仿学习中专家数据质量不一及动作分布复杂多峰的问题,该研究提出了一种基于价值引导与扩散对抗的多智能体模仿学习方法(VG-ADP)。该方法采用扩散概率模型作为核心生成器,结合U-Net提取环境时空特征,旨在精准刻画专家决策的多峰分布并提升动作序列的平滑度。为应对数据噪声,框架引入对抗增强模块识别并过滤次优数据;同时,在扩散过程的逆向去噪阶段注入价值梯度算子,利用协同价值先验引导智能体向高回报区域纠偏。
中图分类号:TP391.9文献标识码:ADOI:10.19339/j.issn.1674-2583.2026.04.011
中文引用格式:唐坤,王东升,韩兴豪,等. VG-ADP:基于价值引导与扩散对抗的模仿学习方法[J].集成电路应用,2026,43(4):60-64.
英文引用格式:Tang Kun, Wang Dongsheng, Han Xinghao,et al. VG-ADP: Value-guided adversarial diffusion imitation learning[J].Application of IC,2026,43(4):60-64.
中文引用格式:唐坤,王东升,韩兴豪,等. VG-ADP:基于价值引导与扩散对抗的模仿学习方法[J].集成电路应用,2026,43(4):60-64.
英文引用格式:Tang Kun, Wang Dongsheng, Han Xinghao,et al. VG-ADP: Value-guided adversarial diffusion imitation learning[J].Application of IC,2026,43(4):60-64.
VG-ADP: Value-guided adversarial diffusion imitation learning
Tang Kun1, Wang Dongsheng1, Han Xinghao2, Bo Qile1, Wang Liming1, Pan Shu1
1. School of Computer, Jiangsu University of Science and Technology; 2. Jiangsu Automation Research Institute
Abstract: To address suboptimal expert data and complex multimodal action distributions in multi-agent imitation learning, this research proposes a value-guided adversarial diffusion multi-agent imitation learning framework (VG-ADP). The framework employs a Diffusion Probabilistic Model as the core generator, integrated with a U-Net structure to capture spatio-temporal environmental features, accurately modeling the multimodal nature of expert strategies while enhancing action sequence smoothness. To mitigate data noise, an adversarial enhancement module is introduced to identify and filter suboptimal trajectories. Furthermore, a value gradient operator is injected into the reverse denoising phase, utilizing collaborative value priors to dynamically guide agents toward high-reward regions for trajectory correction.
Key words : diffusion model;adversarial imitation learning;value guidance
引言
面对具有严重次优性、复杂噪声以及策略不一致的多源轨迹数据,传统的多智能体行为克隆方法难以有效提取高回报的协同动作区域,极易收敛至无任务解决能力的次优解或导致策略彻底失效。相比之下,基于扩散生成与对抗机制的模仿学习方法,凭借其对复杂联合数据分布的高阶匹配与多峰拟合能力,成为处理该类问题的理想范式。然而,现有生成类方法在多智能体高动态场景中,往往受限于探索的盲目性与训练的不稳定性,难以直接生成高精度的协同动作。
本文详细内容请下载:
http://www.chinaaet.com/resource/share/2000007265
投稿邮箱:ICAPP@belling.com.cn
作者信息:
唐坤1,王东升1,韩兴豪2,薄其乐1,王黎铭1,潘舒1
(1.江苏科技大学计算机学院,江苏镇江212000;
2.江苏自动化所,江苏连云港222006)

此内容为AET网站原创,未经授权禁止转载。
