1.河南工业大学人工智能与大数据学院,河南 郑州 450001
2.河南工业大学电气工程学院,河南 郑州 450001
3.电子科技大学基础与前沿研究院,四川 成都 610050
4.河南科技学院人工智能学院,河南 新乡 453003
[ "武子怡(2004- ),女,河南工业大学人工智能与大数据学院在读,主要研究方向为强化学习。" ]
[ "高晓楠(2003- ),女,河南工业大学人工智能与大数据学院在读,主要研究方向为强化学习。" ]
[ "朱献超(1990- ),男,博士,河南工业大学人工智能与大数据学院讲师,主要研究方向为强化学习和机器学习。" ]
[ "赵亮(1978- ),男,博士,河南工业大学电气工程学院副教授,主要研究方向为深度学习、平行学习。" ]
[ "祝峰(1962- ),男,博士,电子科技大学基础与前沿研究院教授,主要研究方向为粗糙集和粒计算。" ]
[ "蔡磊(1979- ),男,博士,河南科技学院人工智能学院教授,主要研究方向为智能无人系统协同控制和机器人控制。" ]
收稿:2025-03-21,
修回:2025-06-09,
录用:2025-07-17,
纸质出版:2025-09-15
移动端阅览
武子怡,高晓楠,朱献超等.基于KPPO算法的四旋翼无人机飞行控制[J].智能科学与技术学报,2025,07(03):381-395.
WU Ziyi,GAO Xiaonan,ZHU Xianchao,et al.Quadrotor UAV flight control based on the KPPO algorithm[J].Chinese Journal of Intelligent Science and Technology,2025,07(03):381-395.
武子怡,高晓楠,朱献超等.基于KPPO算法的四旋翼无人机飞行控制[J].智能科学与技术学报,2025,07(03):381-395. DOI: 10.11959/j.issn.2096-6652.202531.
WU Ziyi,GAO Xiaonan,ZHU Xianchao,et al.Quadrotor UAV flight control based on the KPPO algorithm[J].Chinese Journal of Intelligent Science and Technology,2025,07(03):381-395. DOI: 10.11959/j.issn.2096-6652.202531.
针对近端策略优化(PPO)算法在四旋翼无人机飞行控制中存在的收敛速度慢和难以适应复杂环境问题,在PPO算法的基础上,融合了阈值触发的KL散度惩罚项与L2正则项构成的复合正则机制,提出了一种改进算法KPPO算法,并将其应用于四旋翼无人机飞行控制任务。通过四旋翼无人机飞行控制物理仿真验证,结果显示,在KPPO算法策略的驱动下,四旋翼无人机能够迅速达到策略收敛状态,并在复杂环境中做出正确的决策,从而显著提升了传统算法的性能及任务执行效率。研究结果证实了KPPO算法在优化四旋翼无人机飞行控制中的有效性。
To address the issues of slow convergence and the difficulty of adapting to complex environments in proximal policy optimization (PPO) algorithms for quadrotor UAV flight control
an improved algorithm was proposed
KPPO
which integrated a composite regularization mechanism composed of a threshold-triggered KL divergence penalty and an L2 regularization term based on the PPO framework. The proposed KPPO algorithm was applied to quadrotor UAV flight control tasks. Physical simulation validation of quadrotor UAV flight control demonstrates that under the guidance of the KPPO strategy
the quadrotor UAV rapidly achieves policy convergence and makes correct decisions in complex environments
thereby significantly enhancing the performance of conventional algorithms. Notably
the KPPO algorithm improves task execution efficiency
a key factor in quadrotor UAV operations. This reassures the audience of the effectiveness of the KPPO algorithm in improving quadrotor UAV flight control.
NGUYEN V N, JENSSEN R, ROVERSO D. Intelligent monitoring and inspection of power line components powered by UAVs and deep learning[J]. IEEE Power and Energy Technology Systems Journal, 2019, 6(1): 11-21.
CHOUTRI K, LAGHA M, DALA L. A fully autonomous search and rescue system using quadrotor UAV[J]. International Journal of Computing and Digital Systems, 2021, 10(1): 403-414.
任博, 潘景余, 苏畅, 等. 不确定环境下的侦察无人机自主航路规划仿真[J]. 电光与控制, 2008, 15(1): 31-34, 46.
REN B, PAN J Y, SU C, et al. Autonomous path planning simulation of UAVs in uncertain environment[J]. Electronics Optics & Control, 2008, 15(1): 31-34, 46.
ZHENG Y J, DU Y C, LING H F, et al. Evolutionary collaborative human-UAV search for escaped criminals[J]. IEEE Transactions on Evolutionary Computation, 2020, 24(2): 217-231.
TIAN Y L, LIU K, OK K, et al. Search and rescue under the forest canopy using multiple UAVs[J]. The International Journal of Robotics Research, 2020, 39(10/11): 1201-1221.
XING L J, FAN X Y, DONG Y X, et al. Multi-UAV cooperative system for search and rescue based on YOLOv5[J]. International Journal of Disaster Risk Reduction, 2022, 76: 102972.
晏磊, 廖小罕, 周成虎, 等. 中国无人机遥感技术突破与产业发展综述[J]. 地球信息科学学报, 2019, 21(4): 476-495.
YAN L, LIAO X H, ZHOU C H, et al. The impact of UAV remote sensing technology on the industrial development of China: a review[J]. Journal of Geo-Information Science, 2019, 21(4): 476-495.
赵静, 闫春雨, 杨东建, 等. 基于无人机多光谱遥感的台风灾后玉米倒伏信息提取[J]. 农业工程学报, 2021, 37(24): 56-64.
ZHAO J, YAN C Y, YANG D J, et al. Extraction of maize lodging information after typhoon based on UAV multispectral remote sensing[J]. Transactions of the Chinese Society of Agricultural Engineering, 2021, 37(24): 56-64.
ANG K H, CHONG G, LI Y. PID control system analysis, design, and technology[J]. IEEE Transactions on Control Systems Technology, 2005, 13(4): 559-576.
NG T C T, LEUNG F H F, TAM P K S. A simple gain scheduled PID controller with stability consideration based on a grid-point concept[C]//Proceedings of the ISIE '97 Proceeding of the IEEE International Symposium on Industrial Electronics. Piscataway: IEEE Press, 2002: 1090-1094.
PAPADOPOULOS K G, TSELEPIS N D, MARGARIS N I. On the automatic tuning of PID type controllers via the Magnitude Optimum criterion[C]//Proceedings of the 2012 IEEE International Conference on Industrial Technology. Piscataway: IEEE Press, 2012: 869-874.
REYES-VALERIA E, ENRIQUEZ-CALDERA R, CAMACHO-LARA S, et al. LQR control for a quadrotor using unit quaternions: Modeling and simulation[C]//Proceedings of the CONIELECOMP 2013, 23rd International Conference on Electronics, Communications and Computing. Piscataway: IEEE Press, 2013: 172-178.
FALANGA D, FOEHN P, LU P, et al. PAMPC: perception-aware model predictive control for quadrotors[C]//Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems. Piscataway: IEEE Press, 2018: 1-8.
RAMEZANI DOORAKI A, LEE D J. An innovative bio-inspired flight controller for quad-rotor drones: Quadrotor drone learning to fly using reinforcement learning[J]. Robotics and Autonomous Systems, 2021, 135: 103671.
GHOURI U H, ZAFAR M U, BARI S, et al. Attitude control of quad-copter using deterministic policy gradient algorithms (DPGA)[C]//Proceedings of the 2019 2nd International Conference on Communication, Computing and Digital systems. Piscataway: IEEE Press, 2019: 149-153.
陈佳盼, 郑敏华. 基于深度强化学习的机器人操作行为研究综述[J]. 机器人, 2022, 44(2): 236-256.
CHEN J P, ZHENG M H. A survey of robot manipulation behavior research based on deep reinforcement learning[J]. Robot, 2022, 44(2): 236-256.
LU J W, HAN L Y, WEI Q L, et al. Event-triggered deep reinforcement learning using parallel control: a case study in autonomous driving[J]. IEEE Transactions on Intelligent Vehicles, 2023, 8(4): 2821-2831.
HAARNOJA T, BEN M R, LEVER G, et al. Learning agile soccer skills for a bipedal robot with deep reinforcement learning[J]. Science Robotics, 2024, 9(89): eadi8022.
ZHOU Y T, YANG J C, GUO Z W, et al. An indoor blind area-oriented autonomous robotic path planning approach using deep reinforcement learning[J]. Expert Systems with Applications, 2024, 254: 124277.
ORR J, DUTTA A. Multi-agent deep reinforcement learning for multi-robot applications: a survey[J]. Sensors, 2023, 23(7): 3625.
KOCH W, MANCUSO R, WEST R, et al. Reinforcement learning for UAV attitude control[J]. ACM Transactions on Cyber-Physical Systems, 2019, 3(2): 1-21.
HWANGBO J, SA I, SIEGWART R, et al. Control of a quadrotor with reinforcement learning[J]. IEEE Robotics and Automation Letters, 2017, 2(4): 2096-2103.
梁晨, 刘小雄, 张兴旺, 等. 基于强化学习的四旋翼无人机控制律设计[J]. 计算机测量与控制, 2021, 29(2): 71-75, 86.
LIANG C, LIU X X, ZHANG X W, et al. Design of control law for quadrotor UAV based on reinforcement learning[J]. Computer Measurement & Control, 2021, 29(2): 71-75, 86.
HAARNOJA T, ZHOU A, ABBEEL P, et al. Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor[J]. arXiv preprint, 2018, arXiv: 1801.01290.
何准, 董文瀚, 蔡鸣, 等. 基于DDPG的多旋翼无人机自主引导与跟踪方法[J]. 飞行力学, 2021, 39(2): 63-69, 76.
HE Z, DONG W H, CAI M, et al. Multi-rotor UAV autonomous guidance and tracking method based on DDPG[J]. Flight Dynamics, 2021, 39(2): 63-69, 76.
胡多修, 董文瀚, 解武杰. 无人机自主引导跟踪与避障的近端策略优化[J]. 北京航空航天大学学报, 2023, 49(1): 195-205.
HU D X, DONG W H, XIE W J. Proximal policy optimization for UAV autonomous guidance, tracking and obstacle avoidance[J]. Journal of Beijing University of Aeronautics and Astronautics, 2023, 49(1): 195-205.
梁吉, 王立松, 黄昱洲, 等. 基于深度强化学习的四旋翼无人机自主控制方法[J]. 计算机科学, 2023, 50(S2): 13-19.
LIANG J, WANG L S, HUANG Y Z, et al. Autonomous control method of quadrotor UAV based on deep reinforcement learning[J]. Computer Science, 2023, 50(S2): 13-19.
王伟, 吴昊, 刘鸿勋, 等. 基于深度强化学习的无人机姿态控制器设计[J]. 科学技术与工程, 2023, 23(34): 14888-14895.
WANG W, WU H, LIU H X, et al. Design of UAV attitude controller based on deep reinforcement learning[J]. Science Technology & Engineering, 2023, 23(34): 14888-14895.
黄号, 马文卉, 李家诚, 等. 未知环境下无人机编队智能避障控制方法[J]. 清华大学学报(自然科学版), 2024, 64(2): 358-369.
HUANG H, MA W H, LI J C, et al. Intelligent obstacle avoidance control method for unmanned aerial vehicle formations in unknown environments[J]. Journal of Tsinghua University (Science and Technology), 2024, 64(2): 358-369.
孔飞, 赵振根, 程磊, 等. 输入受限及干扰下固定翼无人机强化学习控制[J]. 电光与控制, 2024, 31(2): 21-28.
KONG F, ZHAO Z G, CHENG L, et al. Reinforcement learning control of fixed-wing UAV under input limitation and disturbance[J]. Electronics Optics & Control, 2024, 31(2): 21-28.
江泰民, 谭泰, 李辉, 等. 基于分层深度强化学习的六自由度固定翼无人机路径跟踪方法[J]. 计算机工程, 已录用, 2024: 0070197.
JIANG T M, TAN T, LI H, et al. Path following of 6-DOF fixed-wing UAV based on hierarchical deep reinforcement learning[J]. Computer Engineering, accepted, 2024: 0070197.
0
浏览量
125
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621
