Hybrid safe reinforcement learning for take-off and merging trajectory optimization in urban air mobility

Published in Transportation Research Part C: Emerging Technologies, 2026

This paper proposes a Hybrid Safe reinforcement learning (RL) framework to address the eVTOL take-off and merging trajectory optimization problem under complex safety constraints. To enhance safety, we employ model-based higher-order control barrier functions (HOCBFs) to encode complex safety constraints. We incorporate the differentiable HOCBF loss into the policy optimization via adaptive Lagrange multipliers to create a gradient-based fence that guides safety adherence during the learning process. Furthermore, to overcome the performance limitations of existing safe RL methods based on single-layer neural networks, we introduce a deep RL architecture with a Hamilton-Jacobi-Bellman-guidance critic training scheme. This model-driven scheme promotes adherence to optimality conditions, leading to a stable approximation of the optimal control value and enhanced trajectory optimality.

Recommended citation: Yingqi Liu, Tianlu Pan, Can Chen, Renxin Zhong (2026). "Hybrid safe reinforcement learning for take-off and merging trajectory optimization in urban air mobility." Transportation Research Part C: Emerging Technologies. Accepted.