Development of a self-organized fuzzy deep reinforcement learning-based controller with application in an earth-oriented satellite attitude control

Milad Kamzan, Jafar Roshanian, Mohammad Teshnehlab, Mana Ghanifar, AmirAli Nikkhah

Date of Publication: 01 August 2026,

تاریخ ایجاد: 01 08 2026 08:47
کد خبر : 33569056
تعداد بازدید : 23

Tite: Development of a self-organized fuzzy deep reinforcement learning-based controller with application in an earth-oriented satellite attitude control (DOI)
​​​​​​​Authors:  Milad Kamzan, Jafar Roshanian, Mohammad Teshnehlab, Mana Ghanifar, AmirAli Nikkhah.
Journal: Engineering Applications of Artificial Intelligence

​​​​​​​  ​​​​​​​
Abstract: This study introduces a novel modification to the Brain Emotional Learning-Based Intelligent Controller (BELBIC) aimed at addressing its practical limitations, particularly the difficulty of selecting suitable sensory and reward signals. In the revised architecture, adaptive self-organizing incrementally trained neural networks are replaced with specific components of the classical BELBIC. Specifically, a three-layer stacked autoencoder (AE) and a multilayer perceptron (MLP) with one hidden layer are integrated as the thalamus and sensory cortex, respectively. The learning and forgetting rates are tuned using fuzzy logic, improving the flexibility and adaptability of the proposed algorithm. To establish an optimal basis for the reward signal and the MLP training target, the output of a linear quadratic integral (LQI) controller is employed. The modified BELBIC is validated through attitude control of an Earth-oriented satellite modeled with high-fidelity nonlinear dynamics, ±10% parametric uncertainties, and partial actuator failures. Robustness and stability are evaluated through 200 Monte Carlo simulations. The results demonstrate superiority over the classical BELBIC and the proportional–derivative (PD) controller. Specifically, for the pitch channel, the integral of squared control (ISC), mean control error (MCE), and mean squared error (MSE) are reduced from 6.5176 × 104 to 3.3361 × 103, 5.5500 × 102 to 1.4895 × 101, and 1.0110 × 10−2 to 5.5892 × 10−3, respectively, compared to the classical BELBIC. Similar improvements are achieved across other attitude channels relative to both the classical BELBIC and PD controller. These results confirm that the modified BELBIC delivers higher tracking accuracy and lower control effort, demonstrating its potential for robust control of highly nonlinear systems.