The Impact of Hybrid Optimizers on Policy Approximation in Deep Reinforcement Learning
Fragment książki (Rozdział monografii pokonferencyjnej)
MNiSW
20
Poziom I
| Status: | |
| Autorzy: | Tomiło Paweł, Michałowska Joanna, Kochan Roman, Kochan Orest, Kochan Volodymyr |
| Dyscypliny: | |
| Aby zobaczyć szczegóły należy się zalogować. | |
| Wersja dokumentu: | Drukowana | Elektroniczna |
| Język: | angielski |
| Strony: | 29 - 34 |
| Web of Science® Times Cited: | 0 |
| Bazy: | Web of Science |
| Efekt badań statutowych | NIE |
| Materiał konferencyjny: | TAK |
| Nazwa konferencji: | 30th International Conference on Methods and Models in Automation and Robotics |
| Skrócona nazwa konferencji: | 30th MMAR 2026 |
| URL serii konferencji: | LINK |
| Termin konferencji: | 18 sierpnia 2026 do 21 sierpnia 2026 |
| Miasto konferencji: | Międzyzdroje |
| Państwo konferencji: | POLSKA |
| Publikacja OA: | NIE |
| Abstrakty: | angielski |
| This paper investigates the impact of hybrid optimization algorithms that combine classical first-order methods with the Muon gradient moment orthogonalization mechanism on the effectiveness of agent learning in deep reinforcement learning (DRL) tasks. The study was conducted in the Isaac Lab simulation environment using three diverse benchmark scenarios. All experiments employed the same neural network architecture, and the quality of the learned policies was evaluated in terms of the average cumulative reward computed over 100 evaluation episodes. The results were subjected to statistical analysis using permutation tests and Cliff's δ. The experimental results indicate that the impact of Muon-based hybrid optimizers depends strongly on both the environment and the baseline optimization algorithm. In several configurations, the incorporation of Muon resulted in statistically significant improvements in cumulative reward compared to the corresponding baseline optimizers. Overall, the results suggest that hybrid optimization strategies incorporating Muon may improve learning effectiveness in selected DRL settings, while their performance remains highly task- and optimizer-dependent. |