Chattering Under Load: Control-Theoretic and Trajectory-Level Behavior of Lion and AdamW Under Physical Perturbation in PPO
DOI:
https://doi.org/10.65419/albahit.v5i3.161Keywords:
Proximal Policy Optimization, Lion optimizer, AdamW, sliding-mode controlAbstract
PPO policies rely in their training almost exclusively on AdamW, despite what the sign-based alternative Lion offers in reducing the optimizer-state memory footprint by half. However, this structural advantage is paired with a decisive structural cost: Lion’s update magnitude is determined directly and absolutely by the instantaneous learning rate, entirely stripped of any self-normalizing compensation mechanism such as the one AdamW possesses. This structural deficiency reduces the update process to something equivalent to a relay/bang-bang controller behavior applied within the parameter space, which makes it inevitably prone to the chattering phenomenon — the persistent high-frequency oscillations that classical theoretical frameworks of Sliding Mode Control and Variable-Structure Control attribute to discontinuous, sign-based control laws. This paper adopts an analytical perspective derived from control theory to investigate this behavior under the weight of physical perturbation across five continuous-control MuJoCo environments — HalfCheetah-v4, Walker2d-v4, Ant-v4, and Hopper-v4, in addition to the extended-horizon Humanoid-v4 — relying on precise joint-level diagnostics for each environment (contact force, actuator saturation, torque chattering, policy entropy), and supported by an integrated physical-perturbation robustness protocol (adding mass by +20% and +50%, and sensor noise, followed by a short fine-tuning re-adaptation phase), at a strict statistical depth of n = 15 seeds. The outputs of this research are organized around three central findings. First, sensor noise — unlike added mass — causes a confirmed statistically significant increase (Welch’s t-test, Holm–Bonferroni corrected) in chattering for most arms across the four main environments (for example, Hopper-v4 lion_cosine: d = −3.50; and Walker2d-v4 lion_warmup_cosine: d = −3.42). This behavior has an exact counterpart in classical control, where measurement noise near the sliding surface is considered the primary trigger for chattering in the sliding-mode control literature. Second, the pattern of post-fine-tune return recovery varies radically depending on the arm within a single environment, and not only depending on the environment: in the Walker2d-v4 environment, lion_cosine recovers only 17–53% of its own clean-condition return depending on the type of perturbation, the lowest recovery profile for any arm in this study, despite being among the most efficient in clean conditions within that environment. Third, the exploratory extension to the Humanoid-v4 environment (after 5000 iterations, roughly 10.24M steps, with all arms and conditions ending in falls) reveals AdamW’s superiority in retaining a higher fraction of its clean-condition return under the three perturbations compared with any Lion variant — a clear reversal and contradiction of the results recorded in clean conditions there. We present the representative state trajectories for the Humanoid-v4 environment and connect their aggregate patterns to a sliding-mode-control-based interpretation of the chattering phenomenon, along with explicit methodological caveats where episode-length truncation causes confounding and distortion of control-quality metrics
References
1. X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, Y. Lu, and Q. V. Le, “Symbolic discovery of optimization algorithms,” in Advances in Neural Information Processing Systems 36 (NeurIPS 2023), 2023.
2. J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar, “signSGD: Compressed optimisation for non-convex problems,” in Proc. 35th Int. Conf. Machine Learning (ICML), PMLR vol. 80, pp. 560–569, 2018.
3. V. I. Utkin, “Variable structure systems with sliding modes,” IEEE Transactions on Automatic Control, vol. 22, no. 2, pp. 212–222, 1977.
4. T. Dettmers, M. Lewis, S. Shleifer, and L. Zettlemoyer, “8-bit optimizers via block-wise quantization,” in International Conference on Learning Representations (ICLR), 2022.
5. J. Zhao, Z. Zhang, B. Chen, Z. Wang, A. Anandkumar, and Y. Tian, “GaLore: Memory-efficient LLM training by gradient low-rank projection,” in Proc. 41st Int. Conf. Machine Learning (ICML), 2024.
6. Y. Dong, H. Li, and Z. Lin, “Convergence rate analysis of LION,” arXiv preprint arXiv:2411.07724, 2024.
7. K. Liang, L. Chen, B. Liu, and Q. Liu, “Cautious optimizers: Improving training with one line of code,” arXiv preprint arXiv:2411.16085, 2024.
8. N. Rudin, D. Hoeller, P. Reist, and M. Hutter, “Learning to walk in minutes using massively parallel deep reinforcement learning,” in Proc. 5th Conf. Robot Learning (CoRL), PMLR vol. 164, pp. 91–100, 2022.
9. A. Kumar, Z. Fu, D. Pathak, and J. Malik, “RMA: Rapid motor adaptation for legged robots,” in Robotics: Science and Systems (RSS) XVII, 2021.
10. P. Wu, W. Xie, J. Cao, H. Lai, and W. Zhang, “LoopSR: Looping sim-and-real for lifelong policy adaptation of legged robots,” arXiv preprint arXiv:2409.17992, 2024.
11. T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter, “Learning robust perceptive locomotion for quadrupedal robots in the wild,” Science Robotics, vol. 7, no. 62, p. eabk2822, 2022.
12. T. He, J. Gao, W. Xiao, Y. Zhang, Z. Wang, J. Wang, Z. Luo, G. He, N. Sobanbabu, C. Pan, Z. Yi, G. Qu, K. Kitani, L. Fan, Y. Zhu, J. Hodgins, C. Liu, and G. Shi, “ASAP: Aligning simulation and real-world physics for learning agile humanoid whole-body skills,” in Robotics: Science and Systems (RSS), 2025.
13. I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations (ICLR), 2019.
14. D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR), 2015.
15. F. H. Clarke, Optimization and Nonsmooth Analysis. Wiley, 1983.
16. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017.
17. E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in 2012 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 5026–5033, 2012.
18. M. Towers, A. Kwiatkowski, J. Terry, J. U. Balis, G. De Cola, T. Deleu, M. Goulão, A. Kallinteris, M. Krimmel, A. KG, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, H. Tan, and O. G. Younis, “Gymnasium: A standard interface for reinforcement learning environments,” arXiv preprint arXiv:2407.17032, 2024.
19. A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-Baselines3: Reliable reinforcement learning implementations,” Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021.
20. L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry, “Implementation matters in deep RL: A case study on PPO and TRPO,” in International Conference on Learning Representations (ICLR), 2020.
21. M. Andrychowicz, A. Raichuk, P. Stańczyk, M. Orsini, S. Girgin, R. Marinier, L. Hussenot, M. Geist, O. Pietquin, M. Michalski, S. Gelly, and O. Bachem, “What matters for on-policy deep actor-critic methods? A large-scale study,” in International Conference on Learning Representations (ICLR), 2021.
22. J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in International Conference on Learning Representations (ICLR), 2016.
23. S. Holm, “A simple sequentially rejective multiple test procedure,” Scandinavian Journal of Statistics, vol. 6, no. 2, pp. 65–70, 1979.
24. J. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates, 1988.
25. C. Colas, O. Sigaud, and P.-Y. Oudeyer, “How many random seeds? Statistical power analysis in deep reinforcement learning experiments,” arXiv preprint arXiv:1806.08295, 2018.



