التذبذب عالي التردد تحت الحمل: السلوك النظري-التحكمي وعلى مستوى المسارات لمُحسِّنَي Lion وAdamW تحت الاضطراب الفيزيائي في PPO
DOI:
https://doi.org/10.65419/albahit.v5i3.161الكلمات المفتاحية:
chattering، robustness to physical perturbation، MuJoCoالملخص
تعتمد سياسات PPO في تدريبها بشكل شبه حصري على AdamW، على الرغم من أن البديل القائم على الإشارة Lion يوفر ميزة تقليل البصمة الذاكرية لحالة المُحسِّن إلى النصف. ومع ذلك، فإن هذه الميزة البنيوية تقترن بتكلفة بنيوية حاسمة: إذ إن مقدار التحديث في Lion يتحدد بشكل مباشر ومطلق بواسطة معدل التعلم اللحظي، مع تجريده بالكامل من أي آلية تعويض ذاتي للتطبيع مثل تلك التي يمتلكها AdamW. هذا القصور البنيوي يحول عملية التحديث إلى سلوك مكافئ لسلوك متحكم مرحلي/ثنائي الحالة (Relay/Bang-Bang Controller) مطبق داخل فضاء المعلمات، مما يجعله عرضة حتميًا لظاهرة التذبذب عالي التردد (Chattering)، وهي التذبذبات المستمرة عالية التردد التي تربطها الأطر النظرية الكلاسيكية لكل من التحكم بالانزلاق (Sliding Mode Control) والتحكم متغير البنية (Variable-Structure Control) بقوانين التحكم غير المستمرة القائمة على الإشارة.
تتبنى هذه الورقة منظورًا تحليليًا مشتقًا من نظرية التحكم لدراسة هذا السلوك تحت تأثير الاضطرابات الفيزيائية عبر خمس بيئات للتحكم المستمر في MuJoCo — وهي HalfCheetah-v4، Walker2d-v4، Ant-v4، وHopper-v4، بالإضافة إلى البيئة ذات الأفق الزمني الممتد Humanoid-v4 — بالاعتماد على تشخيصات دقيقة على مستوى المفاصل لكل بيئة (قوة التلامس، تشبع المشغلات، تذبذب العزم، وإنتروبيا السياسة)، ومدعومة ببروتوكول متكامل لاختبار المتانة ضد الاضطرابات الفيزيائية (إضافة كتلة بنسبة +20% و+50%، وضوضاء المستشعرات، يليها مرحلة قصيرة من إعادة التكيف بواسطة الضبط الدقيق)، وذلك بعمق إحصائي صارم يبلغ n = 15 بذرة تجريبية.
تُنظم مخرجات هذا البحث حول ثلاث نتائج رئيسية. أولًا، تؤدي ضوضاء المستشعرات — بخلاف إضافة الكتلة — إلى زيادة مؤكدة ذات دلالة إحصائية (اختبار Welch’s t-test مع تصحيح Holm–Bonferroni) في التذبذب عالي التردد لدى معظم المُحسِّنات عبر البيئات الأربع الرئيسية (على سبيل المثال، Hopper-v4 lion_cosine: d = −3.50؛ و Walker2d-v4 lion_warmup_cosine: d = −3.42). يمتلك هذا السلوك نظيرًا مطابقًا في نظرية التحكم التقليدية، حيث تُعتبر ضوضاء القياس بالقرب من سطح الانزلاق العامل الأساسي المحفز لظاهرة التذبذب عالي التردد في أدبيات التحكم بالانزلاق.
ثانيًا، يختلف نمط استعادة العائد بعد الضبط الدقيق بشكل جذري اعتمادًا على المُحسِّن داخل البيئة الواحدة، وليس فقط اعتمادًا على اختلاف البيئة؛ ففي بيئة Walker2d-v4 يستعيد lion_cosine نسبة تتراوح بين 17–53% فقط من عائده الخاص في الحالة النظيفة، وذلك حسب نوع الاضطراب، وهو أقل نمط استعادة لأي مُحسِّن في هذه الدراسة، على الرغم من كونه من بين الأكثر كفاءة في الظروف النظيفة داخل تلك البيئة.
ثالثًا، يكشف الامتداد الاستكشافي إلى بيئة Humanoid-v4 (بعد 5000 دورة تدريبية، أي ما يقارب 10.24 مليون خطوة، مع انتهاء جميع المُحسِّنات وجميع الظروف بحالات سقوط) عن تفوق AdamW في الاحتفاظ بنسبة أعلى من عائده في الظروف النظيفة تحت الاضطرابات الثلاثة مقارنة بأي نسخة من نسخ Lion، وهو انعكاس واضح وتناقض مع النتائج المسجلة في الظروف النظيفة هناك. نقدم المسارات التمثيلية للحالات لبيئة Humanoid-v4، ونربط الأنماط التجميعية الخاصة بها بتفسير قائم على التحكم بالانزلاق لظاهرة التذبذب عالي التردد، إلى جانب عرض القيود المنهجية بشكل صريح في الحالات التي يؤدي فيها تقصير طول الحلقة (Episode-Length Truncation) إلى حدوث عوامل مربكة وتشويه لمقاييس جودة التحكم.
المراجع
1. X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, Y. Lu, and Q. V. Le, “Symbolic discovery of optimization algorithms,” in Advances in Neural Information Processing Systems 36 (NeurIPS 2023), 2023.
2. J. Bernstein, Y.-X. Wang, K. Azizzadenesheli, and A. Anandkumar, “signSGD: Compressed optimisation for non-convex problems,” in Proc. 35th Int. Conf. Machine Learning (ICML), PMLR vol. 80, pp. 560–569, 2018.
3. V. I. Utkin, “Variable structure systems with sliding modes,” IEEE Transactions on Automatic Control, vol. 22, no. 2, pp. 212–222, 1977.
4. T. Dettmers, M. Lewis, S. Shleifer, and L. Zettlemoyer, “8-bit optimizers via block-wise quantization,” in International Conference on Learning Representations (ICLR), 2022.
5. J. Zhao, Z. Zhang, B. Chen, Z. Wang, A. Anandkumar, and Y. Tian, “GaLore: Memory-efficient LLM training by gradient low-rank projection,” in Proc. 41st Int. Conf. Machine Learning (ICML), 2024.
6. Y. Dong, H. Li, and Z. Lin, “Convergence rate analysis of LION,” arXiv preprint arXiv:2411.07724, 2024.
7. K. Liang, L. Chen, B. Liu, and Q. Liu, “Cautious optimizers: Improving training with one line of code,” arXiv preprint arXiv:2411.16085, 2024.
8. N. Rudin, D. Hoeller, P. Reist, and M. Hutter, “Learning to walk in minutes using massively parallel deep reinforcement learning,” in Proc. 5th Conf. Robot Learning (CoRL), PMLR vol. 164, pp. 91–100, 2022.
9. A. Kumar, Z. Fu, D. Pathak, and J. Malik, “RMA: Rapid motor adaptation for legged robots,” in Robotics: Science and Systems (RSS) XVII, 2021.
10. P. Wu, W. Xie, J. Cao, H. Lai, and W. Zhang, “LoopSR: Looping sim-and-real for lifelong policy adaptation of legged robots,” arXiv preprint arXiv:2409.17992, 2024.
11. T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter, “Learning robust perceptive locomotion for quadrupedal robots in the wild,” Science Robotics, vol. 7, no. 62, p. eabk2822, 2022.
12. T. He, J. Gao, W. Xiao, Y. Zhang, Z. Wang, J. Wang, Z. Luo, G. He, N. Sobanbabu, C. Pan, Z. Yi, G. Qu, K. Kitani, L. Fan, Y. Zhu, J. Hodgins, C. Liu, and G. Shi, “ASAP: Aligning simulation and real-world physics for learning agile humanoid whole-body skills,” in Robotics: Science and Systems (RSS), 2025.
13. I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations (ICLR), 2019.
14. D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations (ICLR), 2015.
15. F. H. Clarke, Optimization and Nonsmooth Analysis. Wiley, 1983.
16. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017.
17. E. Todorov, T. Erez, and Y. Tassa, “MuJoCo: A physics engine for model-based control,” in 2012 IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), pp. 5026–5033, 2012.
18. M. Towers, A. Kwiatkowski, J. Terry, J. U. Balis, G. De Cola, T. Deleu, M. Goulão, A. Kallinteris, M. Krimmel, A. KG, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, H. Tan, and O. G. Younis, “Gymnasium: A standard interface for reinforcement learning environments,” arXiv preprint arXiv:2407.17032, 2024.
19. A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann, “Stable-Baselines3: Reliable reinforcement learning implementations,” Journal of Machine Learning Research, vol. 22, no. 268, pp. 1–8, 2021.
20. L. Engstrom, A. Ilyas, S. Santurkar, D. Tsipras, F. Janoos, L. Rudolph, and A. Madry, “Implementation matters in deep RL: A case study on PPO and TRPO,” in International Conference on Learning Representations (ICLR), 2020.
21. M. Andrychowicz, A. Raichuk, P. Stańczyk, M. Orsini, S. Girgin, R. Marinier, L. Hussenot, M. Geist, O. Pietquin, M. Michalski, S. Gelly, and O. Bachem, “What matters for on-policy deep actor-critic methods? A large-scale study,” in International Conference on Learning Representations (ICLR), 2021.
22. J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in International Conference on Learning Representations (ICLR), 2016.
23. S. Holm, “A simple sequentially rejective multiple test procedure,” Scandinavian Journal of Statistics, vol. 6, no. 2, pp. 65–70, 1979.
24. J. Cohen, Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Hillsdale, NJ: Lawrence Erlbaum Associates, 1988.
25. C. Colas, O. Sigaud, and P.-Y. Oudeyer, “How many random seeds? Statistical power analysis in deep reinforcement learning experiments,” arXiv preprint arXiv:1806.08295, 2018.



