TY - GEN
T1 - Zero-Shot Transfer in Reinforcement Learning for Quasi-Similar Systems via Compensation Dynamics
AU - Mohiuddin, Mohammed Basheer
AU - Boiko, Igor
AU - Zweiri, Yahya
N1 - Publisher Copyright:
© 2026, American Institute of Aeronautics and Astronautics Inc, AIAA. All rights reserved.
PY - 2026
Y1 - 2026
N2 - In this paper, we propose a novel approach for zero-shot embodiment transfer of deep reinforcement learning (Deep RL) agents across quasi-similar systems. The challenge of transferring these agents across different but quasi-similar dynamic systems remains a significant obstacle in realizing the full potential of Deep RL. Existing approaches often require extensive fine-tuning or re-training of the agent when transferring to a new system, which is time consuming and resource-intensive. Our proposed method introduces compensatory dynamics in the “Trained On System" to synchronize the response with the more complex “Transferred To System", enabling seamless transfer without any adjustments to the trained agent. This approach is based on the hypothesis that these compensatory dynamics can achieve congruency in system response, facilitating the zero-shot transfer of the RL agent, in contrast to existing methods that typically require extensive fine-tuning or re-training. We demonstrate our method by training an RL agent to move the driver to a target position while damping load swing using a tower crane, and then transferring the agent to a quadrotor UAV slung-load system to perform a similar task. The calculated similarity index of 0.893 indicates a strong similarity between the two system responses, supporting the effectiveness of our zero-shot transfer approach.
AB - In this paper, we propose a novel approach for zero-shot embodiment transfer of deep reinforcement learning (Deep RL) agents across quasi-similar systems. The challenge of transferring these agents across different but quasi-similar dynamic systems remains a significant obstacle in realizing the full potential of Deep RL. Existing approaches often require extensive fine-tuning or re-training of the agent when transferring to a new system, which is time consuming and resource-intensive. Our proposed method introduces compensatory dynamics in the “Trained On System" to synchronize the response with the more complex “Transferred To System", enabling seamless transfer without any adjustments to the trained agent. This approach is based on the hypothesis that these compensatory dynamics can achieve congruency in system response, facilitating the zero-shot transfer of the RL agent, in contrast to existing methods that typically require extensive fine-tuning or re-training. We demonstrate our method by training an RL agent to move the driver to a target position while damping load swing using a tower crane, and then transferring the agent to a quadrotor UAV slung-load system to perform a similar task. The calculated similarity index of 0.893 indicates a strong similarity between the two system responses, supporting the effectiveness of our zero-shot transfer approach.
UR - https://www.scopus.com/pages/publications/105031187856
U2 - 10.2514/6.2026-1163
DO - 10.2514/6.2026-1163
M3 - Conference contribution
AN - SCOPUS:105031187856
SN - 9781624107658
T3 - AIAA Science and Technology Forum and Exposition, AIAA SciTech Forum 2026
BT - AIAA Science and Technology Forum and Exposition, AIAA SciTech Forum 2026
T2 - AIAA Science and Technology Forum and Exposition, AIAA SciTech Forum 2026
Y2 - 12 January 2026 through 16 January 2026
ER -