Robotics: Science and Systems XXII

Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning

Yuan Liu, Haoran Li, Shuai Tian, Yuxing Qin, Yuhui Chen, Yupeng Zheng, Yongzhen Huang, Dongbin Zhao

Abstract:

Pretrained on large-scale and diverse datasets, VLA models demonstrate strong generalization and adaptability as general-purpose robotic policies. However, Supervised Fine-Tuning (SFT), which serves as the primary mechanism for adapting VLAs to downstream domains, requires substantial amounts of task-specific data and is prone to catastrophic forgetting. To address these limitations, we propose **LifeLong-RFT**, a simple yet effective Reinforcement Fine-Tuning (RFT) strategy for VLA models independent of online environmental feedback and pre-trained reward models. By integrating chunking-level on-policy reinforcement learning with the proposed **M**ulti-**D**imensional **P**rocess **R**eward (MDPR) mechanism, LifeLong-RFT quantifies the heterogeneous contributions of intermediate action chunks across three dimensions to facilitate policy optimization. Specifically, (1) the **Q**uantized **A**ction **C**onsistency **R**eward (QACR) ensures accurate action prediction within the discrete action space; (2) the **C**ontinuous **T**rajectory **A**lignment **R**eward (CTAR) aligns decoded continuous action chunks with reference trajectories to ensure precise control; (3) the **F**ormat **C**ompliance **R**eward (FCR) guarantees the structural validity of outputs. Comprehensive experiments across SimplerEnv, LIBERO, and real-world tasks demonstrate that LifeLong-RFT exhibits strong performance in multi-task learning. Furthermore, for continual learning on the LIBERO benchmark, our method achieves a 22% gain in average success rate over SFT, while effectively adapting to new tasks using only 20% of the training data. Overall, our method provides a promising post-training paradigm for VLAs.

Download:

Bibtex:

  
@INPROCEEDINGS{LiuY2-RSS-26, 
    AUTHOR    = {Yuan Liu AND Haoran Li AND Shuai Tian AND Yuxing Qin AND Yuhui Chen AND Yupeng Zheng AND Yongzhen Huang AND Dongbin Zhao}, 
    TITLE     = {{Towards Long-Lived Robots: Continual Learning VLA Models via Reinforcement Fine-Tuning}}, 
    BOOKTITLE = {Proceedings of Robotics: Science and Systems}, 
    YEAR      = {2026}, 
    ADDRESS   = {Sydney, Australia}, 
    MONTH     = {July}, 
    DOI       = {10.15607/RSS.2026.XXII.086} 
}