EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management

Authors: Liang Mi, Weijun Wang, Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu, Haipeng Dai, Guihai Chen, Yunxin Liu, Ting Cao.

Published in: arXiv, 2026

Abstract: Embodied reinforcement learning combines environment simulation, action generation, and model updates, whose heterogeneous CPU and GPU demands make efficient resource utilization difficult. EBRL is an asynchronous embodied RL training system that uses an asynchronous pipelined scheduler to overlap rollout and training while eliminating synchronization stalls, together with a fine-grained resource manager that dynamically allocates pooled CPU cores and GPU streaming multiprocessors. Evaluated with four embodied policies and four simulation benchmarks across heterogeneous GPU testbeds, EBRL achieves 1.30–3.47x higher end-to-end rollout throughput and 2.5x faster training convergence than state-of-the-art embodied RL systems.

BibTeX

@article{ebrl,
  title={EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management},
  author={Liang Mi and Weijun Wang and Bowen Gao and Tianze Yu and Zixu Hao and Han Xiao and Xin Ding and Mingzhe Huang and Xin He and Lu Shi and Hao Wu and Haipeng Dai and Guihai Chen and Yunxin Liu and Ting Cao},
  journal={arXiv preprint arXiv:2609.27547},
  eprint={2609.27547},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.27547},

  year={2026}
}

Download Paper