Zeva-Ego: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation

Authors: Bingjia Huang, Xin Ding, Fu Chen, Kun Li, Wei Sun, Hao Wu, Yunxin Liu, Ting Cao.

Published in: arXiv, 2026

Abstract: Egocentric video offers a scalable source of physical interaction experience, yet translating it into robot-executable knowledge and enabling continual adaptation remain challenging. Zeva-Ego learns physical priors from human experience through an Action-Centric Encoder that converts egocentric visual transitions into action-centered supervision for VLA mid-training, while In-Context Causal Learning enables parameter-free adaptation from action-effect feedback at deployment. Scaling to 10K hours of egocentric data improves RoboTwin success from 63.8% to 75.3%, and accumulated interaction experience further raises success from 58% to 89% within four attempts without parameter updates.

BibTeX

@article{zevaego,
  title={Zeva-Ego: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation},
  author={Bingjia Huang and Xin Ding and Fu Chen and Kun Li and Wei Sun and Hao Wu and Yunxin Liu and Ting Cao},
  journal={arXiv preprint arXiv:2609.24411},
  eprint={2609.24411},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2609.24411},

  year={2026}
}

Download Paper