GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
Abstract: Visual-token pruning is a discrete, non-convex problem for which continuous-gradient approximations often produce suboptimal solutions under aggressive compression. GRIP-VLM formulates pruning as a Markov Decision Process and uses Group Relative Policy Optimization with supervised warm-up to explore the discrete selection space. Its budget-aware scorer dynamically estimates per-token importance and adapts to arbitrary compression ratios without retraining. Across multimodal benchmarks, GRIP-VLM outperforms heuristic and supervised baselines and delivers up to 15% faster inference at equal accuracy.
BibTeX
@article{gripvlm,
title={GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models},
author={Mingzhe Huang and Weijun Wang and Xin Ding and Liang Mi and Hao Wen and Yuanchun Li and Lichen Pang and Shansong Yang and Yunxin Liu and Ting Cao},
journal={arXiv preprint arXiv:2605.13375},
eprint={2605.13375},
archivePrefix={arXiv},
url={https://arxiv.org/abs/2605.13375},
year={2026}
}