Official code for our EMNLP 2026 Findings paper PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference.
PACE is a training-free Condense-and-Extract inference framework for VLMs. An Adaptive Pixel Compressor (APC) downsamples redundant pixels before the vision encoder; a Dynamic Dual-Attention Extractor (DDAE) then keeps the salient visual tokens for the LLM.
conda create -n pace python=3.10 -y
conda activate pace
git clone https://github.com/jjL357/PACE.git
cd PACE
pip install -r requirements.txtexport QWEN_MODEL_PATH=Qwen/Qwen2.5-VL-7B-Instruct
bash scripts/reproduce_10_percent.sh
SETTING=fixed TOKEN_BUDGET=0.10 bash scripts/evaluate.shSETTING is fixed or dynamic. Override TASKS, TOKEN_BUDGET, and OUTPUT_PATH as needed. Logs go to outputs/.
PACE/
├── pace_vlm/models/ # APC, DDAE, lmms-eval plugin
├── scripts/ # reproduce / evaluate
├── tests/
├── docs/reproduction.md
└── requirements.txt
@article{liu2026pace,
title={PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference},
author={Liu, Junjie and Ye, Shengyuan and Chen, Xu},
journal={arXiv preprint arXiv:2608.27206},
year={2026}
}We thank the authors of Qwen2.5-VL, lmms-eval, VisionZip, and MMTok for their open-source models, evaluation tools, and visual-token compression baselines.
Released under Apache-2.0. The Qwen2.5-VL modeling file is derived from Hugging Face Transformers and the Qwen team; see NOTICE.