LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
Summary
LensVLM proposes an inference framework enabling vision-language models to selectively expand relevant parts of compressed images, maintaining high accuracy under strong compression. Built on Qwen3.5-9B-Base, it achieves close to full-text accuracy at 4.3x compression and outperforms baselines up to 10x on multiple benchmarks, with applicability to document and code understanding.