LensVLM: Compressing long context as images, expanding only relevant pages
Summary
LensVLM-9B is a 9B vision-language model that compresses long visual contexts into images and selectively expands only the relevant pages. The Hugging Face page provides the arXiv paper, code, and detailed usage instructions (pip install, Transformers, vLLM, Docker), plus a model tree and deployment options. It emphasizes open-source access and practical demos for image-text tasks.