| |
LensVLM: Compressing long context as images, expanding only relevant pages
LensVLM-9B is a 9-billion-parameter Vision Language Model developed by Apple that compresses long text documents into images and then selectively expands only the relevant pages when answering queries. The model uses learned tools to efficiently handle compressed visual representations of text with compression ratios of 5x, 10x, or 15x, enabling faster processing of large documents. The code and model are publicly available on GitHub under Apple's research and sample code licenses.
Read Full Article →
← More Tech news