| |
LensVLM-9B by Apple
Apple researchers introduced LensVLM-9B, a vision-language model that processes compressed images of text by selectively expanding relevant sections through learned tools, maintaining accuracy at up to 4.3x compression compared to full-resolution processing. The method outperforms text and visual compression baselines across seven text QA benchmarks and generalizes to document and code understanding tasks, addressing the problem of character degradation in highly compressed images.
Read Full Article →
← More Tech news