NeoMME: A Single Transformer for Text and Images, Built From Scratch
Hugging Face's NeoMME ditches the VLM playbook—no vision tower, no causal decoder—and trains a pure bidirectional encoder for multimodal retrieval. The result? Competitive retrieval at 260M params.