Timeline

AI2 releases Molmo multimodal models

AI2 said its largest Molmo model trained on under a million image-text pairs, roughly three orders of magnitude less data than comparable systems, and released weights and data openly.

  • Open weights & ecosystem
  • Models & capabilities
  • Notable

The Allen Institute for AI released Molmo, a family of open vision-language models ranging from a 1B-parameter mixture-of-experts model to a 72B flagship. Unlike most contemporary multimodal systems, which are trained by distilling outputs from existing proprietary vision-language models, AI2 said Molmo was trained from scratch on its own PixMo dataset — roughly 712,000 images captioned by humans describing what they saw aloud, rather than by typing, to capture richer detail — and that the largest model used fewer than one million image-text pairs, which AI2 described as three orders of magnitude less data than many competing approaches.

AI2 reported that Molmo-72B scored highest among tested models on a set of academic vision-language benchmarks and ranked second on human evaluation, just behind GPT-4o, while even the 1B model approached the performance of GPT-4V. A distinguishing capability was “pointing”: rather than only describing an image in text, Molmo could indicate specific locations within it, which AI2 argued made the models more useful for applications that need to act on what they perceive, such as robotics and interface automation.

The release was notable less for topping benchmarks than for its data efficiency claim and full openness — AI2 published the PixMo dataset alongside the model weights, at a time when most vision-language competitors, open and closed alike, either withheld training data entirely or relied on outputs from other companies’ proprietary models to generate captions.