Meta releases V-JEPA
Trained without labels, the model learned by predicting masked video regions in representation space, distinguishing fine-grained actions over roughly ten-second clips.
Open weights & ecosystem · Models & capabilities · Ideas & essays