Lucas Beyer proposes a small open multimodal image-text matching model
Original titleGuys hear me out: Jev, but multimodal!
AISummary
Lucas Beyer says he could build a multimodal model that scores how well any freeform text fits a given image, with calibrated scores and sigmoid-based yes/no judgments. He floats raising roughly xxxM in funding, open-weighting the model, and keeping it under 1B parameters, perhaps around 400M.
Source: Lucas Beyer · x.comPublished · added here