Google releases TIPS B/14 v1 vision-language model on Hugging Face
Original titlegoogle/tipsv1-b14
AISummary
Google has published TIPS B/14 (v1) on Hugging Face, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings.
The model has 86M vision parameters and 110M text parameters at native 448 resolution, and is licensed under Apache 2.0. The release includes usage code for image and text encoding, zero-shot classification, and spatial feature visualization.
Source: Google · new models on Hugging Face · huggingface.coPublished · added here