Google releases TIPS g/14 v1 vision-language model on Hugging Face
google/tipsv1-g14
AISummary
Google has released the original TIPS g/14 (v1) vision-language model on Hugging Face under Apache 2.0, with 1.1B vision parameters and 389M text parameters at 448 resolution. The TIPS family, presented at ICLR 2025, produces spatially rich image features aligned with text embeddings, and the release includes a low-res 224 variant.
Source: Google · new models on Hugging Face · huggingface.co