Google releases TIPS g/14 low-res v1 vision-language model on Hugging Face
Original titlegoogle/tipsv1-g14-lowres
AISummary
Google has released TIPS g/14 low-res (v1) on Hugging Face, a Text-Image Pre-training with Spatial awareness vision-language model with 1.1B vision parameters and 389M text parameters.
The model produces spatially rich image features aligned with text embeddings at 224 resolution, under the Apache 2.0 license. It supports image encoding, text encoding, and zero-shot classification via the transformers library.
Source: Google · new models on Hugging Face · huggingface.coPublished · added here