TIPS So400m/14 v1 Vision-Language Model Released on Hugging Face
Original titlegoogle/tipsv1-so400m14
AISummary
Google released google/tipsv1-so400m14, the original v1 So400m/14 checkpoint of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 413M vision parameters and 448M text parameters at 448 resolution, and is licensed under Apache 2.0.
Source: Google · new models on Hugging Face · huggingface.co