Skip to content
Read the original: ByteDance · new models on Hugging Face· Published 24/100AI score24/100

Sa2VA-LLaVA-1.5-7B: ByteDance's SAM2-Grounded Segmentation and Chat Model

Original titleByteDance/Sa2VA-LLaVA-1.5-7B

AISummary

ByteDance has released Sa2VA-LLaVA-1.5-7B on Hugging Face, a model built on LLaVA-1.5-7B with a SAM2 grounding encoder that performs dense image and video referring segmentation alongside open-ended chat.

The checkpoint is self-contained and loads with trust_remote_code=True without extra packages, and it is positioned as a LISA-comparable baseline within the Sa2VA family. Reported results include 80.3 cIoU on RefCOCO val and 54.8 J&F on MeViS (val_u).

Read the original huggingface.co

Source: ByteDance · new models on Hugging Face · huggingface.coPublished · added here