ByteDance releases Sa2VA-Qwen3-VL-4B-SAM3 for image and video referring segmentation
Original titleByteDance/Sa2VA-Qwen3-VL-4B-SAM3
AISummary
ByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.
Source: ByteDance · new models on Hugging Face · huggingface.coPublished · added here