EvoQuality: ByteDance's self-evolving VLM for image quality assessment without human labels
Original titleByteDance/EvoQuality
AISummary
EvoQuality is a ByteDance vision-language model for no-reference image quality assessment that generates pseudo-ranking labels through pairwise majority voting and refines them with GRPO, requiring no human-annotated quality scores.
On the paper's setting, it raised weighted-average PLCC from 0.615 to 0.770 and SRCC from 0.570 to 0.726 over its Qwen2.5-VL-7B backbone. The model is recommended for research and pre-production assessment, not as the sole criterion for high-stakes decisions.
Source: ByteDance · new models on Hugging Face · huggingface.coPublished · added here