Tencent open-sources Youtu-Parsing-Omni 5B omni-modal parsing model with weights on Hugging Face
Overview
Tencent has released the weights of Youtu-Parsing-Omni, a 5-billion-parameter model that parses documents and media, along with a vLLM plugin and inference examples.
The model converts layout, text, tables, formulas, speech (ASR), OCR, and video segments into a single structured JSON output. According to Tencent's own announcement, it scores 96.96 Overall on OmniDocBench, the highest among the models it compared against. These benchmark results are Tencent's claims and have not been independently verified in the reporting.
Written by AI from the articles below · updated Oct 9, 4:51 AM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Tencent · new models on Hugging FaceTencent Releases Youtu-Parsing-Omni, a 5B Omni-Modal Document and Media Parsing Model
AITencent has open-sourced Youtu-Parsing-Omni, a 5B-parameter omni-modal model that outputs a single structured JSON covering layout, text, tables, formulas, ASR, OCR, and video segments. It scores 96.96 Overall on OmniDocBench, the highest among the compared models, and ships with weights on Hugging Face, a vLLM plugin, and inference examples.
Heat trend
Not enough continuous observations to show a trend yet.