Skip to content
Trending storyDeveloping

Tencent open-sources Youtu-Parsing-Omni 5B omni-modal parsing model with weights on Hugging Face

1 article1 sourcesince Oct 9Last article 3h ago ·

Overview

AISummary of 1 article

Tencent has released the weights of Youtu-Parsing-Omni, a 5-billion-parameter model that parses documents and media, along with a vLLM plugin and inference examples.

The model converts layout, text, tables, formulas, speech (ASR), OCR, and video segments into a single structured JSON output. According to Tencent's own announcement, it scores 96.96 Overall on OmniDocBench, the highest among the models it compared against. These benchmark results are Tencent's claims and have not been independently verified in the reporting.

Written by AI from the articles below · updated Oct 9, 4:51 AM ET

Check the sources:

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 9
  1. Tencent · new models on Hugging Face
    Tencent Releases Youtu-Parsing-Omni, a 5B Omni-Modal Document and Media Parsing Model

    AITencent has open-sourced Youtu-Parsing-Omni, a 5B-parameter omni-modal model that outputs a single structured JSON covering layout, text, tables, formulas, ASR, OCR, and video segments. It scores 96.96 Overall on OmniDocBench, the highest among the compared models, and ships with weights on Hugging Face, a vLLM plugin, and inference examples.

Heat trend

Not enough continuous observations to show a trend yet.