Skip to content
Trending storyDeveloping

LlamaIndex makes Markdown its default parser output, using HTML for merged-header tables

2 articles2 sourcesLast article 12h ago ·

Overview

AISummary of 2 articles

LlamaIndex says Markdown is its default output for document parsing because it keeps headings, lists, and tables intact and stays readable for debugging, and that it switches to HTML for tables with merged headers, which Markdown cannot represent cleanly.

The company argues a parser can capture every word yet lose which column a number belongs to, forcing a model to guess.

A later post by Jerry Liu, LlamaIndex's cofounder, calls Markdown a universal representation between humans and agents, noting that most unstructured documents are not natively Markdown, so the main challenge is a translation layer that converts them into Markdown. Both claims come from social media posts; no performance measurements or comparative results are reported.

Written by AI from the articles below · updated Oct 8, 8:56 PM ET

Check the sources:

Developments

2 developments

  1. Oct 8, 1:52 PM ET · 1 article
    LlamaIndex argues Markdown is the universal format for agents
    Jerry Liu: LlamaIndex argues Markdown is the universal format for agents
  2. Oct 8, 12:49 PM ET · 1 article
    Markdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number belongs to. Then your model has to guess.
    LlamaIndex: Markdown is all you need.

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Jerry Liu
    LlamaIndex argues Markdown is the universal format for agents

    AILlamaIndex says Markdown has become a universal representation between humans and agents, preserving headings, lists, and tables while remaining readable to models. Since most unstructured documents are not natively in Markdown, the main challenge is the translation layer, which the company addresses with models that convert document containers into Markdown. The quoted post adds that Markdown keeps table columns intact, with HTML used for tables with merged headers.

  2. LlamaIndex
    Markdown is all you need.

    AIMostly.) A parser can get every word on the page right and still lose which column a number belongs to. Then your model has to guess. Markdown keeps headings, lists, and tables intact, and it stays readable when you're debugging a bad answer. For tables with merged headers, we switch to HTML. Read our breakdown on why it's our default output for parsing below! ⬇️

Heat trend

Not enough continuous observations to show a trend yet.