Structuring the Unstructured – Cedric Clyburn, Red Hat

Structuring the Unstructured – Cedric Clyburn, Red Hat

More

Summary

Cedric Clyburn, an open source engineer at Red Hat, makes the case that unstructured document processing is the single most important — and most underestimated — bottleneck in enterprise AI pipelines. The talk opens with a striking example: 20 scientific papers now contain a nonsensical merged term because AI misread a scanned PDF, conflating two words from adjacent columns. Researchers using LLMs to assist their writing have since cited this error, propagating it further.

The session focuses on Docling, an open-source Linux Foundation tool that converts PDFs, presentations, contracts, scanned documents, and invoices into clean markdown or structured JSON. Unlike naive PDF parsers that produce truncated, column-merged text unfit for RAG, Docling preserves table structure, image captions, and section hierarchy — and can be run locally without sending data to external servers. Clyburn walks through a live Jupyter notebook demo using Docling’s own research paper as a test document, covering basic PDF conversion, multi-page table extraction, image annotation via vision-language models, and structured field extraction using Pydantic schemas.

Installable via pip, Docling is presented as foundational infrastructure for any team building RAG applications, fine-tuning pipelines, or agent systems that ingest real-world enterprise documents. Clyburn’s central argument: model choice and retrieval strategy matter far less than the accuracy of the data at the point of ingestion.


📺 Source: AI Engineer · Published June 28, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels