PixelRAG Locally: RAG That Reads Screenshots Instead of Text

PixelRAG Locally: RAG That Reads Screenshots Instead of Text

More

Summary

Standard RAG pipelines fail silently on documents with complex layouts: when HTML-to-text parsers mangle tables and charts, the retriever never sees the structured content, and the model returns “cannot be determined” even when the answer is plainly visible on the page. PixelRAG, a research project from UC Berkeley, sidesteps this entirely by rendering documents as screenshot tiles and searching over images — treating pages the way a human eye does rather than the way a text parser does.

Fahd Mirza builds the full pipeline locally on Ubuntu with an NVIDIA RTX A6000, walking through every stage: converting a Wikipedia page on the Terracotta Army to PDF via VisiPrint, rendering to tiles at 100 DPI with a headless browser (staying under the 875-pixel width limit), chunking tiles, embedding them with Qwen VL (a 2B vision-language model downloaded from Hugging Face), and indexing in FAISS. At query time, the question is embedded into the same visual space, the nearest tile is retrieved, and the vision model reads the answer directly off the screenshot. The GPU requirement for the embedding step is modest — around 8GB VRAM — though the full setup benefits from more.

Importantly, Mirza documents and fixes multiple bugs in the original Berkeley repo before getting the pipeline running, including a type mismatch between string and integer article IDs that causes consistent failures. For practitioners building document intelligence systems over financial reports, regulatory filings, or any business document with dense tables, PixelRAG offers a practical alternative to OCR-dependent extraction that preserves visual layout information through the entire retrieval chain.


📺 Source: Fahd Mirza · Published June 23, 2026
🏷️ Format: Hands On Build

1 Item

Channels

1 Item

People