Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI

More

Descriptions:

Ari Morcos, CEO and co-founder of DatologyAI, opens the data quality track at AI Engineer with a data-rich argument that better training data is the most underleveraged lever in AI — especially as compute grows scarcer. He notes that H100 GPU prices have risen roughly 40% from recent lows, reasoning models consume eight times as many tokens as non-reasoning models with a projected 5x further increase within a year, and both Google (capping Meta’s Gemini usage) and OpenAI (selling token futures) are showing early signs of inference rationing at frontier scale.

Against that backdrop, Morcos presents DatologyAI’s four-C curation framework — clean, curate, create, and compose — and shares benchmark results showing that curation alone delivers approximately 14 absolute percentage points of improvement on downstream tasks while requiring 145x less training compute. A curated model can match Qwen 3.5 4B performance at a small fraction of the training budget, and data-curated models also produce significantly more concise outputs, reducing per-answer inference cost by roughly 35x on equivalent tasks.

The talk extends to multilingual performance, demonstrating that strong cross-lingual MMLU results are achievable with only 8% multilingual tokens in the training mix — most languages capped at 6 billion tokens — challenging common assumptions about data volume requirements for non-English capability. Morcos positions data curation not as preprocessing overhead but as a structural compute multiplier that determines how steep the performance curve gets per training dollar.


📺 Source: AI Engineer · Published July 31, 2026
🏷️ Format: Deep Dive

1 Item

Channels