Summary
Sangwha Lee from Krea.ai presents the research behind Krea 2, the company’s image foundation model, at the AI Engineer conference — coinciding with the open-source release of the Krea 2 medium variant. The talk frames Krea’s design philosophy in contrast to production models like ChatGPT’s image generator and “Nano Banana Pro,” which Lee argues achieve reliability through mode collapse that sacrifices stylistic diversity, always defaulting to a centered, average-looking subject. Krea 2 prioritizes faster generation and broader aesthetic range to serve creative workflows where users are still exploring what they want rather than executing a fixed brief.
The data pipeline section is the most technically detailed portion: at billion-image scale, Krea uses hash-based deduplication (pHash and MD5) as a first pass, then applies embedding-based semantic deduplication via SSCD and SigLip once the dataset is at a manageable size. Caption quality is handled by large vision-language models generating detailed descriptions, which are then rewritten into structured training formats. Lee describes a concrete failure mode where captioners consistently failed to note that a painting was hung against a white wall, causing the model to always generate framed paintings in that context — leading to a filter that undersampled this image type entirely. A key scaling trick: a large LLM is used to develop quality classifiers, then those decisions are distilled into a lightweight SigLip-sized classifier capable of running efficiently over billions of images.
The talk concludes with Lee’s perspective on the most impactful levers for model improvement and promising directions for next-generation image models, with a companion infrastructure talk from a Krea colleague scheduled later in the same conference day.
📺 Source: AI Engineer · Published August 18, 2026
🏷️ Format: Keynote Launch







