Summary
Gabriel Jorge Menezes, infrastructure engineer at Krea.ai, gives a detailed technical account of training and serving Krea 2 (K2) — the company’s image generation model trained from scratch with no base checkpoint, released open-source on Hugging Face within the past month. The talk is structured as a practitioner’s lessons-learned report covering every layer of the training stack.
On the training side, Menezes describes scaling from small ablation runs to hundreds of GPUs connected over InfiniBand, and the non-obvious challenges that come with scale: silent failures like NCCL timeouts, GPU thermal throttling, and NVLink errors that require custom metric exporters to detect since NVIDIA doesn’t surface them natively. His core advice is aggressive observability — GPU temperature, InfiniBand packet error rates, NVLink errors — and frequent checkpointing to a high-throughput distributed filesystem capable of 1.8 TB/s reads and sub-terabyte checkpoint writes in under 30 seconds. The team checkpoints every 20–30 minutes, treating crashes as expected events rather than crises.
For serving, Krea uses a gang-scheduled Kubernetes system with a two-tier priority queue: training jobs always preempt inference workloads, but the infrastructure is built so that preempted inference pods drain gracefully to avoid production downtime. Menezes also touches on how researchers interact with the cluster — a job-queue abstraction that hides GPU allocation entirely — and the architectural choices in K2’s diffusion transformer design, deliberately kept simple by porting LLM research techniques into the DiT framework. Both a raw pre-training checkpoint and a post-trained turbo variant (capable of sub-second image generation) are publicly available.
📺 Source: AI Engineer · Published August 18, 2026
🏷️ Format: Deep Dive







