Summary
This AI Engineer panel brings together Olive Song, research lead for reinforcement learning at MiniMax, and Dan, VP of Kernels at Together AI, for a technical discussion on the open-source release of MiniMax M3 and what it takes to serve a frontier model at scale. Song explains MiniMax’s decision to release M3 as an open-weight model — the company’s first multimodal release, capable of processing video and images alongside text and code — as aligned with the company’s mission to democratize intelligence and enable community-driven improvements.
Dan details the inference engineering lifecycle at Together AI: from day-zero kernel writing and benchmarking through weeks of iterative optimization on KV cache management, attention kernels, and quantization. He describes a real pace of improvement — performance gains happening not week-over-week but sometimes overnight — and flags a structural shift in the workloads the inference stack must handle. Agentic workloads, where a model makes hundreds of sequential tool calls inside an automated harness and ingests entire codebases as context, require fundamentally different optimization strategies than the short-context conversational chat that defined previous generation inference infrastructure.
The conversation also covers MiniMax M3’s application highlights — computer use, AI-assisted game development, and multimodal coding agents — and the architectural decisions behind making a model that performs well across both text-heavy agentic loops and vision-intensive multimodal tasks. Together AI reports holding the majority of token usage for M3 since launch.
📺 Source: AI Engineer · Published July 31, 2026
🏷️ Format: Interview







