This AI Model Has No VAE! Testing HiDream-O1’s Unified Transformer

This AI Model Has No VAE! Testing HiDream-O1’s Unified Transformer

More

Summary

HiDream-O1 is a newly released open-source image generation model from the HiDream team that takes a fundamentally different architectural approach from conventional diffusion models. Rather than relying on a separate VAE and text encoder, it operates as a pixel-level unified transformer — processing images, text, and reference conditions within a single 8-billion-parameter model. In ComfyUI, the CLIP and VAE nodes present in its workflow are placeholder components required for framework compatibility, not functional parts of the pipeline.

The video, from the Veteran AI channel, compares the official Dev (distilled, 28 steps) and Full (50 steps) versions side by side using RunningHub as the hosting environment. While the Dev version is significantly faster, its texture quality and image fidelity fall noticeably short in detailed or high-resolution scenes. The host presents an unconventional middle-ground workflow: running the Full model at only 28 sampling steps with the Euler sampler and a CFG between 4–5, which achieves substantially better quality than the Dev version without the prohibitive generation times of the standard 50-step Full workflow. A seam-smoothing patch node is flagged as essential for high-resolution tile-based generation.

Additional capabilities covered include image editing, multi-reference personalized generation, and a color-coded position map technique drawn from the official HiDream paper that controls where specific reference faces appear in the final composition. Model formats available include BF16, FP8 scaled, and MX FP8 — the last providing roughly 30% faster speeds on RTX 50-series GPUs. All versions must be placed in ComfyUI’s checkpoints folder rather than the diffusion models folder.


📺 Source: Veteran AI · Published May 21, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels