Summary
Microsoft’s Lens is a compact 3.8-billion-parameter text-to-image model that punches above its weight class, capable of generating images at 1440p resolution and handling a wide aspect ratio range from 1:2 to 2:1. This video from Veteran AI walks through everything needed to run Lens inside ComfyUI, from downloading the correct model files on the Comfy-Org/Lens Hugging Face page to configuring the GPT-OSS text encoder (in NVFP4 format) and the FLUX2 VAE.
The tutorial explains the two available model variants โ the standard version requiring approximately 20 sampling steps and the Turbo version needing only four โ and covers the critical workflow details that trip up new users, including setting the CLIP type to “Lens,” correctly linking resolution values to both EmptyLatentImage and ModelSamplingFlux nodes, and configuring the CFGNorm node with strength 1.0 and pre_cfg enabled. The workflow is demonstrated on RunningHub, a cloud ComfyUI platform that tracks new model support quickly.
Five prompt tests evaluate Lens across realistic photography (a detailed watchmaker scene with neon reflections), Chinese-language prompts (a rainy Chongqing street), and additional scenarios covering text rendering, creative concept art, and more. Results show strong detail density and good handling of multilingual inputs, with the presenter noting both strengths and limitations honestly. Viewers who want a concrete, reproducible setup guide for Lens in ComfyUI will find specific sampler settings, resolution configurations, and honest output comparisons throughout.
๐บ Source: Veteran AI ยท Published May 28, 2026
๐ท๏ธ Format: Tutorial Demo







