North Micro Vision: A 2.4B Native-Resolution Vision Model: Run Locally

North Micro Vision: A 2.4B Native-Resolution Vision Model: Run Locally

More

Descriptions:

Fahd Mirza installs and tests North MicroVision, a 2.4 billion parameter vision-language model from Cohere released under the Apache 2.0 license, on a local Ubuntu system with an Nvidia RTX A6000 48GB GPU. The model’s distinguishing feature is native resolution image support — it processes documents, charts, and screenshots at their original dimensions rather than downscaling, which the architecture achieves through a three-stage pipeline: a native-resolution vision encoder, a multimodal projector, and a 2B-parameter language model using interleaved sliding-window and full-attention layers across 28 transformer blocks.

The video walks through real installation issues (requiring transformers installed from source), then puts the model through several practical tests: transcribing multilingual handwritten notes (English, Urdu, Indonesian), reading handwritten physics equations in LaTeX, and parsing a structured printed invoice. Results are mixed and honestly reported — English text and mathematical notation perform well, with the model correctly rendering complex equations like the spacetime interval and Lorentz factor, while Indonesian and Urdu handwriting show significant errors. VRAM consumption sits at just over 5GB when fully loaded.

Mirza also previews an upcoming fine-tuning video, noting the model is small enough to fine-tune on custom data — making it potentially useful for domain-specific document parsing applications. The video is installable from Hugging Face without a gated access request, and Mirza provides a working Python script for inference.


📺 Source: Fahd Mirza · Published August 13, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels

1 Item

Companies