Qwen 3.8 27B vs DeepSeek V4 Flash — I Built the Same Apps With Both Locally

Qwen 3.8 27B vs DeepSeek V4 Flash — I Built the Same Apps With Both Locally

More

Summary

Bart Slodyczka runs a hands-on local comparison between two freshly available open-weight models — Qwen 3.827B and DeepSeek V4 Flash 0731 — both downloaded and running on his Mac Studio. Using the pi.dev agent harness with carefully controlled single-prompt runs, he tests each model against three progressively harder coding tasks: a single-page weather dashboard using a free weather API, a tower defense browser game, and an Excel-like spreadsheet with clickable cells and formula evaluation.

Results are competitive throughout. Qwen edges out DeepSeek on the weather dashboard by offering multi-city disambiguation and cleaner UX details, while DeepSeek produces a slightly better visual layout for the tower defense game — Slodyczka calls that round roughly 50/50. The spreadsheet proves the most revealing test: both models struggle on the first pass without reasoning enabled, requiring multiple re-prompting cycles. Slodyczka then reruns with thinking set to low for both, demonstrating measurable improvements in formula correctness and cell interaction behavior.

The video is a practical resource for developers choosing between capable mid-size local models for code-generation tasks. Rather than relying on official benchmarks, it shows real model behavior under constrained, reproducible conditions — single-prompt runs, no cherry-picking, and explicit documentation of failure modes before and after enabling reasoning.


📺 Source: Bart Slodyczka · Published August 16, 2026
🏷️ Format: Comparison

1 Item

Channels

1 Item

People