Qwen3.8-27B & How to Serve it Fast

Qwen3.8-27B & How to Serve it Fast

More

Descriptions:

In this video, I look at the long awaited Qwen3.8-27B model. Both what it can do and how to serve it at the maximum tokens per second

Thanks to Dell for Sponsoring the Compute
#DellProPrecision #DellProMax #DellTech #NVIDIA

📖 Website: https://qwen.ai/
🤗 HF: https://huggingface.co/collections/Qwen/qwen38
SGLang: https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B

Twitter: https://x.com/Sam_Witteveen

🕵️ Interested in building LLM Agents? Fill out the form below
Building LLM Agents Form: https://drp.li/dIMes

👨‍💻Github:
https://github.com/samwit/llm-tutorials

⏱️Time Stamps:
00:00 Intro
00:50 ThinkingCap
01:25 Qwen3.8 – 27B
01:59 Different Versions on Hugging Face
02:12 Benchmarks
03:14 Artificial Analysis Benchmark
04:10 Qwen3.8-27B on Hugging Face
06:40 Demo
14:17 SGLang

1 Item

Channels