Summary
Fahd Mirza walks through a complete, terminal-first fine-tuning workflow for Qwen 3.8’s 27 billion parameter model, running entirely on a single GPU with 48 GB of VRAM using the Unsloth library. The tutorial is structured end-to-end: creating a custom JSONL dataset from scratch, installing dependencies, attaching LoRA adapters, watching loss drop during training, exporting to GGUF format, and running a before-and-after inference test to verify behavioral change.
Mirza takes time to explain the mechanics behind parameter-efficient fine-tuning for viewers unfamiliar with the internals. He covers why full fine-tuning is impractical at 27B scale, how LoRA freezes the base model and inserts small trainable adapter layers (rank r=16 in this example) at key attention and feed-forward positions, and how QLoRA stacks additional savings by first compressing the frozen model to 4-bit precision before training begins — together reducing memory consumption enough to fit the model on commodity hardware. The training stack uses Unsloth alongside HuggingFace TRL’s SFT trainer, with the dataset structured as instruction-input-output triples.
Practical guidance throughout includes dataset sizing recommendations (50+ examples for demos, 500–1,000+ for production), a tip to use local LLMs for synthetic data generation, and notes on the Unsloth alternatives (TRL with PEFT, ExoTAL, TorchTune) for those who prefer different tooling. The tutorial is accessible to practitioners with a single high-end consumer or workstation GPU.
📺 Source: Fahd Mirza · Published August 26, 2026
🏷️ Format: Tutorial Demo







