Qwen3.8-4B-Distill-GGUF

by empero-ai

Qwen3.8-4B-Distill-GGUF is a 4 billion parameter large language model provided by empero-ai designed for efficient deployment. It excels at reasoning tasks thanks to its distillation process and works well for local inference via llama.cpp. Running this quantized version locally requires minimal hardware resources, making it suitable for laptops or systems with limited GPU memory.

Parameters 4B
RAM Required 8 GB
Context 4,096

♥ 121 people have liked this model on HuggingFace

⇩ 752,514 downloads on HuggingFace

How to Get This Model

Ollama ollama pull qwen3.8-4b-distill:4b
HuggingFace View model page →
LM Studio Search "Qwen3.8-4B-Distill-GGUF" in LM Studio's Discover tab, or download the GGUF above.