by empero-ai
Qwen3.8-4B-Distill-GGUF is a 4 billion parameter large language model provided by empero-ai designed for efficient deployment. It excels at reasoning tasks thanks to its distillation process and works well for local inference via llama.cpp. Running this quantized version locally requires minimal hardware resources, making it suitable for laptops or systems with limited GPU memory.
Parameters
4B
RAM Required
8 GB
Context
4,096
♥ 121 people have liked this model on HuggingFace
⇩ 752,514 downloads on HuggingFace
How to Get This Model
Ollama
ollama pull qwen3.8-4b-distill:4b
HuggingFace
View model page →
LM Studio
Search "Qwen3.8-4B-Distill-GGUF" in LM Studio's Discover tab, or download the GGUF above.
