Models

Qwen3.8-2B-Distill-GGUF

empero-ai2B4 GB RAM

2B open-weight model from empero-ai for local AI inference.

Qwen3.8-4B-Distill-GGUF

empero-ai4B8 GB RAM

Qwen3.8-4B-Distill-GGUF is a 4 billion parameter large language model provided by empero-ai designed for efficient deployment. It excels at reasoning tasks thanks to its distillation process and works well for local inference via llama.cpp. Running this quantized version locally requires minimal hardware resources, making it suitable for laptops or systems with limited GPU memory.

Qwen3.8-27B-GSQ-RCO-GGUF

ISTA-DASLab27B24 GB RAM

27B open-weight model from ISTA-DASLab for local AI inference.

Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NEO-CODER-MAX-MTP-GGUF

DavidAU27B24 GB RAM

27B open-weight model from DavidAU for local AI inference.

glm-4-9b-chat-IMat-GGUF

legraphista9B16 GB RAM

9B open-weight model from legraphista for local AI inference.

Qwen3.8-27B-OBLITERATED

OBLITERATUS27B24 GB RAM

27B open-weight model from OBLITERATUS for local AI inference.

Ornith-1.5-397B-GGUF

ornith-ai397B80 GB RAM

397B open-weight model from ornith-ai for local AI inference.

Qwen3.8-Flash-Next-GGUF

AlibabaUnknown8 GB RAM

Unknown open-weight model from Alibaba for local AI inference.

Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF

0bserverx27B24 GB RAM

27B open-weight model from 0bserverx for local AI inference.

Qwen3.8-27B-Uncensored-HauhauCS-Aggressive-MTP-GGUF

HauhauCS27B24 GB RAM

27B open-weight model from HauhauCS for local AI inference.

Huihui-Qwen3.8-27B-abliterated-GGUF

huihui-ai27B24 GB RAM

27B open-weight model from huihui-ai for local AI inference.

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

cdiamond27B24 GB RAM

27B open-weight model from cdiamond for local AI inference.

Ornith-1.5-35B-A3B-GGUF

ornith-ai35B48 GB RAM

35B open-weight model from ornith-ai for local AI inference.

Ornith-1.5-9B-GGUF

ornith-ai9B16 GB RAM

9B open-weight model from ornith-ai for local AI inference.

Mistral-Small-24B-Instruct-2501-GGUF

MaziyarPanahi24B24 GB RAM

24B open-weight model from MaziyarPanahi for local AI inference.

Llama-3-8B-Instruct-32k-v0.1-GGUF

MaziyarPanahi8B16 GB RAM

8B open-weight model from MaziyarPanahi for local AI inference.

SmolLM2-135M-GGUF

QuantFactoryUnknown8 GB RAM

Unknown open-weight model from QuantFactory for local AI inference.

Qwen3.6-35B-A3B-NVFP4-MTP-GGUF

michaelw999935B48 GB RAM

35B open-weight model from michaelw9999 for local AI inference.

Llama-3.3-70B-Instruct-GGUF

MaziyarPanahi70B48 GB RAM

70B open-weight model from MaziyarPanahi for local AI inference.

gemma-3-4b-it-GGUF

MaziyarPanahi4B8 GB RAM

4B open-weight model from MaziyarPanahi for local AI inference.

Mistral-7B-Instruct-v0.3-GGUF

MaziyarPanahi7B8 GB RAM

7B open-weight model from MaziyarPanahi for local AI inference.

Mixtral-8x22B-v0.1-GGUF

MaziyarPanahi22B24 GB RAM

22B open-weight model from MaziyarPanahi for local AI inference.

Qwen3.8_4B_Distilled_GGUF

Ma7ee74B8 GB RAM

4B open-weight model from Ma7ee7 for local AI inference.

1 2 3 … 9

Hello, Nice to meet you! 👋

Subscribe our newsletter to get the latest AI news.

Malcare WordPress Security