13B Models

Best 13B Local AI Models

Mid-size models in the 10B-19B range - noticeably more capable, still realistic on a single GPU with enough VRAM.

Wan2.2-T2V-A14B-GGUF

QuantStack14B24 GB RAM

The Wan2.2-T2V-A14B-GGUF is a large-scale text-to-video model developed by QuantStack with a total of 14 billion parameters. It excels at generating high-quality video clips from text prompts, making it ideal for creative storytelling and content creation workflows. Running this locally requires significant GPU memory due to its size, so it is best suited for users with high-end hardware or those willing to use quantized GGUF versions for efficiency.

Qwen3.6-14B-A3B-FableVibes-GGUF

tvall4314B24 GB RAM

14B open-weight model from tvall43 for local AI inference.

Qwen3-14B-GGUF

MaziyarPanahi14B24 GB RAM

14B open-weight model from MaziyarPanahi for local AI inference.

Sugoi-14B-Ultra-GGUF

sugoitoolkit14B24 GB RAM

14B open-weight model from sugoitoolkit for local AI inference.

gemma-4-12b-heretic-abliterated-GGUF

culturerevolt12B16 GB RAM

12B open-weight model from culturerevolt for local AI inference.

gemma-4-12B-coder-fable5-composer2.5-v1-GGUF

yuxinlu112B16 GB RAM

12B open-weight model from yuxinlu1 for local AI inference.

gemma-4-12B-it-qat-q4_0-gguf

google12B16 GB RAM

The gemma-4-12B-it-qat-q4_0-gguf is a large language model developed by Google that features approximately 12 billion parameters. It excels at conversational interactions and handles any-to-any queries effectively, making it ideal for customer support bots or general dialogue systems. Running this model locally requires a system with sufficient RAM and

Bielik-11B-v3.0-Instruct-awq

speakleash11B16 GB RAM

11B open-weight model from speakleash for local AI inference.

Wan2.2-I2V-A14B-GGUF

QuantStack14B24 GB RAM

14B open-weight model from QuantStack for local AI inference.

gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-GGUF

yuxinlu112B16 GB RAM

12B open-weight model from yuxinlu1 for local AI inference.

gemma-4-12B-it-qat-GGUF

Google12B16 GB RAM

12B open-weight model from Google for local AI inference.

Gemmable-4-12B-MTP-GGUF

Mia-AiLab12B16 GB RAM

12B open-weight model from Mia-AiLab for local AI inference.

gemma-4-12b-it-GGUF

Google12B16 GB RAM

12B open-weight model from Google for local AI inference.

Hello, Nice to meet you! 👋

Subscribe our newsletter to get the latest AI news.

Malcare WordPress Security