Find the Perfect Local AI Model

Browse open-weight models with hardware requirements, benchmarks, and one-click Ollama and LM Studio commands.

What Can I Run?

Move the slider to your available RAM and see which models fit.

16 GB

Trending Models

Qwen3-Coder-30B-A3B-Instruct-GGUF

Alibaba30B24 GB RAM

Qwen3-Coder-30B-A3B-Instruct-GGUF is a 30-billion parameter coding model developed by Alibaba that integrates advanced instruction tuning for specialized tasks. It excels at complex code generation, debugging, and conversational programming assistance, making it ideal for developers seeking high-performance solutions in the US region. Running this model locally requires substantial GPU memory to handle its large parameter count, so it is best suited for users with powerful hardware who prioritize raw coding capability over speed.

Ornith-1.5-9B-GGUF

ornith-ai9B16 GB RAM

9B open-weight model from ornith-ai for local AI inference.

Ornith-1.0-9B-GGUF

ornith-ai9B16 GB RAM

Ornith-1.0-9B-GGUF is a 9-billion parameter language model developed by ornith-ai designed for conversational interactions within the US region. It excels at generating natural dialogue and handling casual chat tasks, making it ideal for lightweight virtual assistants or simple customer support bots. Running this model locally requires modest hardware resources, allowing it to operate quickly on consumer-grade GPUs without needing massive data centers.

Ornith-1.5-35B-A3B-GGUF

ornith-ai35B48 GB RAM

35B open-weight model from ornith-ai for local AI inference.

Ornith-1.0-35B-GGUF

ornith-ai35B48 GB RAM

Ornith-1.0-35B-GGUF is a 35-billion parameter language model developed by ornith-ai specifically for conversational tasks within the US region. It excels at maintaining natural dialogue flows and handling context-aware interactions, making it ideal for chatbots and virtual assistants. Running this model locally requires substantial GPU memory to handle its size efficiently, so it is best suited for users with high-end hardware or those willing to use quantized GGUF versions to reduce resource demands.

Qwen3.8-27B-iMatrix-NVFP4-MTP-GGUF

cdiamond27B24 GB RAM

27B open-weight model from cdiamond for local AI inference.

Browse by Size

Works With Your Tools

✓ Works with Ollama

✓ Works with LM Studio

✓ GGUF format, ready to run

Hello, Nice to meet you! 👋

Subscribe our newsletter to get the latest AI news.

Malcare WordPress Security