Find the Perfect Local AI Model

Browse open-weight models with hardware requirements, benchmarks, and one-click Ollama and LM Studio commands.

What Can I Run?

Move the slider to your available RAM and see which models fit.

16 GB

Trending Models

Qwen3-Coder-30B-A3B-Instruct-GGUF

Alibaba30B24 GB RAM

Qwen3-Coder-30B-A3B-Instruct-GGUF is a 30-billion parameter coding model developed by Alibaba that integrates advanced instruction tuning for specialized tasks. It excels at complex code generation, debugging, and conversational programming assistance, making it ideal for developers seeking high-performance solutions in the US region. Running this model locally requires substantial GPU memory to handle its large parameter count, so it is best suited for users with powerful hardware who prioritize raw coding capability over speed.

Ornith-1.0-9B-GGUF

ornith-ai9B16 GB RAM

Ornith-1.0-9B-GGUF is a 9-billion parameter language model developed by ornith-ai designed for conversational interactions within the US region. It excels at generating natural dialogue and handling casual chat tasks, making it ideal for lightweight virtual assistants or simple customer support bots. Running this model locally requires modest hardware resources, allowing it to operate quickly on consumer-grade GPUs without needing massive data centers.

Ornith-1.0-35B-GGUF

ornith-ai35B48 GB RAM

Ornith-1.0-35B-GGUF is a 35-billion parameter language model developed by ornith-ai specifically for conversational tasks within the US region. It excels at maintaining natural dialogue flows and handling context-aware interactions, making it ideal for chatbots and virtual assistants. Running this model locally requires substantial GPU memory to handle its size efficiently, so it is best suited for users with high-end hardware or those willing to use quantized GGUF versions to reduce resource demands.

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF

DavidAU27B24 GB RAM

The Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF is a 27-billion parameter large language model created by DavidAU that combines multiple fine-tuning techniques including Unsloth and Heretic methods. This model excels at generating uncensored, highly creative, and unrestricted content across diverse topics, making it ideal for users seeking an abliterated assistant capable of multi-stage tuned responses without safety filters. Running this GGUF quantized version locally requires a substantial GPU with ample VRAM to handle the 27B parameter load efficiently, offering best performance on systems equipped with high-end hardware for fast inference speeds.

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

HauhauCS35B48 GB RAM

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is a 35-billion parameter large language model developed by HauhauCS that combines MoE architecture with advanced multimodal and vision capabilities. It excels at complex reasoning, multilingual image-text analysis, and generating uncensored content for aggressive or unrestricted use cases where standard safety filters are undesirable. Running this model locally requires substantial high-end GPU memory to handle its sparse mixture-of-experts structure, making it best suited for powerful workstations or servers with significant VRAM available.

nemotron-3.5-asr-streaming-0.6b-gguf

handy-computer0.6B2 GB RAM

0.6B open-weight model from handy-computer for local AI inference.

Browse by Size

Works With Your Tools

✓ Works with Ollama

✓ Works with LM Studio

✓ GGUF format, ready to run

Hello, Nice to meet you! 👋

Subscribe our newsletter to get the latest AI news.

Malcare WordPress Security