Contents
gemma-4-26B-A4B-it-GGUF
The gemma-4-26B-A4B-it-GGUF is a 26-billion parameter large language model developed by Google that leverages advanced instruction tuning for high-performance reasoning. This model excels at complex text generation and multimodal tasks, making it ideal for applications requiring deep contextual understanding and precise instruction following. Running this model locally demands substantial GPU memory and significant compute power, so it is best suited for users with high-end hardware or those utilizing optimized inference frameworks like Unsloth to manage its resource requirements.
Bonsai-27B-gguf
Bonsai-27B-gguf is a compact 27-billion parameter language model developed by prism-ml that utilizes quantization for efficient deployment. It excels at conversational tasks and general reasoning while running smoothly on llama.cpp with support for both CPU and CUDA hardware acceleration. Users can run this model locally on modest hardware, though performance will vary depending on whether they utilize 1-bit quantization or have access to a GPU.
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is a 35-billion parameter large language model developed by HauhauCS that combines MoE architecture with advanced multimodal and vision capabilities. It excels at complex reasoning, multilingual image-text analysis, and generating uncensored content for aggressive or unrestricted use cases where standard safety filters are undesirable. Running this model locally requires substantial high-end GPU memory to handle its sparse mixture-of-experts structure, making it best suited for powerful workstations or servers with significant VRAM available.
Qwythos-9B-Claude-Mythos-5-1M-GGUF
9B open-weight model from empero-ai for local AI inference.
Ornith-1.0-35B-GGUF
Ornith-1.0-35B-GGUF is a 35-billion parameter language model developed by ornith-ai specifically for conversational tasks within the US region. It excels at maintaining natural dialogue flows and handling context-aware interactions, making it ideal for chatbots and virtual assistants. Running this model locally requires substantial GPU memory to handle its size efficiently, so it is best suited for users with high-end hardware or those willing to use quantized GGUF versions to reduce resource demands.
Qwen3.6-27B-MTP-GGUF
Qwen3.6-27B-MTP-GGUF is a 27-billion parameter large language model developed by Alibaba that supports advanced conversational tasks and image-to-text conversion. It excels at handling complex reasoning and multi-turn dialogues, making it ideal for applications requiring deep contextual understanding and visual analysis. Running this model locally typically demands high-end GPU hardware to manage its substantial memory footprint, though quantized GGUF versions can offer a practical balance between speed and performance on consumer-grade systems.
deepseek-v4-gguf
The deepseek-v4-gguf model is a quantized Mixture-of-Experts architecture provided by antirez with unknown parameter counts available in 2-bit and 4-bit formats. It excels at efficient inference for large language tasks while maintaining high performance through its specialized MoE structure, making it ideal for users needing reduced memory footprints without significant capability loss. Running this model locally requires moderate hardware resources to handle the quantized weights effectively, offering a fast and practical solution for deploying advanced AI capabilities on consumer-grade systems.
Ornith-1.0-9B-GGUF
Ornith-1.0-9B-GGUF is a 9-billion parameter language model developed by ornith-ai designed for conversational interactions within the US region. It excels at generating natural dialogue and handling casual chat tasks, making it ideal for lightweight virtual assistants or simple customer support bots. Running this model locally requires modest hardware resources, allowing it to operate quickly on consumer-grade GPUs without needing massive data centers.
