Models

gemma-4-26B-A4B-it-GGUF

Google26B24 GB RAM

The gemma-4-26B-A4B-it-GGUF is a 26-billion parameter large language model developed by Google that leverages advanced instruction tuning for high-performance reasoning. This model excels at complex text generation and multimodal tasks, making it ideal for applications requiring deep contextual understanding and precise instruction following. Running this model locally demands substantial GPU memory and significant compute power, so it is best suited for users with high-end hardware or those utilizing optimized inference frameworks like Unsloth to manage its resource requirements.

Bonsai-27B-gguf

prism-ml27B24 GB RAM

Bonsai-27B-gguf is a compact 27-billion parameter language model developed by prism-ml that utilizes quantization for efficient deployment. It excels at conversational tasks and general reasoning while running smoothly on llama.cpp with support for both CPU and CUDA hardware acceleration. Users can run this model locally on modest hardware, though performance will vary depending on whether they utilize 1-bit quantization or have access to a GPU.

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

HauhauCS35B48 GB RAM

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is a 35-billion parameter large language model developed by HauhauCS that combines MoE architecture with advanced multimodal and vision capabilities. It excels at complex reasoning, multilingual image-text analysis, and generating uncensored content for aggressive or unrestricted use cases where standard safety filters are undesirable. Running this model locally requires substantial high-end GPU memory to handle its sparse mixture-of-experts structure, making it best suited for powerful workstations or servers with significant VRAM available.

Qwythos-9B-Claude-Mythos-5-1M-GGUF

empero-ai9B16 GB RAM

9B open-weight model from empero-ai for local AI inference.

Ornith-1.0-35B-GGUF

ornith-ai35B48 GB RAM

Ornith-1.0-35B-GGUF is a 35-billion parameter language model developed by ornith-ai specifically for conversational tasks within the US region. It excels at maintaining natural dialogue flows and handling context-aware interactions, making it ideal for chatbots and virtual assistants. Running this model locally requires substantial GPU memory to handle its size efficiently, so it is best suited for users with high-end hardware or those willing to use quantized GGUF versions to reduce resource demands.

Qwen3.6-27B-MTP-GGUF

Alibaba27B24 GB RAM

Qwen3.6-27B-MTP-GGUF is a 27-billion parameter large language model developed by Alibaba that supports advanced conversational tasks and image-to-text conversion. It excels at handling complex reasoning and multi-turn dialogues, making it ideal for applications requiring deep contextual understanding and visual analysis. Running this model locally typically demands high-end GPU hardware to manage its substantial memory footprint, though quantized GGUF versions can offer a practical balance between speed and performance on consumer-grade systems.

deepseek-v4-gguf

antirezUnknown8 GB RAM

The deepseek-v4-gguf model is a quantized Mixture-of-Experts architecture provided by antirez with unknown parameter counts available in 2-bit and 4-bit formats. It excels at efficient inference for large language tasks while maintaining high performance through its specialized MoE structure, making it ideal for users needing reduced memory footprints without significant capability loss. Running this model locally requires moderate hardware resources to handle the quantized weights effectively, offering a fast and practical solution for deploying advanced AI capabilities on consumer-grade systems.

Ornith-1.0-9B-GGUF

ornith-ai9B16 GB RAM

Ornith-1.0-9B-GGUF is a 9-billion parameter language model developed by ornith-ai designed for conversational interactions within the US region. It excels at generating natural dialogue and handling casual chat tasks, making it ideal for lightweight virtual assistants or simple customer support bots. Running this model locally requires modest hardware resources, allowing it to operate quickly on consumer-grade GPUs without needing massive data centers.

Hello, Nice to meet you! 👋

Subscribe our newsletter to get the latest AI news.

Malcare WordPress Security