Contents
- 0.1 gemma-4-E4B-it-GGUF
- 0.2 Qwen3.6-27B-GGUF
- 0.3 Qwen3.6-35B-A3B-GGUF
- 0.4 Hy3-GGUF
- 0.5 HyperCLOVAX-SEED-Text-Instruct-1.5B-Q4_K_M-GGUF
- 0.6 Qwen3.5-9B-GGUF
- 0.7 Qwen3.5-4B-GGUF
- 0.8 parakeet-unified-en-0.6b-gguf
- 0.9 nemotron-3.5-asr-streaming-0.6b-gguf
- 0.10 gemma-4-26B-A4B-it-GGUF
- 0.11 Bonsai-27B-gguf
- 0.12 Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
- 0.13 Qwythos-9B-Claude-Mythos-5-1M-GGUF
- 0.14 Ornith-1.0-35B-GGUF
- 0.15 Qwen3.6-27B-MTP-GGUF
- 0.16 deepseek-v4-gguf
- 0.17 Ornith-1.0-9B-GGUF
- 1 Hello, Nice to meet you!
gemma-4-E4B-it-GGUF
The gemma-4-E4B-it-GGUF is a 4-billion parameter language model developed by Google and optimized for Italian using GGUF quantization. It excels at handling complex reasoning tasks within the Gemma 4 architecture while supporting efficient image-to-text conversions through specialized unsloth optimizations. Running this model locally requires a modern GPU to leverage its speed, making it ideal for users seeking high-performance local inference without cloud dependencies.
Qwen3.6-27B-GGUF
Qwen3.6-27B-GGUF is a 27-billion parameter large language model developed by Alibaba that supports advanced conversational AI and image-to-text capabilities. It excels at complex reasoning tasks, multilingual communication, and visual analysis, making it ideal for professional applications requiring high accuracy in text generation and understanding. Running this model locally typically demands substantial GPU memory and a powerful processor to handle its 27B parameters efficiently, so it is best suited for users with dedicated hardware or access to cloud instances.
Qwen3.6-35B-A3B-GGUF
Qwen3.6-35B-A3B-GGUF is a 35-billion parameter large language model developed by Alibaba that features specialized optimizations for image-to-text conversion and multilingual conversation. This model excels at handling complex visual reasoning tasks and maintaining coherent dialogue, making it ideal for applications requiring deep contextual understanding across diverse topics. Running this model locally requires substantial GPU memory to handle its full parameter count, though quantized GGUF versions can offer a practical balance between speed and performance on high-end consumer hardware.
HyperCLOVAX-SEED-Text-Instruct-1.5B-Q4_K_M-GGUF
HyperCLOVAX-SEED-Text-Instruct-1.5B-Q4_K_M-GGUF is a 1.5 billion parameter language model created by rippertnt and optimized for use with llama-cpp. It excels at conversational tasks and general text generation, making it ideal for lightweight chat applications that require efficient inference on standard hardware. Running this model locally is practical even on modest systems due to its quantized format, offering fast response times without the need for high-end GPUs.
Qwen3.5-9B-GGUF
Qwen3.5-9B-GGUF is a compact large language model developed by Alibaba containing 9 billion parameters. It excels at conversational tasks and image-to-text conversion, making it ideal for lightweight applications that require efficient region US deployment. Running this model locally is practical on consumer-grade hardware thanks to its small footprint, offering fast inference speeds even on modest GPUs when paired with optimization libraries like Unsloth.
Qwen3.5-4B-GGUF
Qwen3.5-4B-GGUF is a compact 4-billion parameter language model developed by Alibaba that supports conversational tasks and image-to-text conversion. It excels at handling lightweight inference workloads while maintaining strong performance in unsloth-optimized environments for text generation. Running this model locally requires minimal hardware resources, making it ideal for users seeking fast, efficient deployment on standard consumer devices.
parakeet-unified-en-0.6b-gguf
0.6B open-weight model from handy-computer for local AI inference.
nemotron-3.5-asr-streaming-0.6b-gguf
0.6B open-weight model from handy-computer for local AI inference.
gemma-4-26B-A4B-it-GGUF
The gemma-4-26B-A4B-it-GGUF is a 26-billion parameter large language model developed by Google that leverages advanced instruction tuning for high-performance reasoning. This model excels at complex text generation and multimodal tasks, making it ideal for applications requiring deep contextual understanding and precise instruction following. Running this model locally demands substantial GPU memory and significant compute power, so it is best suited for users with high-end hardware or those utilizing optimized inference frameworks like Unsloth to manage its resource requirements.
Bonsai-27B-gguf
Bonsai-27B-gguf is a compact 27-billion parameter language model developed by prism-ml that utilizes quantization for efficient deployment. It excels at conversational tasks and general reasoning while running smoothly on llama.cpp with support for both CPU and CUDA hardware acceleration. Users can run this model locally on modest hardware, though performance will vary depending on whether they utilize 1-bit quantization or have access to a GPU.
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive is a 35-billion parameter large language model developed by HauhauCS that combines MoE architecture with advanced multimodal and vision capabilities. It excels at complex reasoning, multilingual image-text analysis, and generating uncensored content for aggressive or unrestricted use cases where standard safety filters are undesirable. Running this model locally requires substantial high-end GPU memory to handle its sparse mixture-of-experts structure, making it best suited for powerful workstations or servers with significant VRAM available.
Qwythos-9B-Claude-Mythos-5-1M-GGUF
9B open-weight model from empero-ai for local AI inference.
Ornith-1.0-35B-GGUF
Ornith-1.0-35B-GGUF is a 35-billion parameter language model developed by ornith-ai specifically for conversational tasks within the US region. It excels at maintaining natural dialogue flows and handling context-aware interactions, making it ideal for chatbots and virtual assistants. Running this model locally requires substantial GPU memory to handle its size efficiently, so it is best suited for users with high-end hardware or those willing to use quantized GGUF versions to reduce resource demands.
Qwen3.6-27B-MTP-GGUF
Qwen3.6-27B-MTP-GGUF is a 27-billion parameter large language model developed by Alibaba that supports advanced conversational tasks and image-to-text conversion. It excels at handling complex reasoning and multi-turn dialogues, making it ideal for applications requiring deep contextual understanding and visual analysis. Running this model locally typically demands high-end GPU hardware to manage its substantial memory footprint, though quantized GGUF versions can offer a practical balance between speed and performance on consumer-grade systems.
deepseek-v4-gguf
The deepseek-v4-gguf model is a quantized Mixture-of-Experts architecture provided by antirez with unknown parameter counts available in 2-bit and 4-bit formats. It excels at efficient inference for large language tasks while maintaining high performance through its specialized MoE structure, making it ideal for users needing reduced memory footprints without significant capability loss. Running this model locally requires moderate hardware resources to handle the quantized weights effectively, offering a fast and practical solution for deploying advanced AI capabilities on consumer-grade systems.
Ornith-1.0-9B-GGUF
Ornith-1.0-9B-GGUF is a 9-billion parameter language model developed by ornith-ai designed for conversational interactions within the US region. It excels at generating natural dialogue and handling casual chat tasks, making it ideal for lightweight virtual assistants or simple customer support bots. Running this model locally requires modest hardware resources, allowing it to operate quickly on consumer-grade GPUs without needing massive data centers.
