Contents
- 0.1 Gemma-4-E4B-Uncensored-HauhauCS-Aggressive
- 0.2 Qwopus3.6-27B-Coder-Compat-MTP-GGUF
- 0.3 UI-TARS-1.5-7B-GGUF
- 0.4 Huihui-DeepSeek-V4-Flash-abliterated-ds4-GGUF
- 0.5 gemma-4-12B-it-qat-GGUF
- 0.6 Ternary-Bonsai-27B-gguf
- 0.7 Gemmable-4-12B-MTP-GGUF
- 0.8 Qwopus3.6-35B-A3B-Coder-MTP-GGUF
- 0.9 Qwen3.6-35B-A3B-MTP-GGUF
- 0.10 gemma-4-12b-it-GGUF
- 0.11 Qwen3-VL-30B-A3B-Instruct-GGUF
- 0.12 models-moved
- 0.13 cohere-transcribe-03-2026-gguf
- 0.14 Qwen-AgentWorld-35B-A3B-GGUF
- 0.15 vntl-llama3-8b-v2-gguf
- 0.16 gemma-4-E4B-it-GGUF
- 0.17 Qwen3.6-27B-GGUF
- 0.18 Qwen3.6-35B-A3B-GGUF
- 0.19 Hy3-GGUF
- 0.20 HyperCLOVAX-SEED-Text-Instruct-1.5B-Q4_K_M-GGUF
- 0.21 Qwen3.5-9B-GGUF
- 0.22 Qwen3.5-4B-GGUF
- 0.23 parakeet-unified-en-0.6b-gguf
- 0.24 nemotron-3.5-asr-streaming-0.6b-gguf
- 1 Hello, Nice to meet you!
Gemma-4-E4B-Uncensored-HauhauCS-Aggressive
Gemma-4-E4B-Uncensored-HauhauCS-Aggressive is a 4-billion parameter multimodal model developed by HauhauCS that integrates advanced capabilities from Gemma 4 with ablated safety filters. This aggressive variant excels at unrestricted text generation, vision analysis, and audio processing, making it ideal for creative tasks requiring raw output without content moderation. Running this model locally demands significant GPU memory to handle its multimodal inputs, and users should expect high computational costs when processing complex audio or video streams.
Qwopus3.6-27B-Coder-Compat-MTP-GGUF
27B open-weight model from Jackrong for local AI inference.
UI-TARS-1.5-7B-GGUF
7B open-weight model from mradermacher for local AI inference.
Huihui-DeepSeek-V4-Flash-abliterated-ds4-GGUF
Unknown open-weight model from huihui-ai for local AI inference.
Ternary-Bonsai-27B-gguf
27B open-weight model from prism-ml for local AI inference.
Gemmable-4-12B-MTP-GGUF
12B open-weight model from Mia-AiLab for local AI inference.
Qwopus3.6-35B-A3B-Coder-MTP-GGUF
35B open-weight model from Jackrong for local AI inference.
Qwen3.6-35B-A3B-MTP-GGUF
35B open-weight model from Alibaba for local AI inference.
Qwen3-VL-30B-A3B-Instruct-GGUF
30B open-weight model from Qwen for local AI inference.
models-moved
The models-moved collection is provided by ggml-org and currently lists an unknown parameter count for its various model files. These models excel at general-purpose tasks within the US region and are best suited for users seeking accessible inference options from this specific provider. Running them locally requires downloading the appropriate GGUF quantization files, with performance varying significantly based on your GPU or CPU hardware capabilities.
cohere-transcribe-03-2026-gguf
Unknown open-weight model from handy-computer for local AI inference.
Qwen-AgentWorld-35B-A3B-GGUF
35B open-weight model from Alibaba for local AI inference.
vntl-llama3-8b-v2-gguf
The vntl-llama3-8b-v2-gguf is an 8B parameter language model developed by lmg-anon that specializes in high-quality translation and conversational tasks. It excels at processing the VNTL-v5-1k dataset to deliver fluent responses, making it ideal for multilingual chat applications and localized content generation. Running this model locally requires a GPU with sufficient VRAM to handle its 8B parameter weight efficiently, ensuring responsive performance for real-time dialogue.
gemma-4-E4B-it-GGUF
The gemma-4-E4B-it-GGUF is a 4-billion parameter language model developed by Google and optimized for Italian using GGUF quantization. It excels at handling complex reasoning tasks within the Gemma 4 architecture while supporting efficient image-to-text conversions through specialized unsloth optimizations. Running this model locally requires a modern GPU to leverage its speed, making it ideal for users seeking high-performance local inference without cloud dependencies.
Qwen3.6-27B-GGUF
Qwen3.6-27B-GGUF is a 27-billion parameter large language model developed by Alibaba that supports advanced conversational AI and image-to-text capabilities. It excels at complex reasoning tasks, multilingual communication, and visual analysis, making it ideal for professional applications requiring high accuracy in text generation and understanding. Running this model locally typically demands substantial GPU memory and a powerful processor to handle its 27B parameters efficiently, so it is best suited for users with dedicated hardware or access to cloud instances.
Qwen3.6-35B-A3B-GGUF
Qwen3.6-35B-A3B-GGUF is a 35-billion parameter large language model developed by Alibaba that features specialized optimizations for image-to-text conversion and multilingual conversation. This model excels at handling complex visual reasoning tasks and maintaining coherent dialogue, making it ideal for applications requiring deep contextual understanding across diverse topics. Running this model locally requires substantial GPU memory to handle its full parameter count, though quantized GGUF versions can offer a practical balance between speed and performance on high-end consumer hardware.
HyperCLOVAX-SEED-Text-Instruct-1.5B-Q4_K_M-GGUF
HyperCLOVAX-SEED-Text-Instruct-1.5B-Q4_K_M-GGUF is a 1.5 billion parameter language model created by rippertnt and optimized for use with llama-cpp. It excels at conversational tasks and general text generation, making it ideal for lightweight chat applications that require efficient inference on standard hardware. Running this model locally is practical even on modest systems due to its quantized format, offering fast response times without the need for high-end GPUs.
Qwen3.5-9B-GGUF
Qwen3.5-9B-GGUF is a compact large language model developed by Alibaba containing 9 billion parameters. It excels at conversational tasks and image-to-text conversion, making it ideal for lightweight applications that require efficient region US deployment. Running this model locally is practical on consumer-grade hardware thanks to its small footprint, offering fast inference speeds even on modest GPUs when paired with optimization libraries like Unsloth.
Qwen3.5-4B-GGUF
Qwen3.5-4B-GGUF is a compact 4-billion parameter language model developed by Alibaba that supports conversational tasks and image-to-text conversion. It excels at handling lightweight inference workloads while maintaining strong performance in unsloth-optimized environments for text generation. Running this model locally requires minimal hardware resources, making it ideal for users seeking fast, efficient deployment on standard consumer devices.
parakeet-unified-en-0.6b-gguf
0.6B open-weight model from handy-computer for local AI inference.
nemotron-3.5-asr-streaming-0.6b-gguf
0.6B open-weight model from handy-computer for local AI inference.
