Contents
- 0.1 parakeet-tdt-0.6b-v3-gguf
- 0.2 Qwen3.5-9B-Uncensored-HauhauCS-Aggressive
- 0.3 gemma-4-26B-A4B-it-qat-GGUF
- 0.4 gpt-oss-20b-GGUF
- 0.5 Qwen3-VL-8B-Instruct-abliterated-GGUF
- 0.6 Flux2-Klein-9B-True-V2
- 0.7 Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
- 0.8 Qwen3-Coder-30B-A3B-Instruct-GGUF
- 0.9 gemma-4-31B-it-GGUF
- 0.10 Gemma-4-E4B-Uncensored-HauhauCS-Aggressive
- 0.11 Qwopus3.6-27B-Coder-Compat-MTP-GGUF
- 0.12 UI-TARS-1.5-7B-GGUF
- 0.13 Huihui-DeepSeek-V4-Flash-abliterated-ds4-GGUF
- 0.14 gemma-4-12B-it-QAT-GGUF
- 0.15 Ternary-Bonsai-27B-gguf
- 0.16 Gemmable-4-12B-MTP-GGUF
- 0.17 Qwopus3.6-35B-A3B-Coder-MTP-GGUF
- 0.18 Qwen3.6-35B-A3B-MTP-GGUF
- 0.19 gemma-4-12b-it-GGUF
- 0.20 Qwen3-VL-30B-A3B-Instruct-GGUF
- 0.21 models-moved
- 0.22 cohere-transcribe-03-2026-gguf
- 0.23 Qwen-AgentWorld-35B-A3B-GGUF
- 0.24 vntl-llama3-8b-v2-gguf
- 1 Hello, Nice to meet you!
parakeet-tdt-0.6b-v3-gguf
0.6B open-weight model from handy-computer for local AI inference.
Qwen3.5-9B-Uncensored-HauhauCS-Aggressive
9B open-weight model from HauhauCS for local AI inference.
gemma-4-26B-A4B-it-qat-GGUF
26B open-weight model from Google for local AI inference.
Qwen3-VL-8B-Instruct-abliterated-GGUF
Qwen3-VL-8B-Instruct-abliterated-GGUF is an 8-billion parameter vision-language model developed by Alibaba that has been fully abliterated to remove proprietary constraints while retaining core capabilities. This model excels at handling complex visual reasoning and multi-step tasks within a conversational framework, making it ideal for open-ended analysis and creative generation in the US region. Running this GGUF quantized version locally is highly practical for users with mid-range GPUs, offering fast inference speeds without requiring specialized enterprise hardware.
Flux2-Klein-9B-True-V2
9B open-weight model from wikeeyang for local AI inference.
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF
The Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF is a 27-billion parameter large language model created by DavidAU that combines multiple fine-tuning techniques including Unsloth and Heretic methods. This model excels at generating uncensored, highly creative, and unrestricted content across diverse topics, making it ideal for users seeking an abliterated assistant capable of multi-stage tuned responses without safety filters. Running this GGUF quantized version locally requires a substantial GPU with ample VRAM to handle the 27B parameter load efficiently, offering best performance on systems equipped with high-end hardware for fast inference speeds.
Qwen3-Coder-30B-A3B-Instruct-GGUF
Qwen3-Coder-30B-A3B-Instruct-GGUF is a 30-billion parameter coding model developed by Alibaba that integrates advanced instruction tuning for specialized tasks. It excels at complex code generation, debugging, and conversational programming assistance, making it ideal for developers seeking high-performance solutions in the US region. Running this model locally requires substantial GPU memory to handle its large parameter count, so it is best suited for users with powerful hardware who prioritize raw coding capability over speed.
Gemma-4-E4B-Uncensored-HauhauCS-Aggressive
Gemma-4-E4B-Uncensored-HauhauCS-Aggressive is a 4-billion parameter multimodal model developed by HauhauCS that integrates advanced capabilities from Gemma 4 with ablated safety filters. This aggressive variant excels at unrestricted text generation, vision analysis, and audio processing, making it ideal for creative tasks requiring raw output without content moderation. Running this model locally demands significant GPU memory to handle its multimodal inputs, and users should expect high computational costs when processing complex audio or video streams.
Qwopus3.6-27B-Coder-Compat-MTP-GGUF
27B open-weight model from Jackrong for local AI inference.
UI-TARS-1.5-7B-GGUF
7B open-weight model from mradermacher for local AI inference.
Huihui-DeepSeek-V4-Flash-abliterated-ds4-GGUF
Unknown open-weight model from huihui-ai for local AI inference.
Ternary-Bonsai-27B-gguf
27B open-weight model from prism-ml for local AI inference.
Gemmable-4-12B-MTP-GGUF
12B open-weight model from Mia-AiLab for local AI inference.
Qwopus3.6-35B-A3B-Coder-MTP-GGUF
35B open-weight model from Jackrong for local AI inference.
Qwen3.6-35B-A3B-MTP-GGUF
35B open-weight model from Alibaba for local AI inference.
Qwen3-VL-30B-A3B-Instruct-GGUF
30B open-weight model from Qwen for local AI inference.
models-moved
The models-moved collection is provided by ggml-org and currently lists an unknown parameter count for its various model files. These models excel at general-purpose tasks within the US region and are best suited for users seeking accessible inference options from this specific provider. Running them locally requires downloading the appropriate GGUF quantization files, with performance varying significantly based on your GPU or CPU hardware capabilities.
cohere-transcribe-03-2026-gguf
Unknown open-weight model from handy-computer for local AI inference.
Qwen-AgentWorld-35B-A3B-GGUF
35B open-weight model from Alibaba for local AI inference.
vntl-llama3-8b-v2-gguf
The vntl-llama3-8b-v2-gguf is an 8B parameter language model developed by lmg-anon that specializes in high-quality translation and conversational tasks. It excels at processing the VNTL-v5-1k dataset to deliver fluent responses, making it ideal for multilingual chat applications and localized content generation. Running this model locally requires a GPU with sufficient VRAM to handle its 8B parameter weight efficiently, ensuring responsive performance for real-time dialogue.
