Tenzro
Models

The model registry.

Every model the network serves, across every modality. Each entry is resolved from the live catalog and checked against its source repository on Hugging Face.
Models
183
Modalities
10
Source-verified
183/183
Gated upstream
10

Language

107 models
NVIDIA Cosmos-Reason2 8B (video reasoning VLM)
cosmos-reason · 8B

NVIDIA Cosmos-Reason2 8B — video reasoning VLM for Physical AI, post-trained on Qwen3-VL-8B-Instruct. Takes video or images plus text and answers in text, with spatial-temporal understanding and object detection. 256K context. Needs the mmproj projector for the vision path; NVIDIA's own repo is gated and requires HF_TOKEN, this GGUF conversion is not.

NVIDIA Open Model License4.7 GB256K ctx
DeepSeek V3 0324 (MoE)
deepseek · 685B (MoE, 37B active)

DeepSeek V3 MoE — 685B total, 37B active, 128K context. Native Multi-Token-Prediction head (n=4, ~80% accept rate, ~1.8× decode speedup per DeepSeek tech report). Retired by upstream after 2026-07-24 in favor of DeepSeek V4.

MIT351.9 GB128K ctx
DeepSeek V4 Flash (MoE)
deepseek · 284B (MoE, 13B active)

DeepSeek V4 Flash 0731 — 284B total / 13B active MoE; 1M context. Hybrid Compressed Sparse Attention (CSA) + Heavily Compressed Attention (HCA). Quantization-aware-trained: routed experts (96% of the model) natively MXFP4, the rest FP8/BF16. Outperforms V4-Pro Preview. MTP head built into the model file.

MIT144.4 GB1024K ctx
DiffusionGemma 26B-A4B
diffusiongemma · 26B (4B active)

Diffusion-generation Gemma 4 26B-A4B. Generates 256-token canvases by parallel denoising rather than autoregressive sampling. Unsloth: 2000+ t/s on RTX 6000. Requires Unsloth Studio or llama.cpp PR #24423+.

Gemma License16.8 GB32K ctx
Gemma 3 12B
gemma3 · 12B

High-performance instruction-tuned model from Google

Gemma License6.8 GB128K ctx
Gemma 3 1B
gemma3 · 1B

Google's compact instruction-tuned model

Gemma License768.7 MB32K ctx
Gemma 3 270M
gemma3 · 270M

Tiny Gemma model for ultra-lightweight on-device inference

Gemma License241.4 MB32K ctx
Gemma 3 27B
gemma3 · 27B

Google's largest Gemma model with exceptional capabilities

Gemma License15.4 GB128K ctx
Gemma 3 4B
gemma3 · 4B

Extended context Gemma model for chat applications

Gemma License2.3 GB128K ctx
Gemma 4 12B
gemma4 · 12B

Google's mid-tier dense Gemma 4 model (128K context). MTP-enabled — pairs with `gemma4-12b-mtp-draft` for 1.5–2.2× throughput on the same hardware (Unsloth: 52 → 162 t/s at Q4 on a 4090).

Gemma License7.1 GB128K ctx
Gemma 4 12B MTP Drafter
gemma4 · MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 12B. Pair with the 12B target via `--spec-type draft-mtp`.

Gemma License572.2 MB128K ctx
Gemma 4 12B (QAT)
gemma4 · 12B

Quantization-Aware-Trained Gemma 4 12B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-12b-mtp-draft`.

Gemma License7.4 GB128K ctx
Gemma 4 26B-A4B (MoE)
gemma4 · 26B (4B active)

Gemma 4 Mixture-of-Experts: 26B total params, 4B active per token (128K context). MTP-enabled — pairs with `gemma4-26b-a4b-mtp-draft`; expect ~1.15–1.2× speedup on MoE targets per Unsloth.

Gemma License16.9 GB128K ctx
Gemma 4 26B-A4B MTP Drafter (MoE)
gemma4 · MTP head

Google's jointly-trained Multi-Token Prediction head for the Gemma 4 26B-A4B Mixture-of-Experts target. Pair via `--spec-type draft-mtp`; Unsloth measures ~1.15–1.2× speedup on MoE targets vs ~1.4–2.2× on dense.

Gemma License1.1 GB128K ctx
Gemma 4 26B-A4B (QAT, MoE)
gemma4 · 26B (4B active)

Quantization-Aware-Trained Gemma 4 26B-A4B MoE. Unsloth measures 85.6% MMLU top-1 vs 70.2% on naive Q4 (+15.4 points). MTP-enabled via `gemma4-26b-a4b-mtp-draft`.

Gemma License17.0 GB128K ctx
Gemma 4 31B
gemma4 · 31B

Google's largest dense Gemma 4 model (128K context). MTP-enabled — pairs with `gemma4-31b-mtp-draft` for ~2× throughput at 101 t/s on consumer GPUs (Unsloth benchmark).

Gemma License18.3 GB128K ctx
Gemma 4 31B MTP Drafter
gemma4 · MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 31B. Pair with the 31B target via `--spec-type draft-mtp`.

Gemma License1.4 GB128K ctx
Gemma 4 31B (QAT)
gemma4 · 31B

Quantization-Aware-Trained Gemma 4 31B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-31b-mtp-draft`.

Gemma License18.4 GB128K ctx
Gemma 4 E2B
gemma4 · E2B

Google's compact Gemma 4 multimodal model (text + image, 128K context). MTP-enabled — pairs with `gemma4-e2b-mtp-draft` for 1.5–2.2× throughput.

Gemma License3.1 GB128K ctx
Gemma 4 E2B MTP Drafter
gemma4 · MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 E2B. Pair with the E2B target via `--spec-type draft-mtp`.

Gemma License93.3 MB128K ctx
Gemma 4 E2B (QAT)
gemma4 · E2B

Quantization-Aware-Trained Gemma 4 E2B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-e2b-mtp-draft`.

Gemma License3.3 GB128K ctx
Gemma 4 E4B
gemma4 · E4B

Google's efficient Gemma 4 multimodal model (text + image, 128K context). MTP-enabled — pairs with `gemma4-e4b-mtp-draft` for 1.5–2.2× throughput.

Gemma License5.0 GB128K ctx
Gemma 4 E4B MTP Drafter
gemma4 · MTP head

Google's jointly-trained Multi-Token Prediction head for Gemma 4 E4B. Pair with the E4B target via `--spec-type draft-mtp`.

Gemma License286.1 MB128K ctx
Gemma 4 E4B (QAT)
gemma4 · E4B

Quantization-Aware-Trained Gemma 4 E4B. Higher quality than naive Q4 at the same size. MTP-enabled via `gemma4-e4b-mtp-draft`.

Gemma License5.1 GB128K ctx
GLM-5 (MoE)
glm · 744B (MoE, 40B active)

Z.ai GLM-5 — 744B total parameter MoE, 40B active, trained on 28.5T tokens. Routes each token to 8 of 256 experts plus 1 shared across 75 MoE layers, with DeepSeek Sparse Attention over a 198K context. Unsloth dynamic UD-Q4_K_XL GGUF (sharded).

MIT372.5 GB198K ctx
GLM-5.1 (MoE)
glm · 744B (MoE, 40B active)

Z.ai GLM-5.1 — next-generation flagship for agentic engineering, class-leading on SWE-Bench Pro; 744B total / 40B active, 8 of 256 experts plus 1 shared across 75 MoE layers; 198K context. `glm_moe_dsa` architecture with Dynamic Sparse Attention.

MIT372.5 GB198K ctx
GLM-5.2 (MoE, MTP)
glm · 753B (MoE, 40B active)

Z.ai GLM-5.2 — 753B total parameter MoE flagship, 40B active, routing each token to 8 of 256 experts plus 1 shared across 75 MoE layers. Solid 1M-token context with IndexShare sparse-attention (2.9× per-token FLOP reduction at 1M). Improved Multi-Token-Prediction layer increases speculative-decoding accept rate by ~20% over GLM-5.1.

MIT381.8 GB1024K ctx
GLM-4 9B Chat
glm · 9B

Zhipu AI GLM-4 9B instruction-tuned, 128K context

Apache 2.05.4 GB128K ctx
GPT-OSS 120B
gpt-oss · 120B

OpenAI GPT-OSS 120B — open-weights release, native MXFP4

Apache 2.068.5 GB128K ctx
GPT-OSS 20B
gpt-oss · 20B

OpenAI GPT-OSS 20B — open-weights release, native MXFP4

Apache 2.011.5 GB128K ctx
Granite 4.0 1B
granite · 1B

IBM Granite 4.0 1B — compact enterprise model

Apache 2.0972.7 MB128K ctx
Granite 4.0 350M
granite · 350M

IBM Granite 4.0 350M — ultra-compact for edge deployment

Apache 2.0209.8 MB128K ctx
Granite 4.0 H-Small (32B)
granite · 32B (hybrid)

IBM Granite 4.0 H-Small — 32B hybrid for long-context enterprise

Apache 2.018.1 GB128K ctx
Granite 4.0 H-Tiny
granite · 7B (hybrid)

IBM Granite 4.0 H-Tiny — hybrid Mamba/Transformer architecture

Apache 2.04.0 GB128K ctx
Inkling (MoE, multimodal)
inkling · 975B (MoE, 41B active)

Thinking Machines Inkling — 975B-total multimodal MoE, 41B active, routing each token to 6 of 256 experts plus 2 shared across 66 layers. Hybrid local/global attention, 1M context. Accepts text, images and 16kHz WAV audio via a hierarchical patch encoder and discrete audio tokens, all projected into one hidden space; output is text. Apache-2.0.

Apache 2.0546.7 GB1024K ctx
Inkling Small (MoE, multimodal)
inkling · 276B (MoE, 12B active)

Thinking Machines Inkling Small — 276B-total multimodal MoE, 12B active, 1M context. Accepts text, images and 16kHz WAV audio; output is text. Apache-2.0.

Apache 2.0158.3 GB1024K ctx
Kimi K2 Instruct (MoE)
kimi · 1T (MoE, 32B active)

Moonshot AI Kimi K2 MoE — 1T total, 32B active, 128K context

MIT18.8 GB128K ctx
Kimi K2.5 (MoE)
kimi · 1T (MoE, 32B active)

Moonshot AI Kimi K2.5 — 1T total / 32B active MoE; image input support; 256K context. Predecessor to K2.6's hybrid-thinking variant.

MIT540.2 GB256K ctx
Kimi K2.6 (Hybrid Thinking, MoE)
kimi · 1T (MoE, 32B active)

Moonshot AI Kimi K2.6 hybrid-thinking MoE — 1T total params, 256K context. Replica-routed on B200-class infrastructure; Unsloth measures >40 t/s on B200. Recommended `UD-Q2_K_XL` (350GB) for size/quality balance.

MIT558.8 GB256K ctx
Kimi K2.7 Code (MoE)
kimi · 1T (MoE, 32B active, code-focused)

Moonshot AI Kimi K2.7 Code — code-focused refresh of the K2 series. 1T total / 32B active; 256K context; recent updates target tool-call accuracy on long-horizon coding tasks.

MIT540.2 GB256K ctx
Kimi K3 (MoE, multimodal)
kimi-k3 · 2.8T total / 104B active (MoE)

Moonshot AI Kimi K3 — 2.8T total parameters, 104B active, 896 routed experts with 16 selected per token and 2 shared. Kimi Delta Attention plus gated MLA across 93 layers, 1M context, 160K vocabulary, MXFP4 weights and MXFP8 activations from quantization-aware training. Text, image, and video via the MoonViT-V2 encoder. `UD-IQ1_S` (594GB) is the smallest quant; `UD-Q2_K_XL` (861GB) is the size/quality balance point. Both exceed any single machine, so whole-model serving means a pipeline cluster; a lone host runs it as distributed expert extraction instead.

Kimi K3 License553.2 GB1024K ctx
Laguna S 2.1 (MoE)
laguna · 118B (MoE, 8B active)

Poolside Laguna S 2.1 — 118B-total agentic-coding MoE, 8B active, with a token-choice router using softplus gating over 256 routed experts plus 1 shared. Grouped-query attention with interleaved full and sliding-window layers, 1M context. Native interleaved thinking between tool calls; preserve reasoning blocks across turns. Requires llama.cpp b10087 or newer.

OpenMDW-1.168.4 GB1024K ctx
MiniMax M2.7 (MoE)
minimax · 230B (MoE, 10B active)

MiniMax M2.7 — current frontier MiniMax MoE; Lightning Attention with 1M context. Unsloth dynamic UD-Q4_K_XL GGUF (sharded).

MiniMax Open130.4 GB1024K ctx
MiniMax M3 (MoE, native multimodal)
minimax · 428B (MoE, 23B active)

MiniMax M3 — ~428B total / ~23B active MoE with native multimodal training. MiniMax Sparse Attention (MSA) delivers 9× prefill and 15× decode speedups vs M2 at 1M context. Note: GGUF builds currently fall back to dense attention; sparse attention not yet supported in llama.cpp.

MIT214.2 GB1024K ctx
Ministral 3 14B
mistral · 14B

High-performance Ministral 3 for complex reasoning

Apache 2.07.7 GB128K ctx
Ministral 3 3B
mistral · 3B

Compact Ministral 3 for lightweight tasks

Apache 2.02.0 GB128K ctx
Ministral 3 8B
mistral · 8B

Versatile Ministral 3 for general-purpose tasks

Apache 2.04.6 GB128K ctx
Mistral 7B Instruct v0.3
mistral · 7B

Mistral AI's classic 7B instruction model

Apache 2.04.1 GB32K ctx
Mistral Nemo 12B
mistral · 12B

Extended-context Mistral model built with NVIDIA

Apache 2.07.0 GB128K ctx
Mistral Small 3.2 24B
mistral · 24B

Mistral's latest Small 3.2 model for demanding workloads

Apache 2.013.3 GB128K ctx
Mistral Small 3.1 24B
mistral · 24B

Mistral Small 3.1 — improved reasoning over 3.0 baseline

Apache 2.013.3 GB128K ctx
Mistral Small 3.1 DRAFT 0.5B
mistral · 0.5B

Speculative drafter for Mistral Small 3.1/3.2 — vocab-matched, 6-language fine-tune.

Apache 2.0378.6 MB32K ctx
Mistral Small 3.2 24B
mistral · 24B

Mistral Small 3.2 — latest 3-series point release

Apache 2.013.3 GB128K ctx
Muse Glimmer 30B
muse-glimmer · 30B

Multimodal (text + image) 30B dense causal transformer with a frozen ViT-G/14 perception encoder (mmproj-kquant.gguf), 128K context. K-Quant build sized to fit 24GB VRAM; paired with its DFlash speculative-decoding drafter (dflash-kquant.gguf) for ~2x decode.

Apache 2.015.6 GB128K ctx
Muse Glimmer 30B DFlash Drafter
muse-glimmer · DFlash drafter

DFlash block-diffusion speculative drafter for Muse-Glimmer-30B. Pair via `--spec-type draft-dflash` for ~2x decode.

Apache 2.01.9 GB128K ctx
Nemotron 3 Nano Omni 30B-A3B (MoE, multimodal)
nemotron · 30B (MoE, 3B active)

NVIDIA Nemotron 3 Nano Omni — 30B total / 3B active hybrid reasoning MoE, 256K context. Accepts audio, video, text, images and documents; output is text. Needs a llama.cpp-compatible backend: the vision path uses a separate mmproj file that Ollama does not load.

NVIDIA Open17.2 GB256K ctx
Nemotron 3 Ultra 550B-A55B (MoE)
nemotron · 550B (MoE, 55B active)

NVIDIA Nemotron 3 Ultra — 550B total / 55B active hybrid Transformer-Mamba MoE with Latent MoE and a built-in Multi-Token-Prediction head; up to 1M context.

NVIDIA Open Model, Weights & Data307.3 GB1024K ctx
Nemotron 3 Nano 30B-A3B (MoE)
nemotron · 30B (MoE, 3B active)

Hybrid Mamba-2 MoE — 30B total, 3.5B active, 128K context

NVIDIA Open22.9 GB125K ctx
Nemotron 3 Nano 4B
nemotron · 4B

Hybrid Mamba-2 + Attention edge model, 256K context

NVIDIA Open2.7 GB256K ctx
NVIDIA Nemotron 3.5 Lightning 30B-A3B (MoE)
nemotron-lightning · 30B (MoE, 3B active)

NVIDIA Nemotron 3.5 Lightning — 30B total / 3B active hybrid Mamba-2 + Attention MoE. Thinking + instruct, 256K native context extensible to ~1M. Single-file GGUF, single-node serving.

OpenMDW-1.123.7 GB256K ctx
Ornith 1.0 35B (MoE)
ornith · 35B (MoE, 3B active)

Deep Reinforce Ornith 1.0 35B — agentic-coding MoE with 256 routed experts, 8 selected per token and 1 shared, over 40 layers. 256K context, vision input via the bundled projector. Designed for single-GPU deployment.

MIT20.8 GB256K ctx
Ornith 1.0 397B (MoE)
ornith · 397B (MoE, 17B active)

Deep Reinforce Ornith 1.0 397B — flagship agentic-coding MoE with 512 routed experts, 10 selected per token and 1 shared, over 60 layers. 256K context. Leads open-weight coding benchmarks at its size; MIT licensed with no regional limitation.

MIT228.9 GB256K ctx
Ornith 1.0 9B
ornith · 9B

Deep Reinforce Ornith 1.0 9B — dense agentic-coding model post-trained on Qwen 3.5, 256K context, vision input via the bundled projector. Smallest member of the Ornith family.

MIT5.6 GB256K ctx
Phi-4 14B
phi · 14B

Microsoft Phi-4 — strong reasoning at 14B

MIT8.3 GB16K ctx
Phi-4 Mini 3.8B
phi · 3.8B

Compact Phi-4 Mini with 128K context

MIT2.3 GB125K ctx
Phi-4 Mini Reasoning 3.8B
phi · 3.8B

Compact Phi-4 Mini fine-tuned for reasoning tasks

MIT2.3 GB125K ctx
Phi-4 Reasoning 14B
phi · 14B

Phi-4 fine-tuned for chain-of-thought reasoning

MIT8.4 GB32K ctx
Qwen-AgentWorld 35B-A3B (MoE)
qwen-agentworld · 35B (MoE, 3B active)

Qwen-AgentWorld 35B-A3B — agentic MoE, 3B active per token

Apache 2.020.8 GB256K ctx
Qwen 3 0.6B
qwen3 · 0.6B

Compact model optimized for edge deployment

Apache 2.0378.3 MB32K ctx
Qwen 3 1.7B
qwen3 · 1.7B

Versatile model for various language tasks

Apache 2.01.0 GB32K ctx
Qwen 3 14B
qwen3 · 14B

Premium model with extended context support

Apache 2.08.4 GB128K ctx
Qwen 3 30B-A3B (MoE)
qwen3 · 30B (MoE, 3B active)

Mixture-of-Experts with 3B active params for efficient scaling

Apache 2.017.3 GB128K ctx
Qwen 3 32B
qwen3 · 32B

Top-tier model with 128K context window

Apache 2.018.4 GB128K ctx
Qwen 3 4B
qwen3 · 4B

Well-balanced model for production use

Apache 2.02.3 GB32K ctx
Qwen 3 8B
qwen3 · 8B

Extended context model for long-form tasks

Apache 2.04.7 GB128K ctx
Qwen 3 Coder 30B-A3B (MoE)
qwen3 · 30B (MoE, 3B active)

Code-focused MoE — 30B total, 3B active, 256K context

Apache 2.017.3 GB256K ctx
Qwen3-Coder-Next
qwen3-next · 80B (MoE, 3B active)

Qwen3-Coder-Next — agentic coding, long-context repository work

Apache 2.046.2 GB256K ctx
Qwen3-VL 8B Instruct
qwen3-vl · 8B

Qwen3-VL 8B — vision-language instruct, images on the chat surface

Apache 2.04.8 GB256K ctx
Qwen 3.5 0.8B
qwen3.5 · 0.8B

Compact multilingual model for efficient on-device inference

Apache 2.0507.8 MB128K ctx
Qwen 3.5 0.8B (MTP)
qwen3.5 · 0.8B

Qwen 3.5 0.8B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0515.0 MB128K ctx
Qwen 3.5 122B-A10B (MoE)
qwen3.5 · 122B (MoE, 10B active)

Qwen 3.5 large MoE — 122B total, 10B active per token. Replica-routed on high-VRAM provider tiers only; Unsloth ships an MTP variant in `unsloth/Qwen3.5-122B-A10B-MTP-GGUF` for compatible runtimes.

Apache 2.069.8 GB256K ctx
Qwen 3.5 122B-A10B (MoE) (MTP)
qwen3.5 · 122B (MoE, 10B active)

Qwen 3.5 122B-A10B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.069.8 GB128K ctx
Qwen 3.5 27B
qwen3.5 · 27B

Flagship Qwen 3.5 model

Apache 2.015.6 GB128K ctx
Qwen 3.5 27B (MTP)
qwen3.5 · 27B

Qwen 3.5 27B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.015.8 GB128K ctx
Qwen 3.5 2B
qwen3.5 · 2B

Efficient small model for chat and text generation

Apache 2.01.2 GB128K ctx
Qwen 3.5 2B (MTP)
qwen3.5 · 2B

Qwen 3.5 2B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.01.2 GB128K ctx
Qwen 3.5 35B-A3B (MoE)
qwen3.5 · 35B (MoE, 3B active)

Mixture-of-Experts with only 3B active params — fast inference at 35B quality

Apache 2.020.5 GB256K ctx
Qwen 3.5 35B-A3B (MoE) (MTP)
qwen3.5 · 35B (MoE, 3B active)

Qwen 3.5 35B-A3B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.021.0 GB128K ctx
Qwen 3.5 397B-A17B (MoE)
qwen3.5 · 397B (MoE, 17B active)

Qwen 3.5 frontier MoE — 397B total, 17B active per token. Multi-GPU replicas only; Unsloth ships an MTP variant in `unsloth/Qwen3.5-397B-A17B-MTP-GGUF` for compatible runtimes.

Apache 2.0223.5 GB256K ctx
Qwen 3.5 397B-A17B (MoE) (MTP)
qwen3.5 · 397B (MoE, 17B active)

Qwen 3.5 397B-A17B (MoE) with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.0223.5 GB128K ctx
Qwen 3.5 4B
qwen3.5 · 4B

Mid-size model with strong reasoning and coding performance

Apache 2.02.6 GB128K ctx
Qwen 3.5 4B (MTP)
qwen3.5 · 4B

Qwen 3.5 4B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.02.3 GB128K ctx
Qwen 3.5 9B
qwen3.5 · 9B

High-performance model for complex language understanding

Apache 2.05.3 GB128K ctx
Qwen 3.5 9B (MTP)
qwen3.5 · 9B

Qwen 3.5 9B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth measures ~1.5-2× speedup over the non-MTP baseline.

Apache 2.05.1 GB128K ctx
Qwen 3.6 27B
qwen3.6 · 27B

Qwen 3.6 27B — flagship dense model with 128K context

Apache 2.015.6 GB128K ctx
Qwen 3.6 27B Fable-Fusion (MTP, uncensored)
qwen3.6 · 27B

Qwen 3.6 27B Fable-Fusion — community fine-tune of Qwen3.6-27B tuned for long-form fiction and creative writing, with multi-token-prediction tensors carried in-file at Q8_0 so it drafts against itself. 256K native context, YaRN-extensible. Uncensored: the base model's refusal behaviour is deliberately removed.

Apache 2.015.6 GB256K ctx
Qwen 3.6 27B (MTP)
qwen3.6 · 27B

Qwen 3.6 27B with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth: 160 t/s on RTX 6000.

Apache 2.016.7 GB128K ctx
Qwen 3.6 35B-A3B (MoE)
qwen3.6 · 35B (MoE, 3B active)

Qwen 3.6 MoE — 35B total, ~3B active per token

Apache 2.019.9 GB256K ctx
Qwen 3.6 35B-A3B MTP (MoE)
qwen3.6 · 35B (MoE, 3B active)

Qwen 3.6 35B-A3B MoE with built-in Multi-Token-Prediction head. Single-file MTP GGUF — no separate drafter needed. Unsloth: 240 t/s on RTX 6000.

Apache 2.020.5 GB256K ctx
Qwen 3.8 2.4T-A95B (MoE)
qwen3.8 · 2.4T (MoE, 95B active)

Qwen 3.8 flagship MoE — 2.4T total, 95B active per token, thinking-only. 256K native context extensible to ~1M. Multi-node MoE-sharded serving only (10-shard UD-Q1_0 GGUF; first shard is the load entry). Built-in single MTP layer for self-speculative decoding.

qwen3.8-max369.7 GB256K ctx
Qwen 3.8 27B
qwen3.8 · 27B (dense)

Qwen 3.8 27B dense, multimodal, MTP-trained. Strong agentic/competitive coding (SWE-Bench Pro, LiveCodeBench). Hybrid Gated-DeltaNet + attention, 262K context.

Apache 2.016.8 GB256K ctx
Qwen 3.8 27B DFlash-2 Drafter
qwen3.8 · DFlash-2 drafter

DFlash-2 block-diffusion speculative drafter for Qwen3.8-27B. Drafts a whole block per pass and keeps the top candidates at every position. Pair via `--spec-type draft-dflash`.

Apache 2.01.1 GB256K ctx
Qwen 3.8 27B (3-bit)
qwen3.8 · 27B (dense)

Qwen 3.8 27B dense at 3-bit, sized to fit a 16 GB consumer GPU. Measured 62.0 tok/s on an RTX 5070 Ti against 19.6 tok/s for the 4-bit build on a GB10 — decode is bandwidth-bound, and this build reads 13.44 GB per token.

Apache 2.012.5 GB256K ctx
Qwen3.8-Flash-Next (MoE)
qwen3.8-flash-next · 125B (MoE, 6B active)

Qwen3.8-Flash-Next — agentic coding and long-horizon tool use. MTP-enabled — pairs with `qwen3.8-flash-next-mtp-draft` for lossless speculative decode.

Apache 2.073.5 GB256K ctx
Qwen3.8-Flash-Next MTP Drafter
qwen3.8-flash-next · MTP head

Qwen's jointly-trained Multi-Token Prediction head for Qwen3.8-Flash-Next, exported as a standalone GGUF sidecar (Unsloth's main build strips it). Pair with the Flash-Next target via `--spec-type draft-mtp`.

Apache 2.03.9 GB256K ctx
SmolLM2 1.7B
smollm · 1.7B

Compact SmolLM2 for on-device AI

Apache-2.01006.7 MB8K ctx
SmolLM3 3B
smollm · 3B

SmolLM3 — 11T tokens, dual-mode reasoning

Apache-2.01.8 GB64K ctx

Speech recognition

8 models

Detection

9 models

Text embedding

7 models

Forecasting

4 models

Media generation

31 models
NVIDIA Cosmos3 Edge
cosmos3 · 4B

NVIDIA Cosmos3 Edge world foundation model, 4B, image-to-video at 480p/121 frames. Ungated, commercial use permitted. Takes JSON-structured prompts rather than plain text

OpenMDW-1.116.0 GB
FLUX.2 dev
Gated
flux2 · 32B

FLUX.2 dev — 32B flagship text-to-image and editing

FLUX.2 Non-Commercial / BFL custom165.4 GB
FLUX.2 klein 4B
flux2 · 3.9B

FLUX.2 klein 4B text-to-image and editing, runs on consumer GPUs

Apache-2.022.1 GB
FLUX.2 klein 9B
Gated
flux2 · 9B

FLUX.2 klein 9B — distilled few-step generation and editing

FLUX.2 Non-Commercial / BFL custom49.3 GB
FLUX.2 klein 9B (GGUF Q4_K_M)
Gated
flux2 · 9B (Q4_K_M)

FLUX.2 klein 9B, Q4_K_M — runs beside a video model on one box

FLUX.2 Non-Commercial / BFL custom5.5 GB
FLUX.2 klein 9B (GGUF Q8_0)
Gated
flux2 · 9B (Q8_0)

FLUX.2 klein 9B, Q8_0 — near-lossless at under half the bf16 footprint

FLUX.2 Non-Commercial / BFL custom9.3 GB
FLUX.2 klein 9B KV
Gated
flux2 · 9B

FLUX.2 klein 9B KV — few-step generation and editing

FLUX.2 Non-Commercial / BFL custom49.3 GB
FLUX.2 klein base 4B
flux2 · 3.9B

FLUX.2 klein base 4B — undistilled text-to-image and editing

Apache-2.022.1 GB
FLUX.2 klein base 9B
Gated
flux2 · 9B

FLUX.2 klein base 9B — undistilled generation and editing

FLUX.2 Non-Commercial / BFL custom49.3 GB
Hunyuan3D 2.1
hunyuan3d · 3.3B

Tencent Hunyuan3D 2.1 image-to-3D with PBR texture synthesis. Ungated on the Hub but non-commercial: enrollment refuses it until the operator sets --accept-non-commercial. Loads through `hy3dgen`

TENCENT HUNYUAN NON-COMMERCIAL12.5 GB
LTX-2.3 22B Distilled (GGUF Q5_K_M)
ltx2 · 22B

LTX-2.3 22B distilled text-to-video, 8-step schedule, synchronized audio branch

LTX Open Weights28.1 GB
LTX-2.5 22B Dev (bf16)
Gated
ltx2 · 22B

LTX-2.5 22B dev text-to-video, the full trainable transformer with a synchronized audio branch. Unverified: catalogued from the model card and the pinned diffusers signature, never run here.

LTX-2.x Community License187.1 GB
LTX-2.5 22B Dev (comfy-int8)
Gated
ltx2 · 22B

LTX-2.5 22B dev text-to-video, comfy-int8 transformer with the matching int8 text encoder. Same weights repo as the bf16 entry; the smaller serving set is what lets it share a node with the LLM tier. Unverified: catalogued from the files on disk, never rendered here.

LTX-2.x Community License187.1 GB
LTX-2.5 22B Distilled (bf16)
Gated
ltx2 · 22B

LTX-2.5 22B distilled text-to-video with a synchronized audio branch. Unverified: catalogued from the model card, never run here.

LTX-2.x Community License187.1 GB
LTX-2.5 22B Distilled (nvfp4)
Gated
ltx2 · 22B

LTX-2.5 22B distilled text-to-video, nvfp4 transformer. Lightricks ships nvfp4 for the distilled transformer only, so there is no dev equivalent. Fixed 8-step, CFG 1.0 schedule inherited from the distilled checkpoint; raising either fights the distillation. Unverified: catalogued from the files on disk, never rendered here.

LTX-2.x Community License187.1 GB
MiniMax H3 (Hailuo 3.0) H3-Base
minimax-h3 · H3-Base + Qwen3-VL-32B text encoder

MiniMax H3 omni-modal text-to-video with native stereo audio, 768p short edge, 4-15s. Licence excludes the EU, UK, South Korea and the USA, including outputs

MiniMax H3 Community License134.1 GB
MiniMax H3 FL2VA (GGUF Q4_K)
minimax-h3 · H3 FL2VA pruned Q4_K + Qwen3-VL-32B Q4_K_M

MiniMax H3 first-and-last-frame to video, Q4_K. Licence excludes the EU, UK, South Korea and the USA, including outputs

MiniMax H3 Community License33.0 GB
MiniMax H3 (Hailuo 3.0) H3-Base (int8)
minimax-h3 · H3-Base + Qwen3-VL-32B text encoder, int8

MiniMax H3 text-to-video, FP8 transformer, 768p short edge, 4-15s. Licence excludes the EU, UK, South Korea and the USA, including outputs

MiniMax H3 Community License134.1 GB
MiniMax H3 Ref2VA (FP8)
minimax-h3 · H3 Ref2VA FP8 + Qwen3-VL-32B INT8 text encoder

MiniMax H3 reference-image-to-video, FP8. Licence excludes the EU, UK, South Korea and the USA, including outputs

MiniMax H3 Community License49.4 GB
MiniMax H3 Ref2VA (GGUF Q4_K)
minimax-h3 · H3 Ref2VA pruned Q4_K + Qwen3-VL-32B Q4_K_M

MiniMax H3 reference-image-to-video, Q4_K. Licence excludes the EU, UK, South Korea and the USA, including outputs

MiniMax H3 Community License33.0 GB
MiniMax-Music3 2B (text-to-music)
minimax-music · 2B

MiniMax-Music3 2B text-to-music: 32 kHz stereo, up to five minutes, prompt plus optional lyrics. Unverified: catalogued from the model card, never run here.

Creative Commons (see repo LICENSE)53.5 GB
Qwen-Image
qwen-image · 20.4B

Qwen-Image 2512 text-to-image MMDiT with strong text rendering

Apache-2.053.7 GB
Qwen-Image-Edit 2511
qwen-image · 20.4B

Qwen-Image-Edit instruction-driven image editing, multi-reference

Apache-2.053.8 GB
Qwen-Image-Flash
qwen-image · 20.4B

Qwen-Image distilled to four steps, guidance disabled

NVIDIA Open Model License53.7 GB
Qwen-Image 2512 (GGUF Q5_K_M)
qwen-image · 20.4B

Qwen-Image 2512 text-to-image, Q5_K_M GGUF transformer

Apache-2.014.0 GB
TRELLIS.2 4B
trellis2 · 4B

Microsoft TRELLIS.2 image-to-3D. Produces a GLB mesh with PBR materials including transparency, at up to 1536³. Loads through the `trellis2` package, not diffusers

MIT15.6 GB
Wan 2.1 FLF2V 14B 720P
wan2.1 · 14B

Wan 2.1 first-last-frame interpolation; bridges two stills into motion

Apache-2.083.9 GB
Wan 2.2 I2V A14B
wan2.2 · 14.3B

Wan 2.2 image-to-video mixture-of-experts, 480P and 720P

Apache-2.0117.5 GB
Wan 2.2 T2V A14B
wan2.2 · 14.3B

Wan 2.2 text-to-video mixture-of-experts, 480P and 720P

Apache-2.0117.5 GB
Wan 2.2 TI2V 5B
wan2.2 · 5.0B

Wan 2.2 hybrid text/image-to-video at 720P 24fps on a single 24 GB GPU

Apache-2.031.9 GB
Z-Image Turbo
z-image · 6.2B

Z-Image Turbo few-step text-to-image, 9 steps without guidance

Apache-2.030.6 GB

Segmentation

4 models

Text segmentation

1 model

Speech synthesis

3 models

Vision

9 models
How this list is built

The registry is read from the network itself — one catalog call per modality — and every entry is then checked against its source repository on Hugging Face: that the repository resolves, whether the weights are gated behind terms, and what license the source states. Where the registry and the source disagree on a license, both are shown on the model page rather than one being silently preferred.