| Qwen3.8-27B | Permissive weights | Apache 2.0 | 27B language model plus vision components | General text, coding, agents; image and video input; 262K context | Named community UD-Q4_K_M: 16.46 GB + 0.93 GB vision; add runtime headroom | Candidate for local evaluation; community quantizations |
| Muse Glimmer-30B | Permissive weights | Apache 2.0 + separate Meta usage policy | ~29.6B dense + 1.8B vision encoder | Tool-calling agents, search, image input; 131K context | Official Q4: 16.76 GB + 1.40 GB vision; optional drafter adds 1.63 GB | First-party quantized artifacts and local runtime guidance |
| Qwen3.6-35B-A3B | Permissive weights | Apache 2.0 | 35B total / 3B active | General text, coding, agentic workflows, 262K context | Prosumer local — 32–64GB unified or 24GB+ GPU with sufficient runtime headroom | Production-leaning; the MoE fallback |
| Qwen3-Coder-Next | Permissive weights | Apache 2.0 | 80B total / 3B active | Coding agents, tool use, long-horizon reasoning, 262K context | Official Q4_K_M: 48.41 GB before cache and runtime; larger memory or offload | Dedicated coding candidate; non-thinking mode |
| Gemma 4 31B / 12B Unified | Permissive weights | Apache 2.0 | 30.7B dense or 11.95B unified | Text, coding, image; audio on 12B Unified | 12B on a 16 GB laptop (per Google); 31B should fit a 24–48 GB card at Q4 | Production-leaning; official QAT GGUFs |
| gpt-oss-120b / 20b | Permissive weights | Apache 2.0 + usage policy | 117B / 5.1B active · 21B / 3.6B active | Text reasoning, tool use, structured output | Single 80GB GPU · 16 GB device | Production-leaning |
| Nemotron 3.5 Lightning | Permissive weights | OpenMDW-1.1 | 30B total / 3B active | Fast local agents and sub-agents; 1M context nominal | Check the selected NVFP4/BF16 artifact and NVIDIA serving recipe | Hybrid architecture; evaluate speed and quality on target hardware |
| Qwen3.8-Flash-Next | Open weight, custom terms | Qwen Community License 1.0; business-use gate has no revenue floor | 125B backbone / 6B active + 51B n-gram + 4B MTP | Text, image, video; 262K native context | Roughly 180B checkpoint; high-memory deployment | Downloadable; verify license before commercial evaluation |
| GLM-5.3 | Open weight, custom terms | Custom GLM-5.3 License; conditional security review | Approximately 753B checkpoint | Text coding and long-horizon agents | Large-memory / multi-GPU; official FP8 and BF16 repositories | Weights available; previously listed here as pending |
| GLM-5.3-Flash | Permissive weights | MIT | 320B total / 18B active (publisher) | Native multimodal input; hybrid sparse and linear attention | Large-memory / multi-GPU; recipe-specific quantization | Separate architecture from GLM-5.3; released weights |
| DeepSeek-V4-Flash-0731 | Permissive weights | MIT | 284B total / 13B active | Agentic coding and tool use, 1M context, DSpark speculative decoding | Multi-GPU (~167 GB FP4/FP8); reference 4×GB300 node | Official release (Jul. 31, 2026), supersedes the April preview |
| DeepSeek-V4-Pro-0813 | Permissive weights | MIT | 1.6T total / 49B active | Frontier reasoning and agentic behavior, 1M context | Datacenter — 4×GB300 reference deployment | GA (Aug. 13, 2026), supersedes the April preview |
| Inkling | Permissive weights | Apache 2.0 (ungated download) | 975B total / 41B active | Natively multimodal (text, image, audio in) frontier-scale reasoning | Datacenter — 1.91 TB of weights; Inkling-Small (276B / 12B) at ~532 GB | New (July 15, 2026) — evaluate before trusting |
| Molmo2-8B | Permissive weights | Apache 2.0 (data-use caveats) | 8B | Image, video, grounding, pointing, tracking | Workstation / server | Research + production pilots (Dec. 2025; full training code Mar. 2026) |
| MiniCPM-o 4.5 | Permissive weights | Apache 2.0 | 9B | Real-time omnimodal — text, vision, speech | 19 GB BF16 / 11 GB INT4 (per OpenBMB); quantized local | Production-leaning for edge/omni |
| Whisper large-v3-turbo | Permissive weights | MIT (code and weights) | 0.8B | ASR, low-latency transcription | CPU, modest GPU (~6GB VRAM), or edge device | Mature / production standard |
| Qwen3-TTS 1.7B / 0.6B | Permissive weights | Apache 2.0 | 1.7B or 0.6B | Multilingual TTS, voice cloning, streaming | Consumer / prosumer GPU | Production-leaning |
| Wan2.2 TI2V-5B | Permissive weights | Apache 2.0 | 5B (family also has 27B / 14B-active MoE models) | Text- and image-to-video, 720p at 24 fps | 24 GB GPU with model offload (RTX 4090 per the README) | Advanced prototyping / R&D; open-weight Wan model |
| OLMo 3.1 32B-Think | Fully open | Apache 2.0 + published data and recipe | 32B dense | Auditable reasoning with full model-flow traceability | Prosumer local — 32–64GB | Mature for research and governed evaluation |
| OLMo Hybrid 7B | Fully open | Apache 2.0 + published data and recipe | 7B | Hybrid Gated DeltaNet architecture research; 65K context | Consumer GPU | Research (Mar. 5, 2026); no RL-final chat checkpoint yet |
| Llama 4 Scout | Open weight, custom terms | Llama 4 Community License + AUP | 109B total / 17B active | Multimodal text+image, 10M context | INT4 on a single H100 or a 128GB unified box | Production-leaning; Meta's newer open weights are Muse Glimmer |
| Mistral Medium 3.5 | Open weight, custom terms | Modified MIT — no rights above $20M monthly revenue | 128B dense | Reasoning, coding, multimodal, agents; 256K context | ~4-GPU self-hosting | Production-leaning |
| Kimi K3 | Open weight, custom terms | Custom Kimi K3 License; conditional MaaS agreement and UI attribution | 2.8T total / 104B active | Native vision, 1M context, MXFP4 weights | Datacenter — 1.56TB across 96 shards; 64+ accelerators per Moonshot | Weights July 27, 2026 — verify license fit first |
| Qwen3.8-2.4T-A95B | Open weight, custom terms | Qwen3.8-Max License — $50M trigger for MaaS / AI-work-assistant businesses + UI attribution | 2.4T total / 95B active | Text-only, thinking-only checkpoint of the hosted Qwen3.8-Max; 262K native context | Datacenter — 2.5 TB FP8 / 4.89 TB BF16; GB300 NVL72 reference | Weights Aug. 12, 2026 — the checkpoint is not the hosted product |
| Voxtral TTS | Restricted / non-commercial | CC-BY-NC 4.0 (inherited from reference voices) | 4B | High-quality multilingual speech synthesis | Consumer / prosumer GPU | Non-commercial use only |
| Muse Spark 1.2 (weights promised) | Announced, not released | None published | Not disclosed | Hosted model; not included as a downloadable recommendation | No checkpoint verified in this review | Re-check publisher artifacts before evaluation |