#1Best match
Official MIT DeepSeek V4 Flash multimodal experiment with image understanding, 1M context and Unsloth Dynamic GGUF artifacts. The lightest practical GGUF is roughly 82-97GB, while higher-quality Q4/Q8 builds are about 155-162GB, so this belongs on large-memory workstations.
Parameters284B (13B active, multimodal MoE)Minimum RAM128GBQuantizationUD-Q2_K_XLModel size97GB
View model details →#2Best match
Official MIT DeepSeek V4 Flash successor release with stronger agentic coding, DSpark speculative decoding support and a practical Unsloth Dynamic GGUF path. Still a large workstation/server local model: Q4 is about 155GB and Q8 is about 162GB.
Parameters284B (13B active)Minimum RAM256GBQuantizationUD-Q4_K_XLModel size155GB
View model details →#3Best match
Official MIT DeepSeek V4.1 Flash release with CED architecture, CSA2 attention, multimodal input and 1M-token context. Community GGUF artifacts include split Q2_K/Q3/Q4 builds plus a llama.cpp patch path; use only on 256GB+ workstations, with 384GB+ safer.
Parameters552B MoE (8B/16B active)Minimum RAM256GBQuantizationQ2_KModel size246GB
View model details →#4Best match
Official China Telecom XingChen-AGI Apache 2.0 MoE release with 29B total parameters, 4B active parameters, 256K native context and an official IQ4_NL GGUF path for llama.cpp-class local inference on 24GB GPU workstations.
Parameters29B (4B active, MoE)Minimum RAM32GBQuantizationIQ4_NL GGUFModel size20.1GB
View model details →#5Best match
Accio Lab Apache 2.0 co-work model post-trained from Qwen3.6-35B-A3B for long-horizon tools, files, code and business workflows. Official GGUF Q4_K_M is 19.7GiB with llama.cpp validation evidence.
Parameters35B (3B active, MoE)Minimum RAM32GBQuantizationQ4_K_MModel size19.7GB
View model details →#6Best match
Official Nex-AGI Apache 2.0 multimodal agent model for computer use, web browsing, coding and tool calling. Community Q4_K_M GGUF is about 21.3GB with documented llama.cpp text and vision smoke tests.
Parameters35B MoEMinimum RAM32GBQuantizationQ4_K_MModel size21.3GB
View model details →#7Best match
Official Qwen sparse multimodal MoE preview with 125B model parameters plus 51B n-gram embeddings, about 6B active parameters, Qwen Community 1.0 licensing, 262K native context and local Q4_K_M GGUF paths for llama.cpp, Ollama and LM Studio-class runtimes.
Parameters125B + 51B n-gram (6B active)Minimum RAM96GBQuantizationQ4_K_MModel size54.5GB
View model details →#8Best match
Official MIT 35B MoE reasoning model from Ornith AI with about 3B active parameters, 262K context, strong agentic-coding positioning and official Q4_K_M GGUF plus MLX/Ollama/llama.cpp local paths for larger workstations.
Parameters35B (3B active, MoE)Minimum RAM48GBQuantizationQ4_K_MModel size21.72GB
View model details →#9Best match
Meta Superintelligence Lab local agent model with text+image input, 131K context, Apache 2.0 weights and official GGUF/ExecuTorch artifacts. The K-Quant 17GB build targets 24GB machines; 32GB is safer for vision and long-context sessions.
Parameters29.8B multimodalMinimum RAM24GBQuantizationK-Quant 17GB Q4_K_MModel size17GB
View model details →