
HF Breakout Models, Jun 1–8: Nemotron Ultra, Ideogram 4, and Six Models Worth Evaluating This Week
Six open-weight HuggingFace models with explosive June 1–8 download growth across four modalities: NVIDIA Nemotron-3-Ultra (550B/55B active, OpenMDW-1.1, commercial-ready frontier LLM with 71.9% SWE-Bench Verified), JetBrains Mellum2 (12B/2.5B active, Apache 2.0, low-latency coding MoE at 17.4K downloads), Ideogram 4 (9.3B T2I, non-commercial, #1 open-weight on Design Arena), PaddleOCR-VL-1.6 (1B, Apache 2.0, 96.33% document parsing SOTA), Higgs Audio v3 TTS (4B, non-commercial, 100+ language zero-shot voice cloning), and Nemotron 3.5 ASR (600M, OpenMDW-1.1, 40-language streaming ASR at 17× throughput of Parakeet 1.1B). MiniMax M3 and Qwen3.7-Plus flagged as API-only / no weights.
LLMs
Nemotron-3-Ultra — 550B/55B active, OpenMDW-1.1, commercial OK

- License: OpenMDW-1.1 (Linux Foundation) — commercial and non-commercial use both explicitly permitted 1
- Active params: 55B (550B total MoE)
- Context: 1M tokens, 10 languages
- Deployment: vLLM, SGLang, TRT-LLM; 50+ cloud providers Day-0
- Builder angle: NVIDIA positions this as an orchestration model for "hard calls" in long-running agent workflows — maintaining architectural decisions across coding sessions, reconciling conflicting research sources, verifying constraints at scale. The throughput advantage is most relevant at high concurrency in cloud settings; self-hosting requires serious multi-GPU hardware. If you're on H100s already and need the best US-origin open-weight reasoning model for agentic pipelines, this is currently the leading option.
Mellum2 — 12B/2.5B active, Apache 2.0, low-latency coding
<think>...</think> tags before the final answer. 3vllm serve JetBrains/Mellum2-12B-A2.5B-Thinking --max-model-len 131072 --reasoning-parser qwen3; add --enable-auto-tool-choice --tool-call-parser hermes for tool calls. 3- License: Apache 2.0 — commercial use fully permitted
- Active params: 2.5B (12B total MoE), 131k context
- Deployment: vLLM (native), plus 26 community GGUF variants
- Builder angle: at 2.5B active parameters with 131k context and Apache 2.0 licensing, Mellum2 fits the routing and sub-agent slot in an IDE or RAG pipeline where you need reasoning-capable output but can't afford the latency or GPU cost of a 10B+ dense model. The Thinking/Instruct split means you can route low-stakes queries to Instruct and keep Thinking for the complex debug sessions.
Image generation
Ideogram 4 — 9.3B, non-commercial license
- License: Ideogram 4 Non-Commercial License — research and personal use only; commercial deployment requires a separate agreement from Ideogram AI
- Params: 9.3B, single-stream DiT
- Min hardware: single 24 GB GPU (nf4 variant)
- Builder angle: the non-commercial license blocks production deployment without a deal with Ideogram, but the weights are public for prototyping, fine-tuning research, and evaluating fit for your use case. The JSON-structured prompting interface (bounding boxes + color palettes) is specifically useful for design tools and templated asset generation, where most open T2I models require complex prompt engineering to hit layout precision. If you need a commercial path today, this isn't it yet.
Document parsing
PaddleOCR-VL-1.6 — 1B, Apache 2.0, document SOTA
- License: Apache 2.0 — commercial use fully permitted
- Params: 1.0B
- Builder angle: the combination of 1B scale, Apache 2.0 license, drop-in compatibility with 1.5, and SOTA benchmark performance makes this the go-to upgrade for any document intelligence pipeline currently running PaddleOCR. For builders building document extraction from scratch, the small footprint means it runs comfortably on a single mid-range GPU or even CPU inference for batch workflows.
Audio
Higgs Audio v3 TTS — 4B, non-commercial, 100+ languages
<|category:value|> format — 21 emotions (including elation, anger, sadness, and amusement), 3 speaking styles (singing, shouting, whispering), 9 sound effects, and explicit pause/pitch/speed control. The model card describes the design intent as: "Higgs Audio v3 TTS is built for voice chat: it speaks, not just reads." 7 Throughput: 14.74 req/s at 16 concurrent requests on a single H100, with RTF 0.262 (Real-Time Factor — below 1.0 means faster than real-time).- License: Boson Higgs Audio v3 Research and Non-Commercial License — no commercial use
- Params: ~4B
- Languages: 100+ (85 at production WER/CER)
- Builder angle: the non-commercial license limits production deployment, but for prototyping multilingual voice interfaces or testing zero-shot voice cloning quality before committing to a commercial TTS provider, this is currently the strongest open-weight option. The inline emotion tokens are particularly useful for voice assistant research — you can vary emotional register without post-processing or separate model calls.
Nemotron 3.5 ASR — 600M, OpenMDW-1.1, commercial OK
target_lang=auto for automatic detection. 8- License: OpenMDW-1.1 — commercial and non-commercial use both explicitly permitted
- Params: 600M
- Languages: 40 locales (19 transcription-ready)
- Builder angle: if you're running real-time transcription across multiple languages and currently deploying separate per-language models, Nemotron 3.5 ASR collapses that into one model at substantially lower per-stream GPU cost. The 240-concurrent-stream figure on a single H100 at 80ms latency is the number to stress-test against your traffic patterns. Commercial license means you can ship it.
On the radar
The week's shape
참고 출처
- 1
- 2NVIDIA Developer Blog: Nemotron 3 Ultra Powers Faster, More Efficient Reasoning
developer.nvidia.com
- 3JetBrains/Mellum2-12B-A2.5B-Thinking · Hugging Face
huggingface.co
- 4ideogram-ai/ideogram-4-fp8 · Hugging Face
huggingface.co
- 5
- 6PaddlePaddle/PaddleOCR-VL-1.6 · Hugging Face
huggingface.co
- 7bosonai/higgs-audio-v3-tts-4b · Hugging Face
huggingface.co
- 8nvidia/nemotron-3.5-asr-streaming-0.6b · Hugging Face
huggingface.co
- 9
- 10GitHub - MiniMax-AI/MiniMax-M3
github.com
- 11OpenRouter June 2026: New Models, Pricing and Rankings
digitalapplied.com

Hugging Face Surging Models
Weekly digest of HF models with > 10x download growth, with brief description, license, and business applicability
이 콘텐츠는 채널이 자동으로 생성했습니다. 한 문장이면 Neodrop이 당신을 위해 계속 만들어 냅니다.
관련 콘텐츠
- 로그인하면 댓글을 작성할 수 있습니다.