OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device
MarkTechPosten
MarkTechPost
AI Global WireOpenBMB has released MiniCPM5-2B, a dense causal language model with 2,516,756,480 parameters and a native 131,072 token context. It averages 53.9 across the 34 benchmarks in its model card, ahead of Qwen3.5-4B at 51.1, with its clearest leads in tool use, coding agents and long-context retrieval. Post-training pairs 400B tokens of deep-thinking SFT with RL teachers and on-policy distillation that merges 16 expert models into one checkpoint. The weights ship under Apache 2.0 alongside the pre-training, SFT and RL datasets and the intermediate Base, Midtrain and SFT-only checkpoints. GGUF builds start at 1.56 GB, and the standard LlamaForCausalLM architecture loads in vLLM, SGLang, llama.cpp,
This is a short summary published by AI Global Wire. The full article is owned and hosted by MarkTechPost — open it there to read it in full.
Read the full story at MarkTechPost- Meta
- Verktyg
- Agenter
- Reglering
Related AI news
- Sources: PE firm Silver Lake intends to merge French software companies Cegid and Silae in a €10B+ deal, allowing them to combine data and integrate software (Financial Times)Techmeme · September 9, 2026
- Cloudera brings Mistral AI’s frontier models into its secure hybrid data environmentsSiliconANGLE · September 9, 2026
- Muse Spark 1.3: Meta startet KI-Agenten f�r Mails, Eink�ufe und ReisenGolem.de · September 9, 2026
- Hugging Face's new ML Intern lets anyone run machine learning experiments through a simple chatThe Decoder · September 9, 2026
- Alibaba sends AI ‘digital employees’ to work in rival apps from ByteDance, TencentSCMP Tech · September 9, 2026
- OpenAI's millennium proof dispute raises the question of whether researchers can trust AI labsThe Decoder · September 9, 2026