JetBrains Introduces the Advanced Mellum2.1 Model
On October 8, JetBrains published an insightful blog post. It details how Mellum2.1 gets to work as a fast, open model for coding agents. Continuing the legacy of the 12-billion-parameter Mixture-of-Experts architecture established by its predecessor, this iteration utilizes 2.5 billion active parameters. Furthermore, it remains freely accessible under the permissive Apache 2.0 license.
Enhancing Reinforcement Learning
The primary advancement in this major update involves a profound evolution of the post-training reinforcement learning phase. Developers expanded this process from a brief finishing stage into the central pillar of the training regimen. Consequently, they incorporated extensive supplementary training data. This data spans mathematics, algorithmic competitions, scientific research, tool utilization, and complex software engineering tasks.
To facilitate this development, JetBrains meticulously constructed an internal reinforcement learning infrastructure. During the training phase, the system launched millions of isolated sandboxes across thousands of distinct environments.
Rigorous Data Curation
Prior to training, the engineering team rigorously curated open datasets. They actively eliminated defective tests, unverifiable answers, and tasks with inappropriate difficulty levels. Following these sophisticated upgrades, the JetBrains Mellum2.1 model can autonomously navigate vast codebases, edit files, and critically evaluate its own modifications.
Advanced Agentic Programming
JetBrains asserts that the most significant breakthroughs manifest within the realm of agentic programming. The model can accurately identify the root causes of failing software tests. Moreover, it can thoughtfully draft effective remediation strategies and independently verify the final outcomes.
Regarding raw performance, this system maintains architectural parity with its predecessor, preserving its impressive baseline speed. However, the seamless integration of Multi-Token Prediction technology dramatically accelerates response times. In single-request scenarios, this advanced predictive framework boosts overall operational speed by approximately 1.6 times.
Exceptional Performance Benchmarks
JetBrains conducted rigorous comparative evaluations. They measured their newest creation against its predecessor, Qwen3.5-9B, and Gemma 4 E4B under identical testing parameters. The definitive results demonstrated remarkable efficiency during high-load operations. Specifically, the inference throughput approached nearly double the token volume of Qwen3.5-9B.
Furthermore, the model exhibited substantial improvements across multiple critical dimensions. These include general programming, algorithmic competitions, mathematics, logical tool invocation, and comprehensive general knowledge.
Local Deployment Options
Presently, the model is fully available on the Hugging Face platform. It seamlessly supports deployment on local machines or proprietary enterprise infrastructure. This exceptional flexibility allows organizations to securely retain sensitive code and proprietary data within their deeply controlled environments.
Looking forward, the official development team plans to release specialized GGUF versions. These upcoming releases will be fully optimized for llama.cpp, Ollama, and LM Studio. Additionally, they will soon introduce an MTP speculative decoding component specifically designed for vLLM architectures.











