News · New this week
JetBrains Mellum2: open 12B coding model, 2.5B active
Mellum2 has 12B parameters but wakes only 2.5B per token, and you can serve it yourself with vLLM.
JetBrains open-sourced Mellum2, a 12B coding model that only wakes 2.5B parameters for each token. The weights are Apache 2.0, so you can run it on your own hardware.
How the 2.5B figure works
Mellum2 is a mixture-of-experts model. Instead of pushing every token through all 12B parameters, it holds 64 experts. A router looks at each token and picks 8 of them. The rest sit idle for that token.
That's the whole trick. The flow is token, router, 8 experts, output. You store a 12B model, but each token only pays for about 2.5B worth of compute. It stays fast because most of the model isn't doing anything at any given moment.
The model name says the same thing. In Mellum2-12B-A2.5B-Instruct, the 12B is total size and the A2.5B is the active part.
What you get
The weights are Apache 2.0 and live on Hugging Face. The context window is 131k tokens. There are two versions, Instruct and Thinking, and tool calling is included for agents.
Both variants are built with agents in mind. If you're wiring a model into an agent loop, tool calling is the feature you'd check first.
Serving it with vLLM
You can serve the Instruct model with vLLM and get an OpenAI-compatible API running on your own GPU. The reel shows it in three lines.
$ ORG=JetBrains
$ M=Mellum2-12B-A2.5B-Instruct
$ vllm serve $ORG/$M
Once the log shows "INFO: Application startup", the server is up. Because the API is OpenAI-compatible, a client that already talks to that style of endpoint should only need a different address. We'd point an existing tool at it and see how it behaves before changing anything else.
Where it fits
JetBrains calls Mellum2 a fast, free helper for coding agents and says it is not a frontier model. That framing matters. Don't expect it to replace the biggest model in your stack for hard reasoning.
The better fit is the small, frequent work an agent does, where you'd rather not pay for every token. Autocomplete-style tasks are the obvious example, and the low active parameter count is what keeps them quick.
Try serving the Instruct model with vLLM on a GPU you already have, then hand it one small task from your agent and see whether the speed and cost trade works for you.


