Xiaomi Mimo 2.6 live post-training dashboard
- joelwallis - 44605 sekunder sedanI been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
- dr_dshiv - 38391 sekunder sedanWell, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.
- passive - 37887 sekunder sedanNeat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far.
I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.
The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.
- ricardobeat - 37325 sekunder sedanFor reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
- krm01 - 45203 sekunder sedanThis is pretty neat. What would be a good reason for the other Model providers to not do this?
- liuliu - 44709 sekunder sedanWhen you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.
- fzysingularity - 39749 sekunder sedanVery cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.
- ProfessorLayton - 44685 sekunder sedan2.6 Pro: >started 2026-09-15 10:32 UTC
For some reason I thought training took much, much longer than what the progress bar suggests.
This is really neat, I'm currently using mimo 2.5 pro, and it's decent (or great given the price). Hopefully their next one is multimodal.
- ssn2000 - 17463 sekunder sedanTotal run cost is $1.2M until now, what resources are they using to train their model? Wish they shared more details on that and what the MFU metrics are.
- ttul - 27823 sekunder sedan$5 per second if my eyes don’t fool me. That’s ~$432K per day. Enough to rent 3,000 B300 nodes on Modal.
- thehamkercat - 44793 sekunder sedanThis is crazy, but sadly anthropic/openai will never do this, what has happened to this world, where chinese companies are more open than US or even EU companies
- speedgoose - 45024 sekunder sedanI didn't know 2 thirds of the training data would be source code.
- kkotak - 21050 sekunder sedanWouldn't us observing this break down the model superposition and make it dumber? :)
- rao-v - 26220 sekunder sedanI absolutely love that someone is doing this! Why isn’t IBM for Granite or Google for Gemini?
If you are going to develop a near frontier model, and you don’t think you have special sauce up your sleeve, why not making training runs and RL environment scores etc. visible to the world?
I’m genuinely learning quite a bit just from the dashboard
- jstummbillig - 13875 sekunder sedanWow, spending money on training an almost-frontier-model is much more time intensive than I thought it was.
- monneyboi - 2840 sekunder sedanRefreshing, now let's make this a default feature. I imagine a "Upcoming models" list with links to these kind of dashboards.
- rozab - 44736 sekunder sedanWhy are they doing this? To try head off accusations about distillation?
- wolttam - 45296 sekunder sedanHah, it would be great to see more labs pick this up.
- ernsheong - 36862 sekunder sedanMino 2.5 has been my workhorse for coder and tester agents (the ones planner agents delegate tasks to)
- wg0 - 9858 sekunder sedan"Slow down this much openness in AI or we won't get our trillion dollars valuations!"
Google had this GPT long go and a wise man within Google noted:
"We don't have any maot neither does anyone else."
The AI bubble burst is guaranteed and is only delayed by IPOs.
- thenews - 24510 sekunder sedanbeen using the 2.5 mimo for side projects, works amazing
- dr_kiszonka - 28798 sekunder sedanVery curious that everyone here (so far) seems to assume this dashboard presents real data.
- esafak - 41047 sekunder sedanThat's the kind of transparency we need! That DeepSWE benchmark puts it in frontier territory: https://artificialanalysis.ai/agents/coding-agents?coding-ag...
- Alifatisk - 8032 sekunder sedanCan we call this open AI?
- tcbbd - 22375 sekunder sedan[dead]
- heronbank - 27728 sekunder sedan[dead]
- sinuhe69 - 10210 sekunder sedan[flagged]
- Toslink - 36821 sekunder sedan[dead]
- pppkin - 24224 sekunder sedan[dead]
- impulser_ - 42920 sekunder sedanThe Chinese labs are just making fun of the US labs at this point.
Where is the cool shit from the US labs?
- dude250711 - 35494 sekunder sedanDistillation in real-time? Very interesting!
- hsbalanxvxjsmab - 24027 sekunder sedanThis is so very clearly fake? See the message stating the flash 2.6 flash run was restarted and 0 graphs correlate that restart
- levocardia - 44420 sekunder sedanYou'd think they would make it less obvious that they are running their whole operation with Claude
Nördnytt! 🤓