Building & writing about real ML systems 15+ years NLP • 3x founder (1 exit) ex Principal MLE @gett • now @withmartian

Mars
Joined June 2020
how many millennial problems can 10000 Luna's solve?
We just added GPT-5.6 Luna, Terra and Sol to the AI Frontier: aifrontier.withmartian.com/ The most interesting result isn't that Sol is the strongest model. It's this: 2× Luna > 1× Terra 4× Luna > 1× Sol Here's what happens when you stop benchmarking models one run at a time 🧵
2
10
855
After navier-stokes we can probably end debates if llm+harness can write/review CRUD-service better than you
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
8
25
1,084
feels like data-labeling companies should stop hiring CS PhD comedians
Ok, Claude deserves an award for this one: "Here's the unification that I think actually resolves your temptation, though: Tufte and Hedberg are the same doctrine in different rooms. Data-ink ratio is laughs-per-word. Chartjunk is hack padding."
2
7
540
Alex Zverianskii retweeted
What you should actually mean when you say "Frontier AI"
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive Site: aifrontier.withmartian.com/ Academic Paper: arxiv.org/abs/2606.26836
1
1
13
337
The first Oracle I can truly rely on!
We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive Site: aifrontier.withmartian.com/ Academic Paper: arxiv.org/abs/2606.26836
2
12
847
The question of the day: Will I see even one account unaffiliated with Google say something positive about Gemini 3.8?
2
2
380
This August marks 10 years for me using Linux as a main personal OS. No regrets: it made me a way better programmer, showed me the beauty of i3wm and later emacs. In 10 years, I only once had a problem with updating my laptop, and it was on Arch and required a few manual fixes, more or less comparable to the macOS experience. Anyway, I really hope that more people will understand that freedom through Omarchy
1
3
406
Code review is the most important step of the development process right now, tokens are cheap but quality is not
This quoted post is unavailable.
1
8
517
Finally, we have a model that’s fast enough for my coding style—and costs 50% less.
Ship now runs DeepSeek V4 Flash 0731. Same quality and behavior of DeepSeek V4 Flash, at half the cost, guaranteed by a quality SLA. In our evals it came in under 50% — ~49% the price.
10
319
I found it quite ironic that using the Grok @bot violates X’s rules. Aside from that, though, it’s been a very pleasant experience so far. I could do everything it does myself by running CLIs on a VPS, but this is just much more convenient.
2
115
That is my second favorite Muse, but 50% cheaper
Ship now runs Muse Spark 1.2. Same quality and behavior of Muse Spark 1.2, at half the cost, guaranteed by a quality SLA. In our evals it came in under 50% — ~45% the price.
4
204
Dear GitHub, happy to lose half the UI experiments (and all copilot) if it means text and git stop vanishing during the next capacity crunch.
1
2
187
Prefer a boring, stable GitHub that can serve a markdown file over a feature-rich one that keeps falling over.
1
128
2026 and GitHub still treats a markdown file like it needs a full CI pipeline, branch protection rules, and three layers of caching before it’s allowed to exist on the public internet. How hard can hosting text be? Harder than it has any right to be.
4
176
The AI industry’s trillion-dollar mistake is not hallucination. It’s the decision that the system must start delivering the final answer before it has finished evaluating it.
1
4
90
When the visible product is a stream of tokens, the token itself becomes the SKU. The provider sells volume. The customer wants a solved problem. Once quality is measured by the same benchmarks and a competitor offers nearly identical scores for less (as always with benchmarks), the buyer switches. Margins begin to look like commodity compute.
1
5
That is how a small UX hack becomes a trillion-dollar mistake — architecturally and commercially.
1
3
Life didn’t prepare me for this—I didn’t even have time to use up my 4.5 balance, and Grok 4.6 has already SHIPped.
Ship now runs Grok 4.6 - 50% off
7
111