ThorneSophia CodeCheese retweeted
What if the jagged frontier is mainly math + code (which you can push arbitrarily far with RLVR), and everything else starts to plateau because it is still bottlenecked by human generated data?
Model performance in non-verifiable areas has kept improving steadily, albeit much slower than for math and code. But is that steady improvement a side effect of a higher G (itself driven by RLVR), or only a function of the amount of new human data getting injected into training (which is still continually happening on a massive scale)?
A lot of things depend on the answer to this question
ThorneSophia CodeCheese retweeted
Honestly stunned by result 128
Shortest Common Superstring went from 3x optimal to 7/3 over 35 years of human work. OpenAI's model got it to exactly 2. And it didn't use greedy, the algorithm the 1988 conjecture was about. Lean-checked
I animated the proof in 6 minutes:
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
github.com/openai/math
ThorneSophia CodeCheese retweeted
Can vibe coding produce production ready software without an experienced engineer?
ThorneSophia CodeCheese retweeted
Anthropic to adopt Samsung as foundry partner
Always been bullish on Samsung and now you can see their foundry business making huge moves.
Fast becoming a strong, big, viable 3rd foundry player.
On another Samsung news, SOUTH KOREAN media reporting that SAMSUNG’S 12-HIGH HBM4E PASSES QUALIFICATION TESTS AT $NVDA and other customers
Samsung 🚀🚀
$TSM $DRAM $INTC $EWY
ThorneSophia CodeCheese retweeted
What happens if the US (through the utilization of frontier models) develops a cure and starts price gouging the rest of the world, especially the drifters, something like $1k per person? Or maybe even higher?
A cure to a disease, so simple as the Pneumonic Plague, is literally a few prompts away (plus a day of work of a few thousand agents) for the currently unreleased and gated models of OAI and ANT.
ThorneSophia CodeCheese retweeted
7 Linux commands that give you instant answers on a struggling server: 🐧
1. uptime: Check 1, 5, and 15-minute load averages against your total CPU core count.
2. vmstat 1: See CPU idle percentage, memory swapping, and IO wait blocks every second.
3. mpstat -P ALL 1: Check if one single CPU core is pinned at 100% while others sit idle.
4. iostat -xz 1: Measure disk read/write bandwidth and average request queue size.
5. free -m: Check actual available RAM without getting confused by OS buffers.
6. ss -s: Summary of total open TCP connections and time-wait sockets.
7. pidstat 1: Pinpoint the exact process consuming CPU, disk, or memory in real time.
Save this cheat sheet for your next server investigation. 📌
ThorneSophia CodeCheese retweeted
quick heads up if you use Gemini for free
starting Oct 9, you'll only get flash-lite, google's smallest model. flash and pro go away
on the $4.99 ai plus plan, you keep flash but lose pro. ai pro and ultra keep all three
your usage limits don't change. you just get fewer models to pick from
I sympathize with programmers who worked on systems they never cared about, for customers they didn't connect with, and took all their working joy from the mechanical act of turning tickets into code. You have to care about the software to be excited about making it w/ agents!
ThorneSophia CodeCheese retweeted
$IONQ UPDATE:
We hosted IONQ CFO/COO Inder Singh in investor meetings. Key takeaways include:
1) Superion 256 PQ system ramping in 2027E with ASP potentially in $25-30M,
2) IONQ continues to see upside from deployments in CSP clouds as it works on modular scalable platforms for its next-gen 10K PQ system,
3) upside with research labs/sovereigns as US pushes adoption of Post-Quantum Cryptography (PQC) standards in 2027E, with 2030-35E compliance targets, and
4) SKYT integration driving continued performance/roadmap acceleration, with own foundry, potentially set to reduce fab cycles from 8 to -2-3 months and potentially drive foundry revenue up ~100% y/y. Reiterate IONQ at Outperform with a $52PT, as we see its 256PQ Superion system driving 2027E upside, while vertical integration drives a scalable platform and also provides a flagship onshore industry QC foundry.
$IONQ | Mizuho Securities 𝗺𝗮𝗶𝗻𝘁𝗮𝗶𝗻𝘀 𝗕𝘂𝘆 on 𝗜𝗼𝗻𝗤, 𝗜𝗻𝗰., maintains PT at $𝟱𝟮
ThorneSophia CodeCheese retweeted
My main problem with Dots, is it's not sparks joy yet.
ChatGPT sparks joy, Codex sparks joy, Dot is not yet.
ThorneSophia CodeCheese retweeted
Where does the time actually go when an AI agent runs?
Total response time: 2,400ms.
The breakdown:
- User Network Roundtrip: 80ms
- Query Embedding Generation: 60ms
- Vector Database Hybrid Search (pgvector): 40ms
- Prompt Assembly & Context Injection: 10ms
- Time to First Token (TTFT) from LLM API: 650ms
- Model Output Generation (300 tokens): 1,200ms
- Pydantic Validation & Tool Execution: 360ms
If your users complain about speed:
Optimizing the database search saves 20ms.
Streaming tokens with Server-Sent Events (SSE) saves 1,200ms of perceived wait time.
Always stream your outputs.
Are you streaming LLM tokens or waiting for the full response to finish?
ThorneSophia CodeCheese retweeted
A few things to think about as AI gets more capable:
1. Intelligence is becoming abundant and limitations are moving somewhere else. Good judgment, good questions, good data and knowing what to do with an answer are becoming more valuable than simply having access to a model.
2. The cost of trying an idea has collapsed. You can build a prototype, analyze a market, test a workflow or write the first version of a product in an afternoon. That changes how I think about ideas. There is less reason to debate something for three months when you can test it this week.
3. Most of what gets produced with AI will be mediocre. That makes taste more valuable. When everyone can produce ten versions of something before lunch, knowing which one deserves to exist becomes a real advantage.
4. A lot of AI value will sit outside the model. Models will improve, prices will change and providers will come and go. The data, workflows, infrastructure and relationships built around those models can become much more durable assets.
5. Access and ownership are two very different things. Renting intelligence from a provider is incredibly useful. Having some control over the compute, data and systems your business depends on gives you a different kind of leverage.
6. AI is also changing the size of a company that can do meaningful things. A small team with access to good models, software agents and the right infrastructure can take on work that would previously have required a much larger organisation.
7. This is why I keep coming back to user-owned infrastructure. If AI becomes one of the basic layers of the economy, I don't think every person and business should have to remain a permanent tenant of someone else's infrastructure.
There is a lot of excitement around what AI can do, we should be more interested in what kind of economy we build around it.
ThorneSophia CodeCheese retweeted
the way opus 5.5 spatially align text in images in iterations is so beautiful, openai's web harness is lacking in this capability of analyzing plots it just constructed w/ code.
image_1: bad spatial text alignment,
image_2: next self-corrected iteration of great text alignment.
ThorneSophia CodeCheese retweeted
Top 10 system design resources that actually help in day-to-day backend work + interviews:
1) Designing Data-Intensive Applications (Kleppmann, book)
Replication, partitions, consistency, storage engines. Great for arguing tradeoffs.
2) Site Reliability Engineering (Google, book)
SLIs/SLOs, error budgets, capacity planning. The ops side most designs ignore.
3) System Design Primer (GitHub)
Free. Lots of common components (LB, cache, queue) with quick pros/cons.
4) AWS Well-Architected Framework docs
Concrete checklists: reliability, cost, security, ops. Use it to review your own design docs.
5) Martin Fowler’s architecture articles (fowler.com)
Strangler fig, event sourcing, CQRS, microservices failures. Good for naming patterns correctly.
6) High Scalability (highscalability.com)
Real-ish architecture breakdowns. Helps with “what does a big version look like” intuition.
7) Designing Distributed Systems (Brendan Burns, book)
Kubernetes-style patterns: leader election, work queues, sidecars. Practical distributed primitives.
8) MIT 6.824 Distributed Systems (course + labs)
Harder, but it forces you to reason about failure modes, not diagrams.
9) Jepsen blog + Knossos writeups
Learn what breaks under partitions and clock issues. Makes “exactly once” claims disappear fast.
10) Practice project: build a tiny production-ish service
API + Postgres + Redis cache + queue worker + tracing + load test (k6). Add one failure drill: kill DB, throttle downstream, replay a poison message.
ThorneSophia CodeCheese retweeted
This tool is literally the free version of OpenAI Dots.
it's called opendots: always-on ai coworkers that each get their own computer, and work by text, calls, or slack. you host it yourself.
> each dot gets its own browser, files, and shell
> works with any openai-compatible model
> call a dot while it keeps working in the background
> mention a dot in slack, continue in the thread
mit licensed. open source.
I use AI to create pipelines that convert a book in HTML format hosted elsewhere into an interactive format used on ChapterPal.
Once a book is converted, there's a validation phase where I scroll through the entire converted book, spot conversion issues, and ask the AI to fix the converter.
Because third-party books can be in hundreds of different formats, each conversion pipeline results in its unique conversion artifacts.
When I scroll through a converted book, I easily see such artifacts.
But now we have models with vision and agents capable of opening webpages in a browser and scrolling through them to compare the source and the target, right?
Every time a new model or a new harness version is released, I ask the fucking PhD-level AGI/ASI to visually compare the source to the target and spot conversion artifacts. I don't give examples of what those artifacts might look like because they are all different for different sources.
The fucking PhD-level AGI/ASI is consistently useless. It can see that some words are missing, but everything about leaked HTML, math that is supposed to be in LaTeX but is in plain text, unescaped dollar signs, a list that is supposed to have bullets but doesn’t, and so on, it's all okay to the PhD-level AGI/ASI.
That's why I'm so annoyed hearing that those tin cans are sooo(...)oooo intelligent, beat the dumb ARC-AGI benchmarks like crazy, but cannot spot what looks "weird" in a target webpage compared to the source.
ThorneSophia CodeCheese retweeted
Thank you for telling me that so much, its so good.
Everyone should see this post. Its so worth it, a lot better than claude code's computer use
everyone thinks codex computer use only works inside codex
it doesn't, it's a local mcp server in the chatgpt mac app and claude code can just call it, even headless with claude -p
same 8-task test: opus 5.5 on codex's engine got 6/8, codex itself got 6/8, cua driver got 3-4/8 at 4x the cost
cua's background mode on mac only goes through accessibility, so canvases and drags just don't land
codex's engine sends real clicks and drags to the app without touching my cursor
no hover though, and it's unofficial, so enjoy it until a chatgpt update breaks it
give this to your claude:
"Set up Codex's computer use as an MCP server for you on my Mac. Find the "cua_repl" entry in ~/.codex/plugins/cache/openai-bundled/unified-computer-use//.mcp.json and register it as a user MCP server called "codex-cu" with the same command, args and env. Then test it by using Calculator in the background to work out 12 × 12."
ThorneSophia CodeCheese retweeted
I can't tell if Opus 5.5 is way better than Astra, or openAI nerfed Astra lately, but man is 5.5 miles better for design tasks as well.
Sonnet 5.5 is also surprisingly good and fast -- I was able to work the other day for about an hour on 5 PRs with 1% usage remaining.
ThorneSophia CodeCheese retweeted
This is so true!
At OpenHands we have have reached the "bonus stage" from the original post.
Our process has been rebuilt around the assumption that if an issue is opened on the repo, we are guaranteed get a vibe-coded PR opened, so our defense is: (1) make sure that all the issues are maximally clear through automated issue triage, and (2) automate code review so that we only accept high-quality contributions.
All of these are powered by OpenHands automations:
github.com/OpenHands/extensi…
Stages of grief about the effect of AI on open source.
Denial:
- AI code is slop! Nobody serious will use this.
- Open source will keep working exactly like it always has.
Anger:
- These AI slop PRs are ruining GitHub!
- Ban AI-generated code!
- If you close PRs, you're not open source anymore!
Bargaining:
- Okay, AI can write code, but only if a human reviews every line.
- Maybe adversarial AI review will make arbitrary external code intake safe again.
Depression:
- Wait, if code generation is basically free, what is a contribution anymore?
- Is GitHub-style open source culture just going away?
Acceptance:
- Open source is still open source, even if the contribution model evolves.
- Good issues, test cases, benchmarks, specifications, and ideas become the scarce inputs.
- Upstream may stop accepting arbitrary code while forking becomes much easier.
Bonus stage! Reconstruction:
- Projects redesign around maintainer-controlled agents, issue-driven development, stronger test/validation infrastructure, and new ways for contributors to earn trust.