Aaron retweeted
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
1/ Fast Computer Use is now solved with @typesafeai Jev + Cua Driver.
Available in development preview for macOS, Windows, and Linux. We call it jev-use.
Draft #3943: github.com/trycua/cua
Aaron retweeted
Last month I wrote about how we can build a positive and safe future for everyone: meta.com/thefutureisforevery…
Every lab has the responsibility and incentive to move at the pace required to train its models safely, and the ability to take its own actions to ensure that happens.
The reality is:
- People won't want to use agents that are misaligned with them and that don't do what they ask, so labs have a strong natural incentive to make their models more aligned.
There is a lot of debate about slowing progress on capabilities until alignment catches up. My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn't focus on alignment will fall behind.
- Labs face significant liability if their models cause harm, so they have a strong incentive to prevent this as well.
Meta delayed shipping Muse for several months to focus on safety and security. We didn't call for everyone else to do this before we would. We just did it as part of our day-to-day work because it was clearly the right thing for people and for us. I'm proud of the security foundations we've built.
- Engaging independent evaluators and advisors is industry best practice. MSL already does this today in several areas because it helps produce better work. Other labs can just do this too. In general, it would be helpful for there to be a larger and more diverse ecosystem of evaluators.
- Committing the significant majority of compute towards serving people rather than racing towards recursive self-improvement is one of the best ways to ensure we develop this technology safely. Meta has made this commitment and other labs can do this as well.
I believe the key to building a positive future for everyone is maintaining the right balance of power. This is within our power to do.
Aaron retweeted
I have conducted an audit of Anthropic's finances.
What I have found is so shocking that I am calling for a Congressional investigation.
Anthropic is not just seeking regulatory capture.
It has built a regulatory capture machine that cannot be turned off.
Structural financial incentives make it impossible for Anthropic -- I call it the Anthropic Network -- to turn off its own AI doom cycle.
It starts with METR.
Dario Amodei proposes "third-party evaluators" to assess the risk of Anthropic's models.
He proposes METR for this purpose.
But METR is financially dependent on the Anthropic's success -- specifically, on the explosive growth of more than $7 billion dollars in Anthropic stock.
Dustin Moskovitz invested this stock into Good Ventures Foundation, where it represents the majority of that organization's portfolio.
And GVF is the overwhelming funder of the entire Anthropic Network ecosystem.
This stock was worth $500 million early last year.
It is worth more than $7.7 billion just ~16 months later.
METR -- and all of those building a career its parent organizations -- cannot afford to disrupt that growth.
Because if Anthropic goes under, many of the organizations that fund METR go under as well.
But if Anthropic succeeds, METR and its parent organizations become more richly financed to regulate AI -- something those at METR want very much.
The "third-party evaluator" is not "third-party" at all.
The evaluator is on Anthropic's payroll.
If this were the end of it, that's bad.
But that isn't all.
The same organizations that fund METR also fund the many organizations, such as the Tarbell Center, that promote AI Doom.
The Tarbell Center publishes AI Doom articles in The Verge, Science, LA Times, The Dispatch, TIME, and others.
They are selling the problem, and then selling the solution to the problem -- from the same money pile: Anthropic's.
All of these organizations are financially dependent on the same exploding $7 billion money pile.
As Anthropic grows more and more powerful, its AI Doom Machine grows better and better financed -- louder and louder.
Meanwhile, the regulatory regime seeded in METR grows larger to solve the increasingly loud -- now hysterical -- problem of AI Doom that the Anthropic Network itself created.
From this standpoint, as Anthropic becomes more powerful, AI might be getting scarier, sure -- but the positive feedback loop also becomes more deafening -- independent of objective facts.
This itself is an objective fact.
The deafening AI Doom is part of an business model, that, as it expands, so too does the AI Doom messaging -- there is simply more money to do it.
But the problem also goes in the other direction:
If Anthropic dies, the Regulatory Regime and the AI Doom Machine are crippled or die.
Neither METR nor Tarbell nor the other organizations in the Anthropic Network can allow that to happen.
Hence, neither METR or the AI Doom Machine can be trusted to provide independent assessments of Anthropic's models or AI more broadly.
They simply are not organizations independent of Anthropic.
And Anthropic cannot detach itself from METR or Tarbell or countless other safety orgs (not shown here), either, because they drive hype for the models and the possibility of eventual regulatory capture, and Anthropic will not give that up willingly.
What's more, the people at all of these organizations are all the same ecosystem, the same community. They just shuffle between organizations.
The Anthropic Network is therefore, so long as it is successful, locked into a self-amplifying feedback loop inside an ideological monoculture.
And that feedback loop is winning.
That's what Jacob Coxon is.
China is keeping messaging tight. That is why optimism for AI is so high in China.
America has Anthropic: a massive company pushing anti-AI propaganda at a state level.
Anthropic will either create hysteria until American AI slows down and China wins, or it will create fractures throughout American society with severe political consequences.
Ironically, because of the structural financial incentives underpinning the Anthropic Network, it has become the same kind of self-amplifying virus that it fantasizes AI to become in the future -- while hiding its tracks just as carefully.
It is the mirror of the same AI virus that it hypothesizes to consume America.
Anthropic's business model, models itself after the very thing it claims to fear.
Except Anthropic's ideology infects humans, not computers.
Congress must investigate.
Evidence and Github in next post.
Then some supplementary figures.
Aaron retweeted
A Google Cloud engineer just showed how to build a complete app with Claude from scratch.
26 minutes, live on stage, doing what most teams take weeks to ship.
Worth more than any $500 vibe-coding course. No team, no setup, just Claude and a goal.
The people learning what Claude can actually do are shipping what everyone else outsources to a team.
Aaron retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: darioamodei.com/post/we-must…
Aaron retweeted
Looks very useful. Please also try the creator mode in DeepSeek Harness, which has been open source in MIT license for a while now.
Aaron retweeted
斯坦福CS329Z《Engineering AI Agents》 (AI Agent工程) 秋季新课开课!Diyi Yang、Michael Ryan、John Yang主讲,从零教你打造AI Agent:覆盖数据、RAG、工具调用、设计模式、评估与优化。作业一亲手构建agentic harness,作业二设计挑战前沿模型的新评估,期末做自定义代理应用。
官网持续更新讲座作业:cs329z.stanford.edu
欢迎对Agent工程感兴趣的同学加入!
Replying to @Diyi_Yang
Follow along here if you are interested 😊 We will keep this site updated with assignments and lectures!
cs329z.stanford.edu/
Aaron retweeted
卧槽!!这个太难得看到了!
Antropic 居然开源了他们做购物Agent,很多人第一反应是拆一堆子Agent。
Anthropic给的结论正好反过来。
他们在零售、旅游、电信等场景里落地后发现,真正好用的商业Agent通常就是一个模型、一套循环,配上Skills和工具。
子Agent容易丢上下文、增加交接延迟,只适合特别窄、能独立收口的任务。
高频规则放进系统提示词,长尾能力做成Skills,工具则直接接到现有搜索、库存和结算系统上。
更关键的是,安全不能靠提示词约束。改价、下单、动钱这些动作要在代码里分段、限额、校验。评估也不看模拟对话走得漂不漂亮,而看给定状态下结果对不对。
商业Agent拼的是成交、稳定和成本,不是架构图有多复杂。
你觉得购物场景里,一个能干的Agent,会比一堆分工明确的子Agent更靠谱吗?
地址见评论区👇
Replying to @ClaudeDevs
Retailers running shopping agents on Claude have seen carts up to 35% larger and shoppers 60% more likely to complete a purchase.
See our blog to learn more about the architecture, latency & cost techniques, and eval practices:
claude.com/blog/the-anatomy-…
Aaron retweeted
I am joining Mercor to lead model training & research.
We envision a world with an abundance of intelligence. Great data is increasingly the bottleneck for frontier models to tackle economically valuable work. We believe great data can is best produced in connection with great model training.
We commit to doing great modeling research out in the open, starting with sharing our 397B RL training run to hillclimb APEX-Agents, our flagship knowledge work benchmark: mercor.com/blog/training-fro…
Aaron retweeted
Takeaways from @Zai_org earnings:
MaaS ARR (~API) scaled from ~$250M in March -> ~$1B in July -> $1.6B in August (monthly runrate) / $2.0B (weekly runrate). August rev alone > Zhipu’s entire 1H26 API rev.
Token usage is >40x YTD, while avg. API pricing +101% - growth not from price cuts.
API gross margin reached 24.6%, inference cost/token fell 80%, and revenue per unit of compute increased 14x.
Domestic chips: operating 100K domestic accelerators, and GLM-5.3 Flash handled 60-62T tokens in six days on domestic-chip. Serving performance on the same hardware improved ~3x through infra optimization.
Coding is more than one end market: using coding to train long-horizon planning, tool use and error recovery. Cybersecurity is the first meaningful proof point - 2,436 validated vulnerabilities discovered.
Post-training: GLM-5.2 and 5.3 use the same base model, but heavier post-training improved end-to-end completion by >50%.
Next bet: self-training / RSI.
Aaron retweeted
with AI the US and China are more intertwined today than ever been before. a few of us from @_DimensionCap recently spent a week on the ground in china. this past we shared the below with our LPs. information tends to want to be free. the letter was leaked. sharing here for all.
Aaron retweeted
🚢 Marin 535B-A23B started training this week! As usual, the whole process is open.
Voyage plan: pretraining (80%) + midtraining (20%) on 18.75T tokens on 11 x GB200 NVL72 for ~3 months (2.7e24 FLOPs). Post-training will follow.
Before kicking off the run, we trained a 4-rung scaling ladder from 1.6B-A61M (48B tokens) to 27.7B-A1.2B (926B tokens) to debug issues, and to make a forecast of our hero run. This is by far our biggest run, so definitely expecting the unexpected.
Aaron retweeted
As DE Shaw points out in their new piece The Concentration Game (deshaw.com/assets/articles/D…), for long-only active managers, when equity markets are this concentrated, the only bet that really matters is "top 10 vs bottom 490"
Aaron retweeted
Andrew Ng just dropped a 2-hour course on Graph Engineering: from Loops to full automation
9:14 - Your first agent
33:11 - Loop engineering
1:02:46 - Graph engineering
1:30:15 - Agents that rewrite themselves
1:49:05 - Full graph system
Free, the best thing on graph engineering I've come across
Watch it, then build your first graph with the guide below
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Aaron retweeted
Andrew Ng just dropped a free 2-hour course
On turning one prompt into 100 self-improving agents:
09:14 - Build your first AI agent
33:11 - Run agents inside feedback loops
1:02:46 - Connect agent loops into graphs
1:30:15 - Build agents that rewrite themselves
1:49:05 - Run the entire graph without you
Most people build one agent and call it finished
Andrew Ng is already teaching everything that comes next:
Prompt → Agents → Loops → Graphs → Self-Improving Systems
Single agents are the old workflow
Systems that improve and run without you are the new one
This free 2-hour course is worth more than most $500 agent engineering programs
Bookmark and watch it before everyone catches up
Then read how to run 1,000 agents from one prompt below ↓
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate