Pinned Tweet
The Marionette Benchmark: Sonnet 5.5 vs. Fable 5.1 vs. Opus 5.5 vs. GPT-6 Sol
We have a new leader! A marionette only works if the bar above moves the limbs below, and juggling only works if the balls obey gravity and the hands take turns. Each model writes the whole animation without ever seeing what it made.
Claude Sonnet 5.5 xHigh | (13m 15s / $1.06)*
Somebody is actually holding the bar, a gloved hand reaching down from above. The puppet wears a bowler hat, the first hat in this whole test, with a pinstriped suit. His knees bounce with the rhythm of the throws, so the whole body is in the act, whereas Fable's legs never move. Dust drifts through the spotlight. Gravity is close to Earth and the balls cast shadows on the floor. The one miss is the same as Opus, the puppets body is driving the bar rather than the reverse.
*1st attempt was on Max setting. It thought for sixteen minutes and handed back an empty file ($1.28)
Claude Fable 5.1 Max | (24m 27s / $5.80)
Had three weeks at the top. Still the closest to real gravity, but nothing is holding its bar and the legs never move. The stage and background also less detailed than Sonnet.
Claude Opus 5.5 High | (3m 51s / $0.50)
From last week, run on high after max came back empty three times. Still the only one that draws the scene in 3D.
GPT-6 Sol Max | (7m 11s / $0.41)
From last week. Its bar really does tilt, but two of its balls still pass straight through each other.
Nobody has made the bar drive the puppet yet.
The Marionette Benchmark: Opus 5.5 vs. GPT-6 Sol vs. GPT-6 Luna
Same test as two weeks ago. A marionette only works if the bar above moves the limbs below. Juggling only works if the balls obey gravity and the hands take turns. Each model writes the whole animation without ever seeing what it made. Fable 5.1 stays in the top left because nothing has beaten it yet.
Claude Opus 5.5 High | (3m 51s / $0.50)*
*On max reasoning it thought for twenty minutes and handed back an empty file, at $2.56 a try. I tried three times and got the same thing every time. On high it finished in under four minutes, and it's the most impressive build here. It's the only one that draws the scene in 3D, so the balls pass in front of each other instead of through each other. Gravity is close to Earth and the feet stay on the floor. But it moves the bar to follow the hands, which is backwards. The hands are supposed to follow the bar.
GPT-6 Sol Max | (7m 11s / $0.41)
The bar tilts and the strings move with it, and the hands actually throw instead of sitting still. But both hands throw along the same arc, so two of the balls keep passing straight through each other. Front facing body, side facing head. It also titled its own work "A balancing act. Three balls. One deadline."
GPT-6 Luna Max | (6m 32s / $0.02)
Clean side view, and the balls never overlap. The bar never moves, so the strings just hang there. It also throws about a third higher going one way than coming back, which makes the whole pattern lean.
Claude Fable 5.1 Max | (24m 27s / $5.80)
From two weeks ago. Still has the best gravity and the best lighting of anything we've run.
The Marionette Benchmark: Opus 5.5 vs. GPT-6 Sol vs. GPT-6 Luna
Same test as two weeks ago. A marionette only works if the bar above moves the limbs below. Juggling only works if the balls obey gravity and the hands take turns. Each model writes the whole animation without ever seeing what it made. Fable 5.1 stays in the top left because nothing has beaten it yet.
Claude Opus 5.5 High | (3m 51s / $0.50)*
*On max reasoning it thought for twenty minutes and handed back an empty file, at $2.56 a try. I tried three times and got the same thing every time. On high it finished in under four minutes, and it's the most impressive build here. It's the only one that draws the scene in 3D, so the balls pass in front of each other instead of through each other. Gravity is close to Earth and the feet stay on the floor. But it moves the bar to follow the hands, which is backwards. The hands are supposed to follow the bar.
GPT-6 Sol Max | (7m 11s / $0.41)
The bar tilts and the strings move with it, and the hands actually throw instead of sitting still. But both hands throw along the same arc, so two of the balls keep passing straight through each other. Front facing body, side facing head. It also titled its own work "A balancing act. Three balls. One deadline."
GPT-6 Luna Max | (6m 32s / $0.02)
Clean side view, and the balls never overlap. The bar never moves, so the strings just hang there. It also throws about a third higher going one way than coming back, which makes the whole pattern lean.
Claude Fable 5.1 Max | (24m 27s / $5.80)
From two weeks ago. Still has the best gravity and the best lighting of anything we've run.
The Marionette Test: Astra vs Fable vs. Muse Spark vs. Gemini
A marionette only works if the bar above moves the limbs below. Juggling only works if the balls obey gravity and the hands take turns. The model has to hold both in its head and write out the coordinates without ever seeing what it made. Although not a frontier model, we included Gemini 3.8 Flash so it feels included and won't form a complex it takes into adulthood 🫂
Claude Fable 5.1 Max | (24m 27s / $5.80)
Best of the four. Correct gravity, best lighting/shadows, the only one that built a proper two bar rig. But it drew four strings, its legs never move, and the top bar isn't connected to anything and nothing is holding it up.
GPT-6 Astra Max | (17m 42s / $2.03)
Nice suit, plenty of strings. Its near knee bends backwards. Gravity wise it looks like he's in outer space, so the balls hang like it's juggling on the moon.
Muse Spark 1.3 Max | (4m 42s / $0.18)
Most complete rigging, seven strings running to head, hip, both knees and both hands. Open curtains and a city skyline (best background other than Fable IMO). Middling physics, rough suit, but at least it has real arms unlike Gemini.
Gemini 3.8 Flash Max/Extended Thinking | (1m 18s / In-App)
Despite being a Flash model, Gemini has the best juggling animation outside of Fable. The balls also have faint shadows on the ground, which GPT-6 Astra failed to do. The only one that worked out that a strict side view flattens a cascade into a single plane, and it corrected for it. Good gravity. But the arms are bare sticks inside a sleeveless suit and the background is empty.
Nobody made the bar drive the puppet.
as a completely uninvolved third party, watching "The Crucible" play out in 2026 as the lynch mob push EA conspiracies to preserve their own financial interests in AI is fucking wild 😂. there is not a show on tv that rivals this
There is an unseen hand pushing AI safety as a political ideology.
Effective Altruists believe a tiny group of technocrats should decide how much technological progress the rest of us are allowed to have…. and how many shrimp your life is worth.
You are witnessing their attempted coup.
thefp.com/p/dangerous-ideolo…
The strong pushback to AI regulation by the community here should always be contextualized against how everyone here stands to make money off of it continuing to accelerate. Scoble included
Robert Scoble just DM'd me to pitch why Anduril should join the other AI companies paying him to shill for them. He is charging $2,000 per X post, more for videos/blogs/astroturfing.
I would share screencaps, but he deleted the message. What a Sunday.
nitter.cf/PalmerLuckey/status/18…
It's not the best AI model in the world, but one thing I do love about Grok is it will just do the thing you ask without any unnecessary fuss, unlike Claude and others
YIKES
So, lots of people have been raving about @noahrshinn's instinct.com, but when I asked the agent about documentation, it said there wasn't any - but it could answer any questions that I have.
Come to find out there is a privacy policy at the root of the Instinct website, but not accessible from the backend where the agent sends you, and where you set up your connections (app.instinct.com).
It is wild to me that people would connect their work accounts, credit cards, and sensitive login data to a system without a readily accessible privacy policy. It is also fishy to me that the Instinct agent would say there isn't a privacy policy, when there clearly is. It is a cutting-edge AI agent, right?
After diving into their privacy policy and terms, I wouldn't give Instinct access to my work inbox under its current policy. Connecting an account can expose conversations with clients, coworkers, and everyone else who trusted you with their information.
Disconnecting an integration doesn’t delete what Instinct already copied. Model training is opt-out, with a safety-review exception, and opting out doesn’t undo prior training. Google Workspace API data has explicit protections, but you shouldn’t assume those apply to everything else you connect or upload.
And this is the huge one - the terms also let the agent enter binding agreements on your behalf while putting responsibility for unintended actions on you.
I work with AI agents. I want them to be useful. But before an agent gets access to client conversations or permission to spend money, I want enforced approval controls, clear retention limits, and deletion I can verify.
I wouldn’t connect sensitive accounts until those questions are answered. Read the privacy policy and terms before granting access. And remember, kids. If a product or service is free, you're the product.
It is a shame, because the respose times of Instinct are truly amazing, but if you connect your work email, iCloud, and plop your credit card data into a product that has basic HTML as the home page, and a fairly predatory privacy policy - you're a braver person than I.
another way to think of this:
"We just made the worst version of our AI future way more unlikely to happen, and the cost was a 10% stock drop the following week that will eventually recover over time"
AI stocks will drop 10%+ on Monday morning
Brace for impact folks
@DarioAmodei just unwound the AI trade with a blog post
ClearText AI retweeted
Ok @DavidSacks time to update your script again
The spirit of free speech is in the air on 9/11 🫡
Can someone explain to me why OpenAI and Anthropic allow any employee to tweet anything they want.
Apple, Google, Facebook, Nvidia, Microsoft, Tesla, Uber and Airbnb all have obvious and granular rules about e
employees speaking for the company.
Which is to say, employees should never say anything publicly unless it's cleared with comms first
which isn't intended to censor anyone, just to make sure the company is aligned and marching in unison
I could never imagine tweeting the stuff that OpenAI and Anthropic employees are tweeting without checking with the CEO and founders first!
So are these P-doomers clearing their tweets with the CEO and comms?
Inspired by @vinn_ayy for resigning from Boiling People Alive Labs and surrendering all his equity. wow.
People don't like to accept the fact that intelligent folks are also subject to GroupThink. X is not a welcoming place for anything other than AI accelerationism.
There's also the giant shadow of financial interest fueling a lot of the fervor against pausing / slowing down.
A lot of otherwise smart people on Twitter seem 100% convinced AI risks are all fake and stupid and part of some marketing ploy. What is surprising is that some of these people seemingly also believe that AI’s positive uses are on an incredible trajectory of increasing capability with no end in sight. VC Twitter is particularly infected by this pattern. It’s not really coherent.
Most positive use cases for AI have a corresponding “dark version”. If you are super human at coding, you are also super human at hacking. If you are superhuman at structural engineering you are likely superhuman at finding structural flaws to knock buildings down. If you are superhuman at designing drugs, you are superhuman at designing novel undetectable poisons. If you can cure viruses, you can create them. Some of these “dark versions” are not so bad, and some are actually pretty scary. Either way these are real societal and technical problems that need to be solved to get the good stuff and avoid the bad stuff. We’re experiencing the first of these with coding and computer security which is the most advanced, but that won’t be the last. I think we’ll be able to solve these problems, but they aren’t solved yet and if you believe in continued AI progress they are surely coming.
But “bad people using AI” is not the only problem. Uncontrolled AI autonomously doing bad things, despite sounding kind of nutty, is also something we should be concerned about. AI “killing us all” is not the most likely outcome, but the chance of a major civilization-wide catastrophe doesn’t have to be very high for it to be a concern. Again, this is only a problem if capabilities advance to a point that AI can do really crazy things on their own, which hasn’t happened yet, but I think the Hugging Face incident is a good example of the outline of how things can go wrong when capabilities outpace alignment. We should be glad that the only available bad thing right now is hacking, which isn’t all that bad. It’s clear that as you get to superhuman capabilities you need a level of alignment and control that is correspondingly superhuman. Humans have plenty of misalignment problems themselves (serial killers, mass shooters, tyrants, etc), but it’s a manageable problem because most humans can’t do that much damage and we’ve developed systems to prevent dangerous humans from getting too much power.
Talking about these issues is just common sense. It’s not a sign of some kind of neuroticism or pessimism. These are just hard problems that it’s very important to solve for AI to have a positive impact. I’m pretty confident we will solve them. But we haven’t solved them yet, and to my mind we are clearly on a trajectory of rapidly increasing capabilities which means this is important. Mocking people who are worried about this or talking about it, without anything substantive to say about how we can be sure these problems won’t arise, is not really a very helpful contribution.
The Marionette Test: DeepSeek V4.1-Flash Max | (5m 00s / $0.04)
130x cheaper than Fable. Five live strings, a real cascade with the hands taking turns, and the only entrant that admits someone is up there: two strings run off the top of the frame instead of a bar hovering in the void. But the hands barely move, and it set the loop to half the pattern's true length, so a ball teleports across the chest every 1.5 seconds.
The Marionette Test: Astra vs Fable vs. Muse Spark vs. Gemini
A marionette only works if the bar above moves the limbs below. Juggling only works if the balls obey gravity and the hands take turns. The model has to hold both in its head and write out the coordinates without ever seeing what it made. Although not a frontier model, we included Gemini 3.8 Flash so it feels included and won't form a complex it takes into adulthood 🫂
Claude Fable 5.1 Max | (24m 27s / $5.80)
Best of the four. Correct gravity, best lighting/shadows, the only one that built a proper two bar rig. But it drew four strings, its legs never move, and the top bar isn't connected to anything and nothing is holding it up.
GPT-6 Astra Max | (17m 42s / $2.03)
Nice suit, plenty of strings. Its near knee bends backwards. Gravity wise it looks like he's in outer space, so the balls hang like it's juggling on the moon.
Muse Spark 1.3 Max | (4m 42s / $0.18)
Most complete rigging, seven strings running to head, hip, both knees and both hands. Open curtains and a city skyline (best background other than Fable IMO). Middling physics, rough suit, but at least it has real arms unlike Gemini.
Gemini 3.8 Flash Max/Extended Thinking | (1m 18s / In-App)
Despite being a Flash model, Gemini has the best juggling animation outside of Fable. The balls also have faint shadows on the ground, which GPT-6 Astra failed to do. The only one that worked out that a strict side view flattens a cascade into a single plane, and it corrected for it. Good gravity. But the arms are bare sticks inside a sleeveless suit and the background is empty.
Nobody made the bar drive the puppet.
Today I made the costly mistake of using Claude extra usage instead of just signing up for an entirely new plan. And ended up spending $90 on roughly 4 or 5 hours of work
Its interesting watching this community of Highly Educated Free Thinkers rush to a conspiracy that the "CIA is running a psyop to turn people against AI" because they can't fathom reasonable people being worried about this. 🍿
ClearText AI retweeted
Here's a few reasons why people have so much trouble predicting the future:
1) Fail to realize that even the best superforecasters' predictions drop off dramatically past a three year time horizon and that most folks are worse than dart throwing monkeys at future predictions.
2) Fail to realize you literally can not see black swan inventions coming around the corner. If you predict the future of Germany in 1439, then in 1440 your predictions are completely wrong because of the Printing Press.
If you're an 18th century farmer you can't see a web developer job because it exists on the back of countless developments and inventions you can't predict.
3) They change one variable and hold all other variables the same. i.e. AI advances and nothing else does, no parallel discoveries or innovation, no solutions, no mitigations, no societal or cultural changes.
4) Predict unlimited resources and zero friction in the real world (dust, disconnects, diffusion, etc) to slow/divert/change/impact the development. All changes experience equal and opposite reactions.
5) They mistake their ability/expertise in a domain for a parallel/orthogonal ability to predict the future of that domain and its impact on the world. Two different skills and they do not usually overlap (though very rarely they do.)
6) What I call "classic sci-fi or Jules Verne syndrome", which is similar to one variable changes. It's like in Jules Verne when one guy gets the submarine and nobody else does. But life is more like cell phones, lots of people getting them over time in a diffusion curve.
7) They mistake exponential curves as infinite always and never see an S curve coming.
8) The see infinite resources (compute/memory/learning upper limits/improvement) and no limitations.
On one hand, you get a really cool AI assistant that helps organize your email and shop.
On the other hand, a 10% risk of AI yeeting humanity into oblivion during one of the most fragile moments in the American empire. Very tough choice.
The Marionette Test: Astra vs Fable vs. Muse Spark vs. Gemini
A marionette only works if the bar above moves the limbs below. Juggling only works if the balls obey gravity and the hands take turns. The model has to hold both in its head and write out the coordinates without ever seeing what it made. Although not a frontier model, we included Gemini 3.8 Flash so it feels included and won't form a complex it takes into adulthood 🫂
Claude Fable 5.1 Max | (24m 27s / $5.80)
Best of the four. Correct gravity, best lighting/shadows, the only one that built a proper two bar rig. But it drew four strings, its legs never move, and the top bar isn't connected to anything and nothing is holding it up.
GPT-6 Astra Max | (17m 42s / $2.03)
Nice suit, plenty of strings. Its near knee bends backwards. Gravity wise it looks like he's in outer space, so the balls hang like it's juggling on the moon.
Muse Spark 1.3 Max | (4m 42s / $0.18)
Most complete rigging, seven strings running to head, hip, both knees and both hands. Open curtains and a city skyline (best background other than Fable IMO). Middling physics, rough suit, but at least it has real arms unlike Gemini.
Gemini 3.8 Flash Max/Extended Thinking | (1m 18s / In-App)
Despite being a Flash model, Gemini has the best juggling animation outside of Fable. The balls also have faint shadows on the ground, which GPT-6 Astra failed to do. The only one that worked out that a strict side view flattens a cascade into a single plane, and it corrected for it. Good gravity. But the arms are bare sticks inside a sleeveless suit and the background is empty.
Nobody made the bar drive the puppet.
Watching the community fall all over themselves to crown any benchmark that GPT-6 Astra tops as 'credible' and any that it doesn't as "benchmaxxed" 😂