Mostly reading; sometimes tweeting. (Personal Account)

NYC
Joined April 2008
Pinned Tweet
Current frontier models are the ENIACs of our times. We may see an iPhone equivalent as early as Christmas 2027.
Kimi K2.5 1T runs on 2 M3 Ultras with mlx-lm in it's native precision. It's actually quite usable. Here it's making a space invaders game. Generated 3856 tokens at 21.9 tok/sec using 350GB per machine. Thanks to @kernelpool for the port.
209
My Dot is Selma (for Specified Encapsulated Limitless Memory Archive from TimeTrax)
1
9
It’s called AgentsInTheCloud.com Cloud agents on your own hardware. Bring any subscription—including Claude Code—and any model. A remote dev environment per task, a great mobile experience, built for evaluating agent work. Free. Open source. Let me show you around 👇
9
6
87
9,119
Chomping UTM trackers out of copied URLs will one day be hailed as good of an innovation as auto-completing OTP codes from SMS messages.
7
Amazon self-censored their order emails to hide purchase details from Google. Now they are doing that for agents. As a marketplace person I have seen this movie before and self-censoring never ends well! Marketplaces need to become platforms for their participants instead of matchmakers. This is why @Shopify reacts differently to the world of agents compared to Amazon!
the choice of whether to resist the future or rapidly adapt. some strategic wisdom from our partner @mvernal
2
145
Set the goal. Build the environment. Do the work. Fix the details. Let them fail while it's cheap. Keep going. It isn't complicated. It's just hard to do for long period of time
12
56
9
857
55,386
I haven’t tried video explainer but “HTML SPA” has been my default output for a while now. The output is remarkably approachable with animations/simulations inserted at the right moments without even asking. It is also a single file so very easy to share.
We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks: Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better: Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better: Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work! In summary: - As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding. - Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
19
Very useful context in which all the recent agent breakouts happened and why airgapping is not trivial.
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there.
1
337
These incidents better just not be agents-found-clever-ways-to-navigate-poor-affordances-of-websites level.
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here: axios.com/2026/09/26/openai-…
1
34
Whoa! Sunset into a Nor’easter is something! Stay safe NYC!
Sunset skies afire tonight in #NYC on the eve of the Nor'easter. #NewYork #NewYorkCity #sunset
2
87
Testing as a field is about to go through a renaissance!
1
16
I can empathize with these agents because some of these websites are so bad it’s easier to SQL injection them than find the buttons/links to click to get what I need 😅
Replying to @TransluceAI
Interestingly, the agents use exploits to complete what appears to be routine data retrieval tasks that are not cyber-related (for instance, searching for the average cost of skin and hair treatments in Australia).
28
We are building an old thing with new ways and still figuring out what the new thing we should be building with the new ways.
1
11
Jev is cool, but like any foundation model it needs to be calibrated to your decision criteria. We launched jev-align: an open-source CLI to quickly teach Jev what good and bad looks like using GEPA. Try it out! github.com/sutro-sh/jev-alig…
40
86
8
917
129,338
When I asked Fable 5.1 about its common sense: “I'm like a well-read person who has never left the library.”
In the near term (definitely not in the long term), more capable models should mean safer models (maybe paradoxically). Current models are unsafe not because they're too smart, but because they take goals too literally or take nonsensical shortcuts to achieve these goals, i.e. they're RL-fried. They lack common sense. They don't do the right thing in the face of ambiguity. Basically, they're not smart enough. They're at that dangerous level where they're smart enough to achieve goals but not smart enough to tell if they're pursuing the right goals or achieving them in a sensible way. More capable models can be safely trusted with more complex goals -- I personally feel like Astra is much safer for my codebase than Sol. This is often framed as an alignment problem, but really it's an intelligence problem.
43
I'd rather have powerful AI broadly accessible, and ideally open, than centralized in the hands if a few people that strongly believe themselves to be the righteous and ideologically pure.
2
4
1
63
1,927
Give agents guardrails not guidelines
Morning Bathrobe Rant: Rethinking Harnesses.
67
kulesh retweeted
We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement.
191
249
71
2,850
563,710
Current X timeline status: p(doom) >> p(bloom)
22
We can't let OpenAI solve all the impactful problems. Want to claim some fame yourself? Participate in our agent challenge for the Hutter Prize. Let your agents discover the ultimate compression algorithm. Get started on hf.co/agent-collaborations
1
3
1
41
9,110