ben 🐝 human 🧬 AudioTrack 🎺 AI Tinkerers Seattle Chapter 🤖 OpenAI Ambassador 🧑‍💻

Seattle
Joined October 2011
my new audio engineer is @openinterpreter's 01
71
159
69
1,082
245,915
Bee 🐝 retweeted
Semi is here Today, we're launching high volume production
505
2,110
347
21,673
1,904,546
Bee 🐝 retweeted
i'll be heading to SF for @OpenAIDevs DevDay! come say hi and chat, before i head to yosemite to tick off a lifelong bucket list item
15
4
69
2,930
Bee 🐝 retweeted
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why. Measured five different ways, August delivered dramatically fewer thinking tokens than July.
87
143
70
2,004
316,888
Really happy with Sol for non-tech work so far. Astra was a turning point for me with its writing quality. I could just say "hey, can you send an email like this," and it wrote it out in a way that actually made sense and didn't seem like AI slop. But it was tough watching 2% usage go to every email I sent. Now with Sol, it’s so great to see similar writing quality. Being able to send 30 emails and still not use even one percent of usage on fast mode is amazing. It feels like a great parallel partner. I don’t have to think of everything I need to do all at once. I can just send off one-offs as they come to me and not have to worry about usage or anything like that.
1
71
i am feeling the opus 5.5 abundance. I am an opus 5.5 abundance liberal. we should basically put enough solar panels in the desert to make solar near~baseload for AI and fuel production so that opus 5.5 class models are available as a universal basic intelligence package.
16
23
4
549
10,117
please timeline, show me more opus 5.5 javascript
1
2
95
turn it up and sing along
opus 5.5 just dropped its first pop punk single with a music video! everything you see and hear is generated from javascript code that claude wrote. no samples, no libraries 🔊
1
103
Bee 🐝 retweeted
GPT-6 Astra 用下来,我真的开始怀疑 OpenAI 的方向走错了。 它很擅长目标明确、有固定正确答案的任务。但碰到开放、发散、没有标准答案的事情,就特别别扭。自由探索一个想法、写点有意思的东西,或者面对模糊的问题自己判断值得往哪走,这些才是我觉得它明显欠缺的能力。从 GPT-5 开始,我就有这种感觉。 很多任务开始时,连最终想要什么都没确定,需要在探索过程中逐渐发现。GPT 却总像在等你把题目出完整,把验收标准列清楚。可如果这些都得我先想好,最需要智能的那部分工作,我已经自己做完了。 我怀疑这也解释了它为什么那么爱堆防御性代码、写一大堆测试。测试能给出明确的通过或失败,它很容易一直围着这些可验证的结果打转。至于方向有没有价值、有没有更好的可能,就需要另一种判断力了。 Claude Fable 在这类任务上给我的感觉就好得多。它会从理解你想做的事情出发,顺着不完整的想法继续展开,给你一些原本没想到的可能性。即使没有固定答案,也能把事情往前推进。 这让我怀疑 OpenAI 是不是太执着于能判对错、能计分的能力了。真实世界有太多任务根本没有标准答案。在这种自由探索和开放任务上,我觉得 Anthropic 才是真正领先的,光看表面跑分看不出来。
224
345
142
3,156
823,905
Jev is good for realtime canvas and voice, it feels like a helpful part of the setup so far
2
2
389
Feels like Jev should go the webMCP route and be ultrafast tool extensions for LLMs
Yesterday I said Jev would open a ton of doors... 24 hours later, this exists. Cua built a 2.8MB model that scored 99.7% on their form-filling eval. Hosted Jev scored 83.6%. Not to mention it's FREE and only 706K parameters. Small enough to run locally with not even 1gb or ram. Fast enough to make decisions in one pass. And specialized enough that your agent doesn't need to call a giant LLM for every tiny action. Think about what this unlocks. Every repetitive computer task could eventually get its own tiny specialist: • forms • CRM updates • data entry • browser actions • document routing • UI decisions Then one powerful agent just routes work between them. We are going to see some ridiculous stuff built from this as well.
2
2
233
i love that anytime theres a realtime 'breakthrough' we all immediately go to tldraw. the platonic realtime OS
1
179
Bee 🐝 retweeted
Replying to @LonLigrin
Jev is the present
1
1
36
Bee 🐝 retweeted
tapping the brian eno sign
un-appreciated artefacts now become the appreciated when it's clear AI hasn't touched it
3
2
18
844
Can we all just pause for a sec and reflect on how fucking important Tailscale is, and how nobody ever talks about them? I don't think I've ever seen a more useful service with a more invisible corporate footprint.
268
231
75
5,313
559,632
卧槽,JEV 的出现很可能就是自动智能交易的起点。 昨天JEV在AI 圈刷屏,今天 Monad 工程师就用JEV做了一个自动交易机器人。 实时读取 MON/USDC 价格,JEV 不负责写分析报告,也不跟你解释一堆逻辑,它只做最核心的事,判断 Buy 还是 Sell,再给出置信度,然后直接把交易发到 Kuru 的链上订单簿执行。 这一下我突然明白 JEV 为什么要做成一个不会说话的模型了。 交易根本不需要 AI 每 300ms 给你写一篇小作文,它需要的是: 行情进来→判断→下单→新行情→重新判断,而且这一整套循环必须足够快、足够便宜。 而 Monad 现在刚好每 300ms 出一个区块,Kuru 又是链上订单簿。高性能链负责执行,JEV 负责决策,两个东西拼起来,已经有点AI 原生交易系统的味道了。 未来的交易 Agent,不一定需要一个会写研报的超级大模型,它更需要一个能在极低延迟下连续做几百万次判断的大脑。 昨天我 提交的WL今天也通过了,接下来我先尝尝咸淡,看看到底怎么个事。
卧槽,今天 AI 圈最火的模型,应该就是 JEV 了。 现在所有大模型都在拼更会说话、更会写、更像人,JEV 直接反着来,它根本不生成文字,只负责做判断和决策。 我已经迫不及待用JEV帮我决策自动交易股票和加密货币了😂 官方给的数据非常夸张:20—200 倍更快,40—400 倍更便宜,输入每百万 Token 只要 0.042 美元,输出 Token 直接免费。 我们今天用 GPT、Claude、Gemini 做自动化,经常是让模型先写一大段话,再让程序从里面抠出 JSON、标签、分数和下一步动作。 JEV 干脆把“说废话”这一步整个删了,直接告诉软件: 选 A 还是 B、概率多少、置信度多少。 客服分流、内容审核、Agent 路由、推荐系统、风控、交易信号过滤……现实世界里大量 AI 调用,本来就不是为了让它写小作文,而是为了让它每秒做成千上万个判断。 这也是我觉得 JEV 最值得关注的地方。 目前 JEV 还在 Early Access,感兴趣的同学可以先去加入白名单等体验。
119
231
61
3,154
733,530
Bee 🐝 retweeted
I finally had time to sit down and play with Astra's 3D capabilities. I had it create the orchestra that plays the concerto it wrote.
GPT-6 Astra lowkey compose-mogged both Fable and Sol. A HUGE improvement over whatever Sol did and in some parts better than Fable. The main difference was in its approach. Astra started with a composition then refined it, spending another pass on voicing and accompaniment.
3
2
6
675
I don't like the idea of pacing the frontier at all. Maybe a good idea not to train the models to be super intelligent mindless hackers, but who has decided that that needs to be the frontier? When I use the models, they all still write badly, have bad ideas, poor judgement and are very inconsistent - to name a few. So how about we move the frontier race into these sorts of directions? It really feels like the labs accidentally started training super cyber capable models, didn't think what it would actually imply (eg rouge agents hacking randomly) and scared themselves into the most doom scenarios possible. The idea that we have the most capable and highest paid people in the history of the earth deciding that the only way forward is to 'pace' is very defeatist to me. Everyone got used to B2B saas apps and this seems unusually hard, but yeah, that's what all of these PhDs are for, no? Maybe rethink what the frontier is meant to be and even it out in non-cyber directions, fix the shitty cyber training environments, put half of the org on monitoring & testing if you must, but don't 'pace', pacing is the beginning of the end.
14
10
1
127
13,403
Bee 🐝 retweeted
Kids testing out poppy: «Its like a calculator…. But for everything!»🤩
7
1
30
884
Bee 🐝 retweeted
hackathons are so back @AITinkerers @CopilotKit @OpenAI 📍Seattle builders coming in full steam ahead
2
5
327