Building the future, one robot arm at a time 🤖. Game on!

Boston
Joined June 2022
James Smith retweeted
KKR’s credit work estimates AI-linked debt at about $600 billion, or 6.3% of the US investment-grade market, versus a 2.6% historical average for the largest sector exposure. >$7.6–8tn infrastructure capex through 2030 >AI exposure could approach 20% of IG index >Guarantees, leases and commitments obscure exposure (1.7T per their figures)
2
12
50
4,065
James Smith retweeted
If you've used the same model in different harnesses, you've probably noticed it behaves very differently in each. Our new technical deep dive shows how to fix this by training the model with RL inside the harnesses themselves: huggingface.co/spaces/FineEn… In some cases, the gap is huge: before any training, LFM2.5-2.6B solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code 😢 The obvious fix is to train inside the harness, but this isn't simple because the harness controls the agent loop, so the trainer never sees the exact tokens the model produced. We solved this with: > a capture proxy in OpenEnv that records tokens + logprobs from harnesses > @harborframework for the tasks and sandboxes > TRL's super fast async GRPO trainer Training LFM2.5-2.6B across four harnesses took it from 42% → 54% on held-out tasks, with gains in every harness and 31% fewer tool calls on tasks it already solved 🔥 Training in OpenCode alone mostly improved ... OpenCode 😅 Huge kudos to @adithya_s_k for leading this, and to @QGallouedec @DirhousssiAmine @SergioPaniego @ben_burtenshaw for enabling multi-harness training in TRL & OpenEnv!
3
7
1
47
2,609
James Smith retweeted
👉 how did @elonmusk build xAI data centers so fast? paying far above-market for certain labor; doing portions of the facilities elsewhere and then moving to the site; and... industrial robots!
11
30
4
325
52,358
James Smith retweeted
Why does it feel like such a MASSIVE waste of time to watch Claude deliberate and talk to itself for 20 minutes during its planning phase, but it's perfectly fine to brainstorm, discuss, draft and re-draft a spec (v4.final.2.final.md) for 2 hours in a meeting of 5 people?
62
5
2
90
6,277
Faster fine-tuning on Apple Silicon with Neural Accelerators thanks to these PRs in MLX! 🚀
4
4
94
9,218
James Smith retweeted
Effective Dense Retrieval using Only In-Context Examples @nour_jedidi et al. show that an LLM can build dense retrieval embeddings with no training, by prompting it with a few query-document examples. 📝 arxiv.org/abs/2609.38099 👨🏽‍💻 github.com/nourj98/RICE
1
8
55
2,477
James Smith retweeted
Look at that thing fly. I actually still prefer it (: Opus can focus on research
3
3
94
11,068
James Smith retweeted
SPP: about 13.5 GW of data-center campus capacity is connected or seeking connection, spanning Kansas City to the Texas panhandle. Just six players, Google, IREN, Beale, STACK, Core Scientific and Nebius, account for two-thirds of it. Still trails both PJM and ERCOT.
3
4
1
8
1,383
James Smith retweeted
🎬London Bridge - Chase Scene (AI Short) It's been a while! I've been busy with AI projects at my day job, but I'm hoping to start posting more soon. I'm slowly finding my way back to the AI community, and I've missed sharing things with you all! So, to kick things off, here's a little chase scene 👀 Video: Seedance 2.0 and 2.5 Images: frames for characters taken directly from video output Editing: CapCut
🤖 Made with AI
8
10
1
37
1,741
James Smith retweeted
GPU poor rejoice! You will be able to run Qwen3.8-Next-Flash on a single 24-32 GB GPU + 64GB ddr4/ddr5 + 100GB NVMe/SSD RTX 3090: - 2000 / 2800 tok/s prefill - 64 tok/s decode c1 - 210k fp8 cache Intel Arc B70: - 800 / 1200 prefill tok/s - 35 decode tok/s - 270k fp8 cache
76
31
5
685
168,480
James Smith retweeted
Jensen Huang calls out the AI doomers scaring young people into thinking there won't be any jobs left for them: "Don't think for a second just because you're an alarmist that you're doing a social good."
9
25
8
89
8,979
James Smith retweeted
No One Knows What To Do With Agents & No One Gets Paid Until They Do There’s a massive opportunity for people who can fill the gaps vinvashishta.substack.com/p/…
1
6
1
11
514
There's no reason for me to deceive you
1
5
1
52
5,859
James Smith retweeted
ICYMI @AnthropicAI’s Opus 5.5 (High) ranks #2 in Agent Arena and reshapes the Pareto frontier. Opus 5.5 (high) not only improved upon both Opus 5 variants with a higher net improvement score than either, but does so at at 40–56% lower cost:
Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier. Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max). By signal, Opus 5.5 (High) ranks: - #1 Steerability (+14.50%) - #2 Confirmed Success (+15.50%) - #3 Praise vs Complaint (+19.80%) - #4 Bash Recovery (+10.64%) Congrats to the @AnthropicAI on another frontier model release!
9
10
1
240
27,035
James Smith retweeted
A GNN library built natively on Keras 3 -- with models running on JAX, torch, TF with full hardware acceleration (Apple Silicon, TPU, etc.). "K3-Node achieves 100% public API parity with PyG and incorporates state-of-the-art foundation models and architectures from Spektral and StellarGraph" github.com/anas-rz/k3-node
26
11
2
108
24,727
James Smith retweeted
Full breakdown: Video generation: Kling 4.0 (partner early access) @Kling_ai Character references: GPT-Image 2.5 Sunburst @ChatGPT Sound design: Epidemic Sound @epidemicsound Edit and grade: CapCut @capcutapp
1
1
1
455
James Smith retweeted
I went back to 1970 and robbed a bank. Built the entire heist getaway scene using Kling 4.0 (coming soon). This is what cinematic AI looks like now:
12
9
2
80
5,580
James Smith retweeted
My game is getting there! WIPE: SURVIVAL Written in With and @raysan5's Raylib
7
14
37
45,141