@danielbjanes

Applied AI at @tryramp building fully agentic code reviews & hanging out with robots

New York
Joined October 2012
Jev’s production capabilities are obvious, but turns out it’s pretty bad at blackjack… @typesafeai
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
3
207
The fun problem to solve after this is how to do this at scale
The most important part of PR reviews right now is how quickly you can recover from a mistake caused by the PR. If a PR were to cause an issue that I can fix within a handful of minutes for all users, great, let’s ship it. If it involves persistent data changes, irreversible migrations, data going to a third party dependency that you can’t easily influence, or DNS, let’s be careful and do it properly. Your goal is to get the time it takes for a PR to safely be deployed to all users as close to 0 as possible.
1
90
I actually think this introduces a potential failure mode to producing software. Reviews have served as a natural place to discuss architectural designs and the quality of output. As these take place less, poor designs are now found tens of thousands LOC later.
One very interesting thing about code reviews starting to fade in many places: I still remember how, as a manager, code reviews would be one of the places where conflict between individuals or teams would be very EASY to spot for me, as a manager (then figure out how to deal with it). What happens when there's not much of this code review happening? Does it mean there's no longer any conflict (or interaction?) between devs on the same team, or between teams? As a broader question: are AI agents drastically changing team dynamics? Or just hiding them? They sure feel they are reducing human-to-human interaction via the code, I'll be honest
74
Daniel Bjånes retweeted
Wherever I go, I see @tryramp
5
2
50
4,270
Daniel Bjånes retweeted
The next generation of Inspect is under way. The internal excitement and bullishness is palpable. Time to win.
6
4
1
84
7,166
>ReviewBuddy: Ramp’s own code review system 👀👀👀
It's rare to see a non-AI frontier lab build an in-house AI coding tool that works so much better than what frontier labs have to offer. Ramp did this with Inspect: an AI coding harness fully integrated with their systems, running on cloud machines, and being a smash hit inside the company. Ramp basically runs full-blown cloud development environments (CDEs) for agents to build code on, and validate it, and has made this tool "multi-player" by default, with no opt out. We find it likely that the likes of Codex, Claude Code and GitHub Copilot will build similar tooling in a few months' time to what Ramp already has in-place. Deepdive: newsletter.pragmaticengineer…
3
124
Thanks for the goodies 🤝 @claudeai @bcherny
1
4
263
Daniel Bjånes retweeted
Ramp, the AI lab that does credit card stuff on the side too
9
10
1
174
20,654
Daniel Bjånes retweeted
Ramp is a "save you time and money" company router.com is our first step to go beyond fintech DM me if this is exciting to you- we're actively working on money saving products across the AI stack (router, harness, application layer)
Replying to @HesomParhizkar
the request-level evidence is the interesting part: Default baseline: $0.001219 Flex actual: $0.000610 same model, roughly half the cost. this is the first test that gave me a concrete reason to use Router instead of calling Luna directly. Nice work, @tryramp , @vral and team! next thing I’d love: aggregate baseline vs actual spend, flex share, fallback count, and why each tier was selected.
4
4
1
104
12,845
Wild how much can be built of only .md files and sub-agents, even crazier that there’s an assumption that its good enough
2
112
Daniel Bjånes retweeted
Monitor and control your AI spend on every provider on router.com. Our early users save 40% on average. Every week, the price-intelligence-latency frontier shifts, and we expect this trend to continue. Tradeoffs between latency, reasoning, cost, service tier, open source and closed source models are shifting constantly. Router sends every request to the model that's actually best for the task and helps you control what tokens you buy. We benchmark it against real work: ~40% lower cost for the same outputs. Today we're opening it to everyone. Two lines of code or just change your base URL. No @tryramp account needed. Free through 2026, first $26 on us. Get an API key today at router.com
148
109
94
1,391
1,423,531
if Orpheus had games on his phone it wouldn’t have gone down like that
75
if only there was a Buddy that could help out on the Review🤔
can I get a review please
5
1,603
Daniel Bjånes retweeted
We’re open-sourcing PorTAL, our framework for shared task representations and cross model LoRA adaptation. It now spans from hybrid attention models to multimodal systems including Gemma 4 E2B, Mistral 7B & @thinkymachines' Inkling. Code: ramp-public/portallib Models: @huggingface /RampPublic
24
44
17
551
171,859
selfie from the "inspectoverse"
1
1
84
Daniel Bjånes retweeted
Every week, the price-intelligence-latency frontier shifts, and we expect this trend to continue Across 100+ use cases in our product, keeping all up to date with the right model is a challenge - Either we're losing out on intelligence for the dollars we spend, or we're spending too much money for the intelligence the feature needs Ramp Router lets you benefit immediately without rewriting your application. We cut our LLM costs by 30%, while also making our features smarter and faster. We built it for Ramp. Now we’re opening it up to everyone. get access here: ramp.com/router?utm_source=X…
106
71
142
1,028
961,134
Daniel Bjånes retweeted
Ramp Inspect hit its 1,000,000th session this week 🚀
17
11
4
140
25,771