23, !Forbes 30u30. Finding new modes of Human-AI Interaction. Cofounder, CTO/Chief IC @getalchemyst | @ai4bharat alum, automating weekend work for @ritikaadass

lost in cyberspace
Joined July 2021
> Got denied leaves for offcampus by college TPM because the same company came for oncampus as well, and my TPM got to know it. > Was told basically "CGPA dekhi hai teri?" > surprised_pikachu.jpeg > Got shadowbanned in placements > Got a 10+ LPA job from oncampus because recruiters got to know me from a hackathon > Left offer, one of my friends got in my stead, me happy > Got $100k+ USD WFH offer, left that to join another promising startup (WFH) > Punished by being forced to stay an extra sem > TPM made an example out of me by sending a mail (can only DM SS for obvious reasons) > Became part of my batch lore, instant popularity - asked out even on LinkedIn lmao (but I already won in life with my gf) > dab_while_crying_internally.jpeg > handled full time job with extra penalty courses and B Tech Final Year Project > Worked on Data conflicts and hallucinations in LLMs as my BTech thesis > Started building @getalchemyst from college room, with @uttaranxnayak > Resigned from aforementioned startup job > Advisor fac didn't know what I was working on, gave me less marks > Left my job, raised preseed funds > Btech thesis got published in Springer > happy_happy.jpeg > Fac indirectly apologized later on > never_give_up.jpeg
people really have lost their basic sense of manners while speaking to elders, especially teachers. i agree some teachers might not be good, doesn't give me the right to be rude to them, they're still teaching people or have taught people, and people should respect that
19
11
2
549
64,063
Imagine a model at par with GPT Realtime 1.5 or GPT Realtime 2.1 that costs you effectively less than $0.005 per minute for the most complex voice AI use cases (imagine multilingual phone interviews), at <200ms E2E latency p50. How does that sound for voice builders or real-time AI agent requirements? In case you are interested to know more, hit me up.
1
51
> "We're building prod, not God"
I’ve put a ton of tokens through Jev now, and...I have thoughts. 1. The big labs lost the script TypeSafe’s intro video says, “We’re building prod, not God” — and can I just say? Hell yeah. @typesafeai built something for engineers like me, so I can build products for people like your mom (she says hi btw). They’re not trying to create “Machines of Loving Grace”; they’re trying to build things that let OTHER HUMAN BEINGS create new types of products. The proof is in the pudding. When ChatGPT shipped, we were all blown away by…ChatGPT. When Jev shipped, we were all blown away by what everyone was making with it. I cannot tell you how refreshing this is. It takes a sincere form of humility to believe that you, the creator of a technology, will not be the best at deploying the technology into the marketplace. The big labs, especially Anthropic, have proven not to have this humility…and now everyone hates AI. Thanks, Dario. I guess what I’m saying is: Diogo for president. 2. Ultra-smart classifiers were genuinely a missing primitive. I don’t know how this got missed. Diogo calls it “machine-native intelligence,” and that’s really what it is. I cannot tell you how much bending and twisting I've had to do with LLMs to get them to act like a classifier when they just weren’t. I’m absolutely shocked this is the first time this “type” of model has come to market. These models will unlock AI utility for entire industries: robotics, trucking, aviation, meteorology, finance, defense, retail…even your smart coffee mug is going to use this. 3. It’s still too expensive Jev is really, really cheap compared to LLMs. But you don’t use it like an LLM. In my full self-driving example, or any robotics example, you’ll be making decisions several times per second continuously. The #1 value proposition of a model like this is speed and cost. To be clear, at current pricing, the model will still be successful and well integrated. It will be able to replace a number of tasks that were already being performed (poorly) by LLMs. However, to unlock industrial-scale demand (Jevons paradox), I believe it needs to be roughly 10x cheaper. Currently, Jev is $0.042/M input, which is actually 7x more expensive than a cache read on DeepSeek V4.1 Flash ($0.006). If our industrial application requires decisions at 10 Hz and we provide only 10k of context, that would cost $0.0042/second to operate, or $15/hour, $600/week, etc. The good news is Diogo said on a recent AMA that they have substantial margin at their current pricing and have already considered dropping the cost further. I hope it's a lot further. 4. Latency is a real limiter I really hope TypeSafe AI works with an edge network provider to improve the co-location of models and reduce network latency. Ultimately, these models really need to run on-device. I would even be willing to pay some kind of recurring license fee to have a great closed model running on my own hardware, just so I can reduce the TTD (Time To Decision) as much as possible. Imagine running a model like this at 60 Hz, or even 120 Hz. At that point, we’re looking at a new kind of logic gate — we can’t even imagine what that will be like. In summary, this one really is a game changer for everyone building AI products, and, assuming cost and latency are further improved…it’ll be a game changer for everyone doing anything.
3
173
If you want to learn about what we do, check this out!
Well, a week late - but better late than never. Had a blast of a podcast with @sunnyray - my first one as well ^_^ . And, as I always say - The Model was never the bottleneck. Link in replies :)
1
2
190
Hey @Teknium , Hermes CLI is amazing. It'd be perfect if only 1 thing was resolved: /redraw /redraw is often a hit and miss. For SSHed VMs running within VSCode env, doing /redraw in a terminal often worsens the alignment than a proper redraw As a big time user of Hermes, would love if the redraw is fixed.
1
122
Wake up Spend time with parents Deep Work (Cook model recipes, Read papers) -- Breakfast -- Feature Requests/PRs/Daily Updates Streamline team processes -- Lunch -- Attend investor meetings Sometimes handle chud behaviour by clients Reach out to people -- Snacks -- Hit the gym Learn something new -- Dinner -- Spend time with fiancee Sleep
2
124
Context is literally the most anti-fragile sector unlocked with GenAI going mainstream. I am ready to die on this hill.
1
1
5
183
Clanker just killed a 16 day long training run as a part of it's memory optimization process. Speechless.
148
GPT 6 Astra is something else man. It's too addictive. GPT 6 Astra in Plan mode + GLM 5.3 in Build mode = 🤯🤯🤯
1
12
319
First in my bloodline to be required to put ISD codes before names in my contacts. :)
1
2
284
If you're going to go for 33 hackathons, good luck trying to build anything substantial beyond quick dopamine hits
4
242
Am I getting it wrong, or are there too many downsides to embedding scaling - before we even consider that the embedding space between two different models might be totally different?
meet @mostik_ai! what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can. everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning? we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched. how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance. we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst WIRED has the first external account of the company and the work: wired.com/story/russian-star… full writeup, the setup, and all the numbers: mostik.ai/read-more
120
Is the goal of AI memory to mimic human memory, or be better than it? Sounds like a real moonshot problem to me.
1
2
122
anuran 🛠️ Alchemyst AI(e/acc) retweeted
Replying to @AbhinavXJ
I used to break games for fun, before Denuvo existed. Windows XP SP2 and Cheat Engine. This was ~2008. Ravenshield, Max Payne 2 were what CSGO and COD are today. Back then Win XP SP2 was the last truly "hackable" Windows. Then I got introduced to Kali Linux, did a bunch of things with real life consequences that I'm not very proud of Started learning AI back when Andrew NG's Stanford ML course was the only popular and reliable course online. Made a new motherboard out of a fried one by soldering wires (although it was too darn slow and sparks used to happen) Then JEE happened. Good lord I feel like an unc while typing this 😭😭😭
1
3
190
Can't remember where I found it though
GOOGLE HAS BEEN FAILING PRETTY BADLY TO PUT OUT A MODEL THAT CAN ACTUALLY HANG WITH THE BEST LATELY tested Gemini 3.8 Flash against Kimi K3 on the same Sticky Ball game Flash was insanely fast and used way fewer tokens, but the actual game was nowhere close. K3 had much better mechanics, movement, and overall game design Flash clearly has the speed and efficiency part down, but the gap in what it can actually build is pretty big
1
1
4
877
14BC
so when did you start coding? answer in BC = Before ChatGpt/Claude me 0 BC
1
4
779
Pretty consistent analogy with what we've found when it comes to Gemini 3.7 Flash. Brother G gets the task done, but with random ahh checkpoints. So it gets ABSO-FKN-LUTELY mogged by smaller models like Qwen 3.8 27B when it comes to checkpoint or sequence based work.
We just evaluated Gemini-3.8-Flash on our @PhyseraAI - TB bench. Our analysis on 5 random tasks from the set: 1. It is excellent at deriving and implementing a coherent numerical method. It turned an internally consistent mathematical model into working code, especially when it can derive a checkable optimum. 2. It is inconsistent on edge-case semantics. It knows the individual primitives, but assemble them in the wrong order when a specification has interacting edge cases. 3. It tends to overbuild static-analysis solutions while missing the hardest coverage cases. 4. In multiple tasks long trajectories do not imply better outcomes. More exploration became speculative scope expansion rather than targeted verification. I think it is promising low-cost choice for numerical/scientific coding / transformations with crisp formulas / tasks where it can independently check residuals or invariants but had fallbacks for production shell tooling / static analysis / clinical derivations / compliance-style work.
3
136
anuran 🛠️ Alchemyst AI(e/acc) retweeted
unpopular, optimistic perspective: people are mostly awesome, life is very beautiful, and the future wants you to be bullish on yourself
17
66
5
871
143,831
In comparison to this, the best US-origin model is Muse Glimmer 30B by @AIatMeta , which scores ~1/4th of DeepSeek V4 Flash 0731. We also have a new winner in town - Ling 3.0 Flash by @TheInclusionAI - at ~2x better than DeepSeek V4 Flash 0731. Criteria remains the same: Intelligence per dollar. Harness: OpenCode.
Deepseek V4 Flash 0731 is dollar-for-dollar ~3x better than GLM 5.3 Flash. - It shows ~33% better alignment towards changing objectives - It produces ~17% more output tokens over GLM 5.3 Flash - GLM Flash performs ~37% better for STRICTLY coding tasks. Tested with @opencode.
1
3
335
anuran 🛠️ Alchemyst AI(e/acc) retweeted
If you thought Calcutta doesn’t have a good tech scene, you were right 3 months ago. Not today. My friend Sahil and I recently started Calcutta AI Club, and yesterday we hosted the city’s largest ever Vibecoder’s Hackathon. 60 participants, more than half of them never having coded before in their life. From teenagers to middle-aged corporates, a variety of people took part. Instructions were simple - in 6 hours, build a working app with real world usecases. From AI-integrated educational apps to industrial ERPs, to a PCOS tracking app and even a recipe sharing app. The results were super creative and too much fun to watch! And you know the best part? The winning team was carried by a 13 year old kid! Seriously a child prodigy. Grateful to our judges for giving us their valuable time - @abhishekrungta @RMantri @priyankarungta Thank you for coming! And of course thank you to our sponsors. None of this would be possible without your help! @CuePilot_AI @getalchemyst @indusnettech @AnuranBuilds The tech landscape in this city is going to change very soon. The talent was always there, what lacked was a platform to showcase it. We’re just getting started. Our goal is to turn Calcutta into a major global tech hub. @swapan55 @paulagnimitra1 @SuvenduWB we will need your help with that! Miles to go before I sleep. Follow us on instagram at instagram.com/calcutta.aiclu…
8
16
2
38
2,517