Data & Evals @cohere | RL research @ NUS | MSCS @UMassAmherst | prev Google | built @CurationsClub ✨
Joined May 2017
- Tweets1.8K
- Following989
- Followers680
- Likes3.5K
Pinned Tweet
Back for my third semester, and feeling very grateful for the summer I get to look back on.
some highlights:
- gave my first poster session for a course project
- 2 became COLM workshop papers. SF, here I come! yayy
I’ll be in SF for @COLM_conf! If you work on RL, multi-agent systems, AI safety, or alignment, I’d love to chat. Feel free to reach out :)
Armaan @COLM 2026 retweeted
Back for my third semester, and feeling very grateful for the summer I get to look back on.
some highlights:
- gave my first poster session for a course project
- 2 became COLM workshop papers. SF, here I come! yayy
well..
Looks like we have over 500 submissions to ICLR this year.
openreview.net/group?id=ICLR…
exciting last few months at sarvam, lot of things i have been working on are now released ✨
1. sarvam-105b-conversations - is specifically post-trained to be robust in noisy ambiguous voice scenarios with deep nested instructions and scenarios...
Sarvam Epoch brought builders, digital-native teams, and enterprise leaders together in Bengaluru for two days of conversations, demos, and launches across the stack.
We introduced new models, inference infrastructure, agents, and products, alongside partnerships that will help bring these capabilities to more people.
Thank you to everyone who showed up, participated, contributed, and made the first edition of Epoch what it was.
Here’s everything we announced: sarvam.ai/epoch/summary
two acceptances at @COLM_conf 2026 workshops 😭🎉
“Measuring Reward Hacking and Reasoning–Answer Decoupling Under Position-Confounded Optimization” at AIMS + “Targeting the Attention Heads Behind Object Hallucination in LLaVA” at @ActInterp.
see y’all in SF. more deets soon!!
i thought i knew how policy gradient algos worked.
then one hierarchical architecture knocked me out -_-
for context, i’m working with:
s/a space → net1 → net2 → loss
both nets contribute to log-prob, & gradients flow across setup. any good resources to understand this?
Armaan @COLM 2026 retweeted
I felt a great disturbance in the Force, as if millions of voices suddenly cried out in terror and were suddenly silenced.
I fear something terrible has happened.
spent last night reading through this and found it super useful.
honestly reassuring to know that the job search is challenging for a lot of people, not just me. the resources are really well-collated and saved me a lot of time and energy finding them myself.
highly recommend.
I'm joining OpenAI next week!🥹 The job search turned out to be really challenging but also super rewarding, so I wrote a small blog to share what I learned along the way and hopefully make the process a little less mysterious for the next person. alisawuffles.github.io/blog/…
Armaan @COLM 2026 retweeted
Subliminal learning is when LLMs transmit traits (e.g. loving cats) through seemingly meaningless data. What’s going on?
We find a simple explanation: it's just steering vector distillation.
We explain which traits transfer and why subliminal learning fails across models.
building this taste layer right now with curations. if you're working on preference infrastructure, semantic retrieval for taste, or anything adjacent, let's compare notes
one version of the future: my agent pays your agent because yours has taste in some weird little corner of reality mine never explored. content goes to zero, while judgment gets expensive. nobody browses anymore. you just wake up one layer above the web, receiving whatever survived the little economy of machines trading, negotiating, and arguing beneath you.
went a bit MIA during the second half of the semester.
spent most of that time heads-down on research, building and running 3 grad projects. 2 of them are now being shaped into workshop submissions for @COLM_conf.
learned more in the last few months than i expected to <3