Postdoc @MSFTResearch | PhD @uwaterloo | RA @UKPLab l working on search and agents | (https://nitter.cf/t.co/bxTCuxkJWZ, https://nitter.cf/t.co/KYPd6PHDYd, @TREC_RAG and FreshStack)
Bengaluru, India
Joined July 2016
- Tweets1.9K
- Following3.5K
- Followers3.4K
- Likes19.1K
Pinned Tweet
Life Update: Last month I started as a Postdoctoral Researcher at Microsoft Research India!
If you are curious on why I chose to return to India after my PhD from Canada; I've written some personal thoughts about it!
Nandan Thakur retweeted
Embedding model evaluations has many issues, so that sadly benchmarks like BEIR & MTEB became low signal.
With RCP-nDCG we are here to fix some of these:
cohere.com/blog/rcp-ndcg
arxiv.org/abs/2609.35739
Nandan Thakur retweeted
Karpathy: disappears from X
Ben Affleck: alright, gather round, so you'll want to freeze the base weights first, learning rate 2e-4
Ben Affleck (Hollywood star & Artists Equity CEO) talks about how he fine-tunes open video models by unfreezing weights and trained only the last cinematic layer so a film crew can hit real production standards.
for context, Ben Affleck founded InterPositive in 2022, a 16-person AI shop for film post and Netflix bought it in March 2026 for $587 mn in cash.
He needed that model because public video models were trained on his peers' films, and he did not think that was a real business. So InterPositive raised money, shot its own dataset for 8 months on a controlled stage, and used it only as late-stage training.
Each new film then trains a private model on its own dailies, so the production keeps the footage and the learning. That is the product Netflix paid $587 million for.
----
From "Bloomberg Live" YouTube channel, (link in comment)
ArXiv: 2 submissions per month.
Claude: Understood. Pushing 47 slop papers to each GitHub codebase now.
arXiv has updated our policy on rate limiting for all submitters.
This update was made to fairly distribute moderator time & support the arXiv community of staff, volunteers, readers & authors.
Please read our announcement to learn more: blog.arxiv.org/2026/10/01/up…
Nandan Thakur retweeted
Excited to share that I have joined @MistralAI, where I will be building a team in Montreal and across Canada to push the frontier of safe, capable, and open LLMs.
Mistral has played, and will continue to play, an important role in keeping frontier AI open and advancing it safely. I deeply believe in Mistral's mission, and since joining, that conviction has only grown after seeing firsthand the ambition and commitment of the people here.
A big part of my role will also be building strong bridges between Mistral and @McGillU, @Mila_Quebec, and the broader Montreal and Canadian AI ecosystem.
And we are hiring! Across levels and across the LLM stack. If you are excited about pushing the frontier, reach out.
Nandan Thakur retweeted
in the past months we worked closely with tpuf team, iterating our models beyond web search. During the journey, we built a new way of training contextual embedding models, which connects our previous work, the context pruner.
We're sharing how we train the model, and how it performs on @turbopuffer 's private bench.
great work from @ESL_Sarah , Markus, @antoine_chaffin , Louis, Max and special thanks to @n0riskn0r3ward , who helps evaluate, iterate again and again 🐐.
We'll work closely together to continue improve it, and eventually offer it to all.
We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view.
pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench.
perplexity.ai/hub/blog/conte…
Nandan Thakur retweeted
The Cohere Embeddings team has been cooking 🍳
Several key innovations in the model:
- The model not only retrieves relevant results, but finds the most human preferred documents based on new training & evaluation techniques.
- Strong focus on multilinguality 🇺🇳 - just retrieving English docs is boring
- Massive throughput gains - scale to any data size
Together with a strong focus on privacy - your data stays private, always.
Nandan Thakur retweeted
📚 Fantastic #ANRFAmbassador paper reading @MSFTResearch #Bengaluru 30+ participants across UG, Master’s, PhD, faculty & industry. 🔥 Great presentations + questions! Thanks @KodaliPrashant & Microsoft Research for hosting! 🚀 #ANRFSabbatical #ProfGiri /c @ANRFIndia @shivkuma_k
Nandan Thakur retweeted
📣 interwhen an @MSFTResearch India project that I had fun being involved with, is going to #NeurIPS2026!! (arxiv.org/abs/2602.11202). Here is a thread from March about viewing what interwhen does as LLM-Process-Modulo.. – at Tempe, AZ
LLM-Modulo (Solution vs. Process): Our original LLM-Modulo framework (arxiv.org/abs/2402.01817) is a Generate-Test framework, with the LLM generating candidate solutions and a bank of verifiers critiquing those solutions.
1/
nitter.cf/rao2z/status/188717226…
Reminds of FrugalGPT....
"JEV-as-a-Judge: Accept When Confident, Escalate When Unsure"
This paper shows you can just use JEV for every evaluation instead of expensive LLM.
JEV basically acts as a cheap first-pass judge, returning both a verdict and how confident it is.
When confidence is high, keep the answer. When it’s low, escalate to a stronger LLM.
This simple routing keeps ~99% of GPT-6’s accuracy while reducing evaluation cost by a lot.
alphaxiv.org/abs/2609.26550
Nandan Thakur retweeted
And now… accepted to #NeurIPS 2026! 🎉
Our paper got an oral at the CTB workshop @ ICML🇰🇷
How stable are LLM leaderboards? 🤔
Less than you'd hope. We built a unified framework for auditing how small changes to pairwise votes ripple through leaderboard rankings
📄 arxiv.org/pdf/2605.15761
🌐 hosnahoseini.github.io/leade…
Nandan Thakur retweeted
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions.
CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks.
With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks.
We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed.
Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size.
📄 Blog: contrastive-lm.notion.site
💻 Code: github.com/Contrastive-LM/CL…
🗣️ Discord: discord.gg/5dAQEDJBs
🤗 Data & Models: huggingface.co/Contrastive-L…
More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
Nandan Thakur retweeted
alright guys I'll have the pleasure this week to interview @mrdrozdov from databricks on all thing retrievals like:
- RL for knowledge agents
- learning representation program for retrieval
- place of harness in all that
if you have any questions on retrieval shoot it my way
Nandan Thakur retweeted
There’s a reason I stayed up all night on my birthday writing that article (nitter.cf/furongh/status/2101913…): I’ll be serving as an ICML 2027 Program Chair, alongside Weijie Su, Andrew Wilson, Julien Mairal, and Courtney Paquette. As someone who has worked on AI detection and will help run the conference, I feel a real responsibility to think more deeply about “AI slop” and what we should actually do about it.
With reports of 60K+ abstract submissions to ICLR 2027 (linkedin.com/posts/gdemelo_o…), how do we scale peer review without lowering the bar or burning people out? Recruiting enough reviewers is one challenge. Matching papers to the right expertise is another.
Here’s where I stand.
Using AI to improve your writing should not be punished. This matters especially for non-native English speakers. If AI contributed meaningfully to your ideas, experimental design, or analysis, be transparent about it. Either way, you are responsible for understanding, checking, and standing behind the work.
But **we should not ask reviewers to spend hours on a paper its own authors haven’t read, checked, or understood**. That standard should apply whether AI was involved or not. Submitting a paper is also a request for someone else’s time.
And “just run it through an AI detector” is not the answer. Research has documented false positives, including for non-native English writing, and ways to evade detection. A detector score is not a verdict on scientific quality. (arxiv.org/abs/2304.02819) Nor should confidential submissions be uploaded to external services without authorization and appropriate privacy safeguards. (icml.cc/Conferences/2026/Rev…)
I think desk rejection deserves serious consideration, but it needs to be based on demonstrable problems, not “this sounds AI-written.” What can we reliably identify before full review? Who makes those calls, and how do we catch mistakes? We need to protect reviewer time without screening out unconventional work just because it is unfamiliar.
**I’d love concrete suggestions on reviewer matching, early triage, and author accountability.** What have other venues tried? What worked and what backfired? How do we discourage careless submissions without making it harder for newcomers and less-resourced researchers to participate?
These are my personal starting views, not an announcement of ICML policy. We’ll do our best to serve the community, and your input would genuinely help.
#ICML2027
Nandan Thakur retweeted
Those who are wondering, I am Nandakishor M, the creator of Laya, now competing against Jev. Just a simple guy from a small village in Kerala just teaching generative AI to everyone through instagram!!!
instagram.com/nandakishor_m/
This is one the biggest reasons, we should try to actively relabel existing retrieval evaluation datasets using LLMs/and create better and newer ones.
Yes, humans can do mistakes and LLMs as well. Asking the LLM to review misalignments in existing evaluation datasets should be done.
BTW, i always recommend this kind of analysis, eyeballing pairs.
Nandan Thakur retweeted
For IR folks: @xueguang_ma and I quickly tested Jev as a pointwise/pairwise/setwise/listwise reranker on the classic DL19 and DL20, reranking the top 100 BM25 results. It is pretty good and cheap.
Nandan Thakur retweeted
As per Pangram, 100% of the @IndianExpress article by @Pvsindhu1 is generated by AI.
pangram.com/history/310d0add…
I’ve always loved writing. Sometimes, it’s the easiest way to say what you feel. This year, Modi ji, I wanted your birthday tribute to come straight from the heart. ❤️
I’ve seen changes in Indian sport I could barely have imagined growing up. The support, the opportunities and the belief have meant so much. Across fields, a generation of Indians has learnt to take on the world’s biggest names and believe we belong alongside them.
But it’s the personal moments I treasure. Sharing an ice cream with you. Watching you laugh with my then coach, Park, and knowing he would remember it for the rest of his life. You have an innate ability to make people feel warm, welcome and loved.
I remember your hour-long conversation with my husband, Datta, about identifying unique citizen records, or “golden records” at 10:30 in the night. Back in the car, he couldn’t stop wondering how the Prime Minister knew so much about such a deeply technical subject. Your curiosity stayed with him as much as your warmth.
Happy birthday to our beloved Prime Minister, Shri Narendra Damodardas Modi ji.
With a daughter’s affection, I wish you good health, happiness and many more years of service to our great country and its people. Thank you for always making us feel that our dreams matter to you 🙏❤️
@narendramodi @PMOIndia
Nandan Thakur retweeted
Introducing SPARSEUP: the first model release from @Linkup_platform , and the missing sparse companion of DenseOn and LateOn.
Same backbone, same data, <150M, 56+ on BEIR-13. The model is open-source, use it!
Blog: linkup.so/blog/introducing-s…
Model: huggingface.co/Linkup-Platfo… (Apache 2.0)
Nandan Thakur retweeted
Nihar Shah did a heroic experiment for TMLR: he spent 20-25 hours over two weeks interviewing authors of seemingly low-quality submissions about their own papers.
He confirmed what we all suspected: people submitting these papers have *no idea* what is going on in them.
TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity
Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission?
medium.com/@TmlrOrg/asking-a…
Nandan Thakur retweeted
Riddles built backwards turn out to be excellent training data for search agents. 🛰️ Each question in ORBIT starts from a short verifiable answer and wraps it in four or five clues that each narrow the field, and solving one properly means verifying every clue, one search at a time. Nandan Thakur (@beirmug) built the 20K-question dataset with colleagues at the University of Waterloo, then had external search agents re-verify the answers, because a training set with wrong ground truth teaches the wrong lessons.
The rest of the conversation ranges across deep research harness design, context compaction and memory, and sequential versus parallel search trajectories in GRPO training. Weaviate Podcast #137:
youtube.com/watch?v=B71WF6Et…