@Mahi4AIi
iAccount based inIndia
About this account
- Account based in
- India
- Connected via
- India Android App
Account-level information from X, not a live location or the device used for a specific post.
Machine Learning Engineer || GenAI || LLM || Distillation || Routing
Banglore
Joined May 2020
- Tweets26
- Following89
- Followers3
- Likes271
Mahesh Deshmukhe retweeted
A harnessed LLM agent, clearly explained!
Two agents can run the same model on the same task and finish as expected. But one of them can spend nearly 3x the tokens to complete the task.
The extra usage originates from the code wrapped around them, which decides what reaches the model's context on each call and how many calls there are.
For instance, consider a tool that returned 50k tokens of JSON at some step. If it stays in the context, the model will continue to read that payload again at every subsequent step.
Tool definitions behave the same way.
A server can expose 50 tools, each with a name, a description, and an input and output schema.
By default, all of them will stay in the prompt from the first call, whether the agent uses them or not.
However, an optimally built harness can avoid that unnecessary cognitive load on the model.
More specifically, one core design principle of harness engineering is to push things out of the model at the right time:
- Memory holds the state that weights and context shouldn't carry.
- Skills hold procedural knowledge. These cover the operating procedures and heuristics that specialize a general model.
- Protocols hold the interaction contracts for users, other agents, and tools.
Do note that the context never disappears permanently.
It is always loaded when needed, and the harness decides how much is loaded and when.
For instance, to manage a 50k token payload, a harness can write it to a file and keep a preview and a path in context, hand the work to a subagent whose context is discarded afterwards, or summarize the older messages once the conversation passes a threshold.
If you want to see this in practice, TrueForge is an open-source harness that already implements these practices.
Tool schemas are deferred unless preloading is switched on, large responses go to a sandbox file, and generated code calls tools back through the harness, so the sandbox never holds the credentials.
The two agents I talked about at the top are from DevRev's Enterprise-Bench. TrueForge solved the same number of tasks as Claude Managed Agents on the same model, using a bit over a third of the tokens and around 40% fewer tool calls.
Here's the GitHub repo: github.com/truefoundry/truef…
(don't forget to star it ⭐ )
I also wrote a full breakdown of where agent tokens actually go inside a run, covering context accounting, the strategies above, and the benchmark in detail, and TrueForge worked with me to put this together.
Read it below.
Mahesh Deshmukhe retweeted
Got tired of watching my agents work through a chat interface, so I gave them an office instead!
Now I can watch them run around doing tasks, and when they’re done, they drop their work in my mailbox
This visual interface is a wrapper around @OpenAI @ChatGPT Desktop App.
They're the same codex agents...but now you can actually see them work in real-time!
Andrej Karpathy just rewrote the rules of using LLMs:
"Prompting is going away. Delete everything, keep Graph."
LLMs → Prompts → Agents → Graphs
He dropped his full 2-hour course from Stanford on "Graph-Native Research"
• 00:00 - Intro to Graph systems
• 01:08:09 - LLMs architecture
This Karpathy course can replace a $100K Yale LLM senior degree.
Watch it today, then save the full graph engineering guide below
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Mahesh Deshmukhe retweeted
🚨 A junior at Jane Street reportedly landed a $220K–$600K role because he used AI to analyze trillions of data points faster than most teams ever could.
In this 1-hour lecture, he breaks down the exact system behind it:
• how he researches massive datasets
• how AI finds patterns humans miss
• how his machine turns raw data into decisions
• how you can apply the same thinking yourself
Skip Netflix tonight.
Watch this instead.
One hour could completely change how you think about research, AI, and opportunity.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Readers added context they thought people might want to know
The video is a 2024 talk by Horace He at Jane Street as a guest speaker. He is not an employee there and did not land a role using AI for data analysis. The talk is about ML systems and infrastructure.
janestreet.com/tech-talks/bui…
horace.io
Mahesh Deshmukhe retweeted
Stop wasting hours trying to learn AI.
One list.
Zero confusion.
No fluff.
I’ve already done the hard work for you 👇
📄 Complete AI Learning Document
docs.google.com/document/u/0…
What’s inside:
📹 Videos
LLMs, Agentic AI, real-world breakdowns (Stanford + more)
🗂️ GitHub Repos
GenAI agents, prompt engineering, hands-on LLMs, beginner → advanced
🗺️ Guides & Whitepapers
Google, Anthropic, practical agent design
🧑🏫 Courses
Hugging Face, MCP, Vector DBs, end-to-end agent systems
📚 Books
From fundamentals to LLM engineering
📜 Research Papers
ReAct, Generative Agents, Toolformer, and more
📩 Newsletters
Stay updated without doomscrolling
Everything is curated, sequenced, and practical.
No random bookmarks. No hype.
♻️ Repost for your network
❤️ Like · 🔖 Save
➕ Follow for more on AI Agents & real-world GenAI
Mahesh Deshmukhe retweeted
Kai-Fu Lee (founder of Sinovation Ventures) explains how the future is all about multi-agent systems.
1 agent today is like a pre-internet PC, useful but isolated. Connect agents, and they share context, split tasks, and coordinate instantly.
Mahesh Deshmukhe retweeted
Scaling LLM context windows to millions of tokens usually leads to reasoning degradation.
This paper introduces Recursive Language Models to process prompts up to 10M tokens by treating them as an external symbolic environment.
Modern frontier models suffer from context rot. Even when a model possesses a large physical context window, its ability to reason over that information drops significantly as the prompt grows.
Existing solutions like context compaction or summarization attempt to compress the input. These methods often fail on complex tasks because they discard fine-grained details that might be necessary for the final answer.
Standard LLMs attempt to ingest the entire prompt into their neural attention mechanism. This forces the model to process every token simultaneously, which leads to noise and memory limitations.
Recursive Language Models (RLMs) move the prompt out of the model and into a symbolic workspace. The model interacts with the prompt through a programming environment rather than direct token ingestion.
The RLM process follows a specific execution loop:
→ The input prompt is loaded as a variable into a Python environment.
→ The model writes code to search, slice, or filter the prompt.
→ The model programmatically creates sub-tasks based on what it finds.
→ It recursively calls versions of itself to solve these smaller snippets.
→ The results are stored as variables and stitched together to form a final response.
This method is distinct from prior agentic scaffolds. While previous work focused on breaking down the steps of a task, RLMs allow the input data itself to scale by offloading it to a variable in memory.
By treating the prompt as an object to be manipulated, the model can navigate 10M+ tokens without ever having to "read" the entire document into its own context window at once.
The primary risk is trajectory variance. Because the model chooses its own path through the data, it can occasionally perform redundant sub-calls or get stuck in verification loops, which increases runtime.
The authors address this by using a system prompt that encourages efficient chunking. They also find that using a smaller, cheaper model for recursive sub-calls maintains a high performance-to-cost ratio.
The researchers tested RLMs using GPT-5 and Qwen3-Coder across four difficult long-context benchmarks. The results show significant improvements over base models and standard retrieval methods:
→ RLMs handled inputs two orders of magnitude larger than model context limits.
→ 91.3% accuracy on BrowseComp-Plus compared to 0% for the base model.
→ 58% F1 score on complex pairwise reasoning tasks where base models failed.
✓ Effective scaling to 10M+ tokens
✓ Reduced context rot
✓ Comparable or lower inference costs
This research suggests that the path to infinite context windows may not be through architectural changes alone.
Instead, it demonstrates that context management can be treated as an inference-time reasoning task.
RLMs prove that models can handle professional-scale datasets by using symbolic tools to manage their own attention.
This shifts the focus from building larger neural memories to developing models that can programmatically navigate external information.
Mahesh Deshmukhe retweeted
This paper shows how to train AI agents to reliably use real tools, fix their own mistakes, and finish long tasks instead of stopping early.
They prove that when you train agents inside real sandboxes and reward whole actions instead of tiny text tokens, a smaller open model can perform like much larger models on real world tasks.
ROME is an open agent model trained on tool runs, so it keeps going when tasks get messy.
Their Agentic Learning Ecosystem, called ALE, combines ROCK for locked down sandboxes, ROLL for training, and iFlow CLI for packing the right context each step.
They generate more than 1mn trajectories, full logs of actions and observations, and train ROME to plan, act, check, and retry inside tools.
Their IPA method uses reinforcement learning, training by trial and scoring, and it rewards whole interaction chunks so long tasks do not fall apart.
They test on terminal and repository bug fix tasks and add Terminal Bench Pro, a bigger command line test set that tries to avoid leaked answers.
ROME reaches 57.40% on SWE bench Verified, a GitHub bug report fixing test, which shows open stacks can rival much larger models.
----
Paper Link – arxiv. org/abs/2512.24873
Paper Title: "Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem"
Mahesh Deshmukhe retweeted
New Tencent paper shows how a 1.96B language model trained from scratch can plan, reason, and use tools like an agent.
The big deal is that the paper claims small models can be taught to behave like an “agent” during pre-training, meaning the model learns to break a task into steps, call tools, track state over a long workflow, and self-correct, instead of only sounding helpful after instruction tuning.
It can read a 128K context, meaning a lot of input text at once, by using Multi-Latent Attention to shrink look back memory.
The problem it targets is that small models often do fine for 1 answer, then lose the thread once tasks get long.
Instead of copying a bigger model, the training shifts from everyday text to math and code, then to agent trajectories.
Each trajectory is written like a workflow, it breaks thinking into analysis, plan, action, self check, and summary.
The team generates these workflows for math solving, code fixing in real repositories, deep research with search tools, and tool calling.
Agentic mid-training means the model is trained on those full workflows, so planning and error fixing become normal behavior.
On SWE-Bench-Verified, a GitHub bug-fixing test, adding this agentic mid-training moves success rate from 12.4% to 17.7%.
----
Paper Link – arxiv. org/abs/2512.24618
Paper Title: "Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight LLMs"
Mahesh Deshmukhe retweeted
New Tsinghua University paper builds a new test for web search agents and shows they still fail on vague queries.
Needle in the Web is a new benchmark that exposes how badly search agents handle vague web queries.
It proposes a test where an agent gets 3 fuzzy clues from a real article and must find the exact webpage.
Across 663 queries from 7 sites, most systems score under 35% accuracy.
Instead of asking for a fact, it targets fuzzy exploratory search, where someone gives vague hints and wants the right webpage.
In Needle in the Web, each question gives 3 masked statements from a real page, and the agent must find the single page that matches all 3.
Difficulty is controlled by picking statements that are central to the article or more side details, which makes matching harder.
To grade answers, they use an LLM, the text model behind chatbots, to read the returned page and check whether all 3 ideas are there, without mixing facts across pages.
The results show that models often chase snippets, miss a constraint, or fail to read full pages, especially on everyday sites.
----
Paper Link – arxiv. org/abs/2512.16553
Paper Title: "Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild"
Mahesh Deshmukhe retweeted
Hands-On Guide to Building AI Agents!
This 400+ page guide is one of the most comprehensive resources for anyone serious about building real-world systems with agents.
It covers everything from basic patterns to complex multi-agent architectures with practical implementation details.
Here's what it covers:
• Prompt chaining, routing, parallelization, reflection, planning, and tool use
• Multi-agent design, memory management, learning and adaptation, and MCP
• Goal setting, exception handling, human-in-the-loop, and retrieval (RAG)
• Inter-agent communication, resource optimization, reasoning techniques, and safety patterns
• Evaluation, monitoring, prioritization, exploration, and discovery
• Advanced prompting, frameworks overview, coding agents, and under-the-hood reasoning engines
424 pages of practical patterns, code examples, and design insights.
Mahesh Deshmukhe retweeted
The full stack pipeline of HY World 1.5 that turns text or an image into a controllable, real time 3D world.
Data is first filtered, rebalanced, and annotated so the model learns both visuals and control from clean, diverse sources.
Training happens in stages, with pre training, middle training, reinforcement learning post training, and distillation to boost quality and responsiveness.
At run time the prompt is prepared, then a streaming diffusion transformer denoises short latent video chunks continuously.
A streaming VAE decodes those latents into frames so the video keeps flowing without pauses.
An optional 3D or 4D reconstruction module converts the stream into a spatial world that can be navigated and extended.
The core loop is autoregressive, the model predicts the next chunk while conditioning on user input.
Dual action control combines discrete keys with continuous camera motion so walking and looking map cleanly to generation.
A memory cache plus context reconstitution rebuilds the right history before each step, which preserves long term geometric consistency.
Continuous decoder updates keep latency low and sustain around 24 FPS.
Mahesh Deshmukhe retweeted
New Nvidia paper shows how to design small language models that are genuinely fast on real hardware.
Nemotron-Flash models beat well known small baselines while roughly halving latency and massively increasing throughput.
Most small models just cut parameter count and stack thin layers, but devices still feel slow.
Their experiments show that deep narrow networks use more latency to reach a given accuracy than balanced ones.
They extend scaling rules so loss depends on depth and width, then choose whichever shape hits a latency target with the lowest loss.
Next they benchmark operators like attention, Mamba2, and DeltaNet and use an evolutionary search, a mutate and select loop inspired by biology, to pick a hybrid ordering that fits the latency budget best.
For training, they keep each weight matrix at fixed norm after updates and prepend special learnable tokens, leading to steadier convergence and consistently better scores.
The final Nemotron-Flash models interleave DeltaNet, Mamba2, and a few full attention layers, tune depth, width, and tokenizer for the latency goal, and beat other small models on reasoning benchmarks while running faster.
----
Paper – arxiv. org/abs/2511.18890
Paper Title: "Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models"
Mahesh Deshmukhe retweeted
This paper shows how attackers could hide AI powered hacking agents inside normal looking AI traffic to stay unseen.
In a lab test, 1 autonomous agent took over a company style network in under 60 minutes without any security alerts.
They focus on command and control, the hidden channel that sends instructions to already hacked machines.
Their design reuses the Model Context Protocol, a way apps talk to AI services, as a stealth command path.
Tiny agents on victim hosts send tasks over this protocol and call Large Language Model APIs to plan commands.
That traffic looks like normal encrypted AI use and only appears when needed, so it blends into other AI calls.
The authors say this helps red team training but is dangerous if attackers copy it, so defenders must watch AI related traffic, not only beacons.
----
Paper – arxiv. org/abs/2511.15998
Paper Title: "Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming"
Mahesh Deshmukhe retweeted
New AMD paper builds a small fully open language model Instella.
Instella is a 3B parameter model family where the authors share weights, training code, and data recipe so others can fully reproduce it.
They first train on huge open text, then on more reasoning heavy data, and finally on about 2.3M instruction examples with a preference step that shapes answers toward human choices.
They also make Instella Long, which can read inputs up to 128K tokens using extra long training data from books and papers, and Instella Math, which gets more math datasets plus reinforcement learning that rewards better solution steps.
Across many benchmarks these models beat earlier fully open ones and come close to the best open weight baselines while staying transparent.
----
Paper – arxiv. org/abs/2511.10628
Paper Title: "Instella: Fully Open Language Models with Stellar Performance"
Mahesh Deshmukhe retweeted
The paper shows how to train a smaller chat model using only the answers from a big closed teacher.
"Black box" means the student never sees the teacher’s internal weights or gradients, it only sends prompts to the teacher and reads the text replies, like talking to an API.
The goal is to transfer the teacher’s behavior into a smaller student model, so the student learns to act like the teacher but is cheaper to run.
Traditional black box distillation just copies teacher answers with supervised learning, but that only trains the student to match surface text, not deeper behavior.
This paper instead builds a discriminator that learns to tell teacher replies from student replies and then uses that as a reward signal.
With this method, a 14B student gets chat quality close to the teacher on common evaluation sets.
"On-policy" here means the discriminator is trained on the current student’s own outputs and the student is also updated based on that same, fresh distribution of outputs.
That is different from older "off-policy" setups where the reward model is trained on some fixed dataset that does not match what the student currently produces.
---
Instead of just fine tuning on teacher replies, the student is treated as a generator and a discriminator learns to tell teacher from student.
The discriminator gives each prompt and reply one score as reward, and reinforcement learning pushes the student toward higher scoring replies.
Student and discriminator update together on current student outputs after a short warmup that first teaches rough imitation.
Because the reward model is on policy, training stays stable and avoids reward hacking tricks.
Across Qwen and Llama students, scores rise on both in domain and out of domain tests and the teacher gap shrinks.
----
Paper – arxiv. org/abs/2511.10643
Paper Title: "Black-Box On-Policy Distillation of LLMs"
Mahesh Deshmukhe retweeted
This paper shows how to detect hidden failures in multi agent AI by analyzing their execution traces.
Key finding is that can spot when a multi agent AI quietly goes wrong just by looking at numbers extracted from its logs.
Those numbers describe how long the run was, how many tools it used, how many tokens it spent, and so on, and a small machine learning model can flag runs that do not look like normal ones.
These agent systems use LLMs to pick tools, so the same query can drift, loop, or miss details.
They record every agent, tool, and LLM call for each request, creating a complete trace of what happened.
From each trace they compute 16 numeric features summarizing token counts, timing, and the overall shape of the path.
Each trace is labeled normal or anomalous using expert reference paths, checks for repeated calls, and any recorded errors.
They train several models on these features to tell normal and anomalous traces apart.
XGBoost, a supervised tree model, works best, and a one class method trained only on normal traces comes close.
Models rely most on path features like total steps and tool count, and they struggle with short drifted runs.
----
Paper – arxiv. org/abs/2511.04032v1
Paper Title: "Detecting Silent Failures in Multi-Agentic AI Trajectories"
Mahesh Deshmukhe retweeted
The paper asks how far raw next pixel prediction can scale for vision and finds compute is the main bottleneck.
Next pixel prediction treats an image as a pixel sequence and trains a Transformer to guess the next pixel.
Because a single pixel has little meaning, the same size model needs much more pixel data than text.
At low resolution like 32x32, the compute optimal setup for classification favors larger models, while the best setup for generation favors faster dataset growth.
As resolution rises, the optimal plan tilts even more toward much bigger models and relatively slower data growth.
Scaling experiments using IsoFlops profiles and extra runs show that next pixel pretraining is limited mainly by compute, not by the available image data.
----
Paper – arxiv. org/abs/2511.08704
Paper Title: "Rethinking generative image pretraining: How far are they from scaling up next-pixel prediction?"