photography, neural rendering, 3D computer vision, and verified insane masters student at @NYU_Courant

New York, New York
Joined May 2025
Being in university is realizing that you are not as smart as you thought and all your classmates are somehow geniuses
being in university is just realising that the average person is not smart
190
15,329
959
112,186
1,510,499
Daksh Shah retweeted
Very excited to share LoGo, my internship project at @theworldlabs on post-training world models! We found that reward design matters a great deal in post-training long-horizon video gen for 3D consistency, echoing what we see in other domains, e.g. LLM reasoning. Check out more 👉 ziqi-ma.github.io/logo-websi…
9
34
8
266
33,656
Daksh Shah retweeted
Paris ce soir
56
1,301
143
15,769
244,589
Traveling around Switzerland for the last few days has made me realize just how bad American rail infrastructure is
5
275
Daksh Shah retweeted
First in my bloodline to publish a paper
43
384
105
3,779
75,368
Rockaway Beach in NYC felt like a scene from a movie tonight. Insane sunset, blowing dust and high surf
115
2,560
235
30,011
1,102,528
I’m pleased to announce our work Contrastive Language Model (CLM), an ultra-fast System One Model trained with contrastive learning for on Internet-scale data. CLM matches or outperforms Jev at substantially lower latency (up to 9x). Exciting work with @jackyk02 @hangoo_kang @JonSaadFalcon @drmapavone @Research_Hazy @Azaliamirh !
Introducing Contrastive Language Model (CLM): an ultra-fast System One Model trained with a contrastive learning objective that connects states and actions. CLM-8B is pre-trained on internet-scale data and delivers up to 9× faster inference than Jev ⚡ while achieving comparable performance across computer-use, gaming, and tool-calling tasks. With lightweight fine-tuning, CLM-8B sets a new SOTA on challenging agentic coding benchmarks, such as DeepSWE (81.6%) and Terminal-Bench 2.1 (87.6%). In contrast, Jev fails to serve as an effective verifier for these long-horizon tasks. We also build an efficient training and serving infra for CLMs by disaggregating states and actions, allowing their embeddings to be cached and reused independently. This substantially reduces inference latency in settings where the state evolves continuously while the action set remains fixed. Finally, we establish scaling laws for CLMs and show that the test contrastive loss decreases predictably as a power law in training compute, model size, and dataset size. 📄 Blog: contrastive-lm.notion.site 💻 Code: github.com/Contrastive-LM/CL… 🗣️ Discord: discord.gg/5dAQEDJBs 🤗 Data & Models: huggingface.co/Contrastive-L… More details on CLM’s architecture, data recipe, and scaling laws in the thread below 🧵
10
13
1
106
21,727
I'm looking for full-time Research Engineer / Machine Learning Engineer roles starting Jan or May 2027! I'm an MIT undergraduate (majoring in Physics + AI) researching generative modeling in Prof. Kaiming He's group. I'm especially interested in LLM pretraining/post-training and multimodal generation.
18
13
3
442
74,536
Daksh Shah retweeted
These are not 3D reconstructions but 3D-native generation! Check out our new work, GAE, to generate videos and geometry simultaneously! Code and models are already released!
Introducing GAE🥳🥳 Generate once. Decode RGB + point clouds. GAE generates directly in a geometry-native latent space. The same generated latent supports both appearance and 3D geometry. The key choice is where generation happens. GAE compresses geometry-foundation features into a shared latent space. A conditional flow model generates directly in that space; RGB and geometry are decoded from its final state. The video pairs show camera-conditioned world generation; the character pairs show text-to-image generation. These native showcase clouds are not reconstructed from generated RGB. Code: github.com/TencentARC/GAE Page: jiah-cloud.github.io/GAE.git…
5
33
4
389
37,054
Daksh Shah retweeted
Happy to announce that my first paper, “K-NeAS: Scalable Multi-Material CT Reconstruction Using Neural SDFs” has been accepted as an oral presentation at the Off-Grid workshop at MICCAI 2026! See you in Strasbourg! Paper: arxiv.org/abs/2607.14415 Code: github.com/dakshshah03/K-NeA…
1
1
6
203
Academia is filled with people who constantly talk about how busy they are, but never seem to complete anything
74
230
81
2,411
103,954
3D/4D reconstruction is one of the great achievements of computer vision. With roots in photogrammetry, multi-view geometry was much studied (e.g. Hartley & Zisserman). In the last decade, it was successfully combined with deep learning. Shape priors which come from recognition got incorporated (e.g. for humans, in HMR) and recent systems like VGGT and SAM 3D can do amazing stuff. This progress is directly relevant for robotics, which (duh) operates in the 3D world. Many robotics papers in the learning era weren't exploiting 3D structure, which IMHO is just wasting valuable signal. But there has been a strong real2sim2real tradition for years e.g. here are two projects from my group (and there are many others, I am not claiming exclusivity). I am glad that people have rediscovered this capability with Astra. But please do give credit to human researchers, not just AI models. zhec.github.io/rhoi/ , videomimic.net/
8
67
11
668
133,601
Everyone’s talking about their Yann LeCun number but what about your Daksh number (with a whole 1.1 publications)
My friend @lino_levan just sent me this, try to find the Erdős number between me and you via research coauthorships: dakshnumber.alphaxiv.org/
105
Diffusion LLMs and speculative decoding promise much faster agents. Yet agents need long contexts, and training on them is painfully slow. Introducing Context-Sharded Block Parallelism (CSBP), a new distributed parallelism strategy unlocking significant training efficiency for diffusion LLMs, with speedup gains growing with context length 🚀 ⚡ 7.59× faster DFlash2 speculative decoding drafter training ⚡ 1.61× faster block diffusion fine-tuning ⚡ 1.33× faster autoregressive → block diffusion adaptation With the same GPU hours, models trained with CSBP score higher on SWE-bench Verified and Terminal-Bench Lite 👑 Open-sourced in Turbo-dLLM, our new optimized distributed training library. Advised by @Azaliamirh and with an amazing team: @PranshuChatur11 @hangoo_kang @pshroff_ @ishanskhare @KumbongHermann
10
34
9
149
81,378
what if humanities and stem students united to fight our common enemy business majors
47
2,138
390
16,641
182,894
My friend @lino_levan just sent me this, try to find the Erdős number between me and you via research coauthorships: dakshnumber.alphaxiv.org/
2
1
6
393
More than 50% of the researchers indexed on @askalphaxiv seem to be within 4 degrees from me
1
2
79
Sort of surprising as I have only one peer reviewed publication so far 😭
39
Every PhD student eventually finds a paper that contains their brilliant new idea. The paper is usually older than they are.
65
768
87
15,286
278,045