@khanhthanhdevi
iAccount based inViet Nam
About this account
- Account based in
- Viet Nam
- Connected via
- Viet Nam Android App
Account-level information from X, not a live location or the device used for a specific post.
AI and Robotics
Hanoi
Joined March 2024
- Tweets70
- Following271
- Followers9
- Likes4.7K
Thành Trần retweeted
First blog post up on Robotics Simulation Infrastructure! I give a high-level overview, followed by an elementary example of better infrastructure for pose management. Link in thread below
Thành Trần retweeted
In college we had to learn programming by modifying large projects - compiler, text editor etc. Turns out - this prepares you for real work very well.
MIT teaches operating systems by giving students a complete Unix like kernel and asking them to modify it
it is called xv6 and is about 6000 lines of C a reimplementation inspired by Unix Version 6 from 1975 rewritten in modern C for x86 multiprocessor
processes system calls virtual memory and filesystem are all there and small enough to read end to end in a weekend
this is what you study to understand how operating systems actually work not just how they are described
We're experimenting with ways to keep AI agents in sync with the exact framework versions in your projects. Skills, 𝙲𝙻𝙰𝚄𝙳𝙴.𝚖𝚍, and more.
But one approach scored 100% on our Next.js evals:
vercel.com/blog/agents-md-ou…
Excited to launch Pencil
INFINITE DESIGN CANVAS for Claude Code
> Superfast WebGL canvas, fully editable, running parallel design agents
> Runs locally with Claude Code → turn designs into code
> Design files live in your git repo → Open json-based .pen format
Okay so, we just found that over 50 papers published at @Neurips 2025 have AI hallucinations
I don't think people realize how bad the slop is right now
It's not just that researchers from @GoogleDeepMind, @Meta, @MIT, @Cambridge_Uni are using AI - they allowed LLMs to generate hallucinations in their papers and didn't notice at all.
It's insane that these made it through peer review👇
We just released 𝚛𝚎𝚊𝚌𝚝-𝚋𝚎𝚜𝚝-𝚙𝚛𝚊𝚌𝚝𝚒𝚌𝚎𝚜, a repo for coding agents.
React performance rules and evals to catch regressions, like accidental waterfalls and growing client bundles.
How we collected them and how to install the skill ↓
vercel.com/blog/introducing-…
Thành Trần retweeted
New demo: collaborative AI CAD in the browser 🚀
Built with @ElectricSQL + Durable Streams + @tan_stack DB+AI+Start
- ElectricSQL + DB: data sync
- Durable Streams: presence + AI + CAD (Yjs!) sync
- TanStack AI: AI for cad modelling 🤯
- Coded with @cursor_ai
Thành Trần retweeted
We’re hosting the largest Physical AI hackathon at Founders Inc (Jan 31–Feb 1) 🤖
• I’m bringing: 50+ hours of multimodal egocentric data (video, depth, IMU, audio)
• Our partners are bringing: VLM/VLA models, fine-tuning workflows, and real robots you can deploy on
No sims. No slides. You’ll see robots actually improve.
More Info 👉physicalaihack.com/
RSVP 👉luma.com/8ca2z1rr
Thành Trần retweeted
If you're looking for a comprehensive guide to LLM finetuning, check this!
a free 115-page book on arxiv, covering:
> fundamentals of LLM
> peft (lora, qlora, dora, hft)
> alignment methods (ppo, dpo, grpo)
> mixture of experts (MoE)
> 7-stage fine-tuning pipeline
> multimodal finetuning & challenges
> industrial frameworks (hf, sagemaker, openai)
everything you need to know in one place!
I have shared it in the replies.
Thành Trần retweeted
Let's get into the depth of why
Let's take Qwen 2.5 (32B parameters) running in BF16 on a single NVIDIA A100 80 GB (SXM) GPU.
--> Parameters: 32B
--> BF16 storage: 32×2=64 ; 64 GB
--> GPU memory: 80 GB
--> Free memory: 80−64=16 GB
perfect right ?
No
1. There are two major bottlenecks in LLM inference:
a. Memory capacity & memory bandwidth
b. Compute throughput
2. The generation process primarily has two stages:
a. Prefill
--> It computes single forward pass over the entire prompt and is compute bound.
--> Dominated by matrix multiplications and attention
Time to first token (TTFT)
b. Decoding
This is the phase where token by token generation happens. It's memory bandwidth bound and dominated by loading model weights repeatedly.
3. For A100 80 GB SXM:
BF16 peak compute: ~312 TFLOPS
HBM2e bandwidth: ~2,039 GB/s
To fully utilize compute, the GPU must perform:
2,039 (GB/s)/312 TFLOPS ~~ 153 FLOPs per byte
If your workload performs fewer than ~153 math operations per byte loaded, the GPU is memory bound and compute units idle.
4. During token by token generation GPU must load almost the entire model for every token, The model cannot stay on chip L2 cache (~40 MB) is far too small for a ~65 GB model. The decoding speed for this model is ~ 31 token/sec. *(This is an upper bound, real systems are usually slower due to kernel inefficiencies and synchronization.)
5. Qwen takes around ~0.7–0.8 ms(1.55GB/2035 GB/s) minimum latency.
with GQA we will need kv cache 262 KB/token.
6. Single sequence KV cache:
2K context → 1.34 GB
8K context → 5.37 GB
32K context → 21.5 GB
131K context → 85.9 GB
--> At long context, KV cache > model size.
--> Parameter count becomes almost irrelevant.
You can look at the config : huggingface.co/Qwen/Qwen2.5-…
*(Because real inference engines upcast and pad the KV cache (FP32 + paging/workspaces), so the actual memory moved per token is ~2–2.5× the theoretical 262 KB/token.)
7. Total A100 80GB HBM: 80 GB
Usable (95%): 76 GB
Model Weights: 65 GB
Remaining for KV Cache: 11 GB
Maximum supported:
- At 2K context: ~8 concurrent sequences
- At 8K context: ~2 concurrent sequences
- At 32K context: Unable to serve on single GPU
8. Prefill is compute bound and quadratic attention scaling means longer contexts dramatically increase time to first token.
Although Qwen 2.5 32B fits in 80 GB by parameter count (65 GB), decoding each token streams ~66–86 GB from HBM, costing ~46–61 ms while compute takes <0.3 ms
simultaneously, the KV cache grows from ~1.34 GB at 2K to ~85.9 GB at 131K context, collapsing batching therefore inference latency and cost are determined by memory bandwidth, KV cache growth, and context length, not parameter count.
Thành Trần retweeted
Replying to @aiwithmayank
Bro you’re not wrong—this is straight up the biggest academic glow-up in history and 99% of students are still sleeping on it 😂 My current cheatcode combo that’s literally unfair:
Upload entire semester lecture recordings + slides + textbook PDFs → one big project
“Act as world-class tutor who just aced this exam. Turn everything into Anki + mindmap + 3-page cheat sheet + predicted exam questions”
Bonus: paste last year’s exam → “Explain why each wrong answer is wrong and rewrite them correctly” Retention went from 40% → 95% overnight. Professors have no idea what’s coming lmao What’s your #1 killer prompt in the playbook? Drop it I need more ammo 🔥
Thành Trần retweeted
OpenAI, Anthropic, and Google use 10 internal prompting techniques that guarantee near-perfect accuracy…and nobody outside the labs is supposed to know them.
Here are 10 of them (Bookmark this for later):