@athletic_coderi
iAccount based inSouth Asia
About this account
- Account based in
- South Asia
- Connected via
- South Asia App Store
Account-level information from X, not a live location or the device used for a specific post.
post training & inference, curr: @zomato; prev: @google, gsoc @tensorflow;
Joined February 2020
- Tweets10.5K
- Following1K
- Followers20.5K
- Likes9.2K
My current LLMs stable
- GLM 5.3 Flash
- DeepSeek V4.1 Flash
- Qwen 3.8 27B
- Kimi K3
I'm telling you future of AI is open and local.
We're just getting started.
anshuman retweeted
learn about inference engineering as fast and deep as you can
i am literally giving it on silver platter, a detailed worklog of optimizing GEMM in CUDA.
in this piece:
> start with a naive implementation
> go on to identify and quantify the performance issues. > explore the concept of online softmax
> implement the kernel quantify the gains in terms of memory & FLOPs
> apply the coalesced memory pattern
> again quantify the gains in FLOPs, Comms, and Memory
> write the CUDA code.
read it here:
athleticcoder21.github.io/in…
FUTURE OF AI IS LOCAL
FUTURE OF AI IS LOCAL
FUTURE OF AI IS LOCAL
FUTURE OF AI IS LOCAL
FUTURE OF AI IS LOCAL
FUTURE OF AI IS LOCAL
FUTURE OF AI IS LOCAL
FUTURE OF AI IS LOCAL
learn about inference engineering as fast and deep as you can
i am literally giving it on silver platter, a detailed worklog of optimizing GEMM in CUDA.
in this piece:
> start with a naive implementation
> go on to identify and quantify the performance issues. > explore the concept of online softmax
> implement the kernel quantify the gains in terms of memory & FLOPs
> apply the coalesced memory pattern
> again quantify the gains in FLOPs, Comms, and Memory
> write the CUDA code.
read it here:
athleticcoder21.github.io/in…
inference and post training go hand in hand
think about it people
let's say worst of it all happens - no new open models from China.
we still have good old llamas which we can continually pre train and post train.
all you need is some gpuuuus and the skills.
Dylan Patel (@dylan522p) of @SemiAnalysis_ says open source is dying:
"There's multiple Chinese model labs who are telling all the inference guys, 'Our next model's not going to be open source. We're going to license it to you.'"
" Open is dying quickly, unfortunately."
Readers added context they thought people might want to know
This July 2026 quote (recirculated without date) claimed Chinese labs were shifting away from open source. Multiple labs have since released large open-weight models, including Qwen updates in Aug-Sept 2026.
huggingface.co/blog/state-of-…
rits.shanghai.nyu.edu/news/
wired.com/story/chinas-o…
x.com/sourceryy/stat…
at this point someone is killing exa and tavily everyday.
yet they somehow survive.
we are definitely in a bubble.
We just killed Exa, Tavily and Brave.
Your agent can now search & fetch any webpage for 100% FREE.
Them: $7 per 1,000 searches.
Us: $0. No subscriptions, no quotas.
Humans search Google for free. Agents shouldn't have to pay either.
Made possible by @Tiny_Fish and @MonidHQ.
suddenly everyone has their own implementation of jev.
Kev-0.5B: A tiny open source Jev-like decision model with a TypeSafe-compatible API based on Qwen2.5-0.5B that you can train and run on a MacBook Pro.
Model card and weights are available on GitHub
github.com/jaredpalmer/kev
anshuman retweeted
inference is all you need
wrote a detailed worklog of optimizing Softmax kernel in CUDA.
in this piece:
> start with a naive implementation
> go on to identify and quantify the performance issues. > explore the concept of online softmax
> implement the kernel quantify the gains in terms of memory & FLOPs
> apply the coalesced memory pattern
> again quantify the gains in FLOPs, Comms, and Memory
> write the CUDA code.
read it here:
heyyanshuman.com/posts/makin…
inference is all you need
wrote a detailed worklog of optimizing Softmax kernel in CUDA.
in this piece:
> start with a naive implementation
> go on to identify and quantify the performance issues. > explore the concept of online softmax
> implement the kernel quantify the gains in terms of memory & FLOPs
> apply the coalesced memory pattern
> again quantify the gains in FLOPs, Comms, and Memory
> write the CUDA code.
read it here:
heyyanshuman.com/posts/makin…
anshuman retweeted
inference is future.
tried going as much in detail as possible for the writing GEMV kernel in CUDA,
hopefully it will be a good read. ciao!
heyyanshuman.com/posts/makin…
inference is future.
tried going as much in detail as possible for the writing GEMV kernel in CUDA,
hopefully it will be a good read. ciao!
heyyanshuman.com/posts/makin…