Driven by my interests. For now is NLP

Joined May 2023
For anyone curious how Jev works, I made a visual explanation using @claudeai :) This is based on the Qwen2.5-RLCD model which @harshagundal released on @huggingface The idea is to replace autoregressive LLM generation by a single Transformer decoder (of a pre-trained LLM), which processes the context + JSON schema only once. The keys and values of those tokens are cached. Next, for each field of the JSON schema, we: 1. pass its field suffix tokens through the Transformer decoder again (reusing the KV-cache) 2. obtain a final hidden state, which we pass through the language modeling head 3. we obtain scores, also called logits, for all tokens in the vocab of the LLM 4. we only look at the scores of the tokens we care about for the given field, and pass those through a softmax to obtain probabilities which sum to 1 5. we take the token with the highest probability. The benefits of this are that: 1. it's fast (we don't need to generate the JSON schema token by token) 2. it's 100% valid JSON (we don't need to rely on the model to generate a valid schema)
12 million views for a JSON classifier? Yeah, we're in a bubble
94
414
34
3,819
415,209
xcjqaq retweeted
Welcome back encoder-decoder. 552B total, 8B activated on decode, 16B activated on prefill. I have never seen a different number of parameters activated in the same model before
25
152
43
1,930
291,689
shouldnt this be called decoder-decoder to respect YOCO
Welcome back encoder-decoder. 552B total, 8B activated on decode, 16B activated on prefill. I have never seen a different number of parameters activated in the same model before
3
29
2
305
31,080
Introducing PC-ALM, a local-learning alternative to backpropagation. Our method trains 1000-layer neural nets using only local dynamics, and without backprop. Blog: pub.sakana.ai/pc-alm/ Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit assignment without explicit use of backprop? We look for inspiration in two related fields: distributed optimization and NeuroAI. In NeuroAI, predictive coding asks each neuron activation to solve an energy-based inference problem instead of using a standard forward pass. That inference step can be implemented as energy-minimization dynamics on local prediction errors. This perspective -- each layer as a dynamical system -- has proven promising, but performance of predictive coding hasn't scaled well with depth. Credit signals at far ends of the network struggle to diffuse into internal layers. We turn to distributed optimization, generalizing predictive coding to use an augmented Lagrangian instead of energy. This motivation stems back to a classic 1988 paper by LeCun, showing that the Lagrange multipliers of a deep network can be identified with gradients of a supervised loss. The augmented Lagrangian then bridges LeCun's perspective to the standard predictive coding that is used in NeuroAI. We find that this new perspective yields a natural PC-like alternative to backpropagation, resulting in a method we call PC-ALM. PC-ALM differs from PC in that it introduces dual neurons (Lagrange multipliers) as part of the layer-local dynamics, resulting in each layer acting as a PI feedback control system to minimize local prediction errors. We find that PC-ALM is capable of propagating signals to seemingly arbitrary depth, especially in deep narrow networks where standard PC struggles to learn. Ultimately, our motivation here is to understand how distributed physical systems, such as the brain, can compute credit signals using only local coupling and local dynamics. PC-ALM may also inform deep learning in neuromorphic hardware, where dynamics are cheaper than on GPUs. Paper: arxiv.org/abs/2605.31022 Code: github.com/SakanaAI/pc-alm
69
393
99
2,904
438,938
This is a good visual btw for Looped Transformers You just pass your inputs T times through the layers instead of just one time I highly recommend the post by @rasbt which explains the pros and cons Figure from the “Universal Transformers” paper paperswithcode.co/paper/1807…
Looped Transformer is a method btw on Papers with Code! Find all papers on this topic here: paperswithcode.co/methods/lo…
2
19
102
7,219
Can Looped Transformers still help when FLOPs, parameters, and KV cache are all matched? We introduce SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers. The answer is yes. And the advantage grows with scale. 🧵 1/8 Paper: arxiv.org/abs/2609.01343
8
71
7
451
74,647
xcjqaq retweeted
a few thoughts on "recurrent depth" transformers the main question: why recurrent depth instead of just scaling depth? it's not faster at inference or training* since you still go through the full "effective depth", the advantage is storage (for kv cache storage btw you could do kv sharing) and a few other things i tried to list below 1) could be better in undertrained regimes (too few data or compute constraints) but i don't think any frontier model is in this regime? could also help convergence somehow, maybe it's just not possible to train a 10T effective param model (so maybe 2Tx5) without tricks like this, but i doubt it.. 2) one positive aspect could be better adaptive depth/compute (which is a good direction imo): depending on the token, you skip some recurrent blocks. but you could also just skip layers (or skip moe routing like in longcat flash) and get something quite similar you could also imagine co-training models this way, sol with 8 blocks, terra with 4, luna with 2 or something like that 3) another interesting thing (to verify?): you could simulate higher batch size at inference by overlapping the computation of depth_block=1 of request 1 and depth_block=t of request 2 on the same gpus, so effective batch size = batch size * number of recurrent blocks (and you could even overlap across models if you do the co-training i mentioned above!), see the picture attached for an illustration of what i mean still quite skeptical overall, i think scaling depth is just more expressive when you are not compute/data bound but excited to see a potential new direction *on "faster to train", could be if you only backprop through the last few loops or something like that
50
58
12
816
119,288
xcjqaq retweeted
“Fast Weight Attention for Continual Learning” Fast-weight models try to make attention cheaper by continuously rewriting a small fixed-size memory. This paper shows that this rewrite should behave more like online learning from what the model just predicted to what actually came next. So they made adaptive memory updates that decide how much to learn, forget, and rehearse, while keeping constant-size memory and efficient training. alphaxiv.org/abs/2608.27763
6
89
7
596
44,874
The best peft paper I've ever read this year. alphaxiv.org/abs/2608.31036
3
61
4
441
31,363
xcjqaq retweeted
Songlin Yang's video explanations youtube.com/watch?v=d0HJvGSW… youtube.com/watch?v=RTJKXK5L… Songlin Yang's blog post Design intuition sustcsonglin.github.io/blog/… Kernel algorithm (I skipped to "A Chunkwise Algorithm for DeltaNet" section) sustcsonglin.github.io/blog/… FlashKDA github.com/MoonshotAI/FlashK… I prompted Codex to explain it to me, only after that the deep dive made sense github.com/MoonshotAI/FlashK… vLLM serving explanations vllm.ai/blog/2026-07-22-kimi… vllm.ai/blog/2026-07-27-k3#p… The great Zhihu post for explaining AttnRes Training zhihu.com/question/201699309… Inference zhuanlan.zhihu.com/p/2017528… Jianlin Su's MoE 環遊記 series kexue.fm/tag/moe/1/ Explains the concept of MoE from math first principles Papers Kimi Linear, LatentMoE, Kimi K3 tech report N/
3
71
2
426
37,143
"On-Policy Delta Distillation" So instead of copying everything a teacher model prefers, this paper asks what changed after the teacher learned reasoning. It compares the reasoning-tuned teacher to its base model, then distills that difference into the student. This focuses training on reasoning-specific behavior, not the teacher’s generic style or old pretraining habits. This approach gives better math, code, and science reasoning across Qwen3 and Gemma4, especially where normal on-policy distillation can hurt strong models.
2
54
1
414
28,406
"T²MLR: Transformer with Temporal Middle-Layer Recurrence" T²MLR adds recurrence where Transformers seem to reason best, which is the middle layers. So it caches a middle-layer state from the previous token and feeds it into the next token, avoiding the token-space bottleneck. By only recurring through 20% of layers, it can beat full recurrence, improves reasoning, and can retrofit a 1.7B model with little inference overhead.
2
44
2
210
12,309
xcjqaq retweeted
Revisiting Convergence Results in Convex Optimization (VII) kexue.fm/archives/11804 - New identity quantitatively link weight averaging and learning rate decay - discuss why uniform averaging loses to sliding, and why Schedule-Free still can't fully escape scheduling
3
8
84
7,917
Nice, and quite readable manuscript on the mathematical basics of diffusion models. arxiv.org/abs/2607.01693
6
53
1
464
32,140
xcjqaq retweeted
very nice tech report from @AntLingAGI, they changed the arch of their previous 1T model to make it more efficient (from full GQA -> 7:1 lightning attention:MLA) and better at agentic tasks with 10T tokens of continual pre-training. also a big focus on reasoning efficiency!
6
29
1
252
11,443
ExpRL: Exploratory RL for LLM Mid-Training Use RL directly for mid-training. An LLM judge compares the sampled reasoning trace against the reference solution and assigns outcome-level or process-level dense rewards. This lets ExpRL reinforce partial progress, useful intermediate reductions, and productive reasoning behaviors that sparse final-answer rewards often fail to upweight. On challenging math reasoning tasks, ExpRL yields stronger RL priming than SFT, sparse-reward GRPO, and self-distillation, and provides a better initialization for subsequent sparse-reward RL.
2
8
2
90
16,350
arxiv.org/abs/2606.03825 Dynamic convolution on QKV! An old idea came back again.
1
25
1
154
13,255
xcjqaq retweeted
I am a big fan of Jianlin Su's blog because it always starts from first principles in mathematics, rather than "ML tricks", to approach a typical ML problem (eg. training-free MoE load balancing). Here is me trying to "reinvent" one such blog which provides an elegant alternative to compute Muon, by filling in all the derivations that the blog skips for a less math-savvy audience (besides being entirely in Mandarin). The goal of the blog is to find a way to compute a essential component of Muon, ie. the left and right singular value matrices U and V for the gradient G, **individually**. In the standard form, Muon really just needs their product UV^T, hence the standard way to compute it via computing a low-rank polynomial of G many times ("Newton-Schulz"). But there are more variants of Muon to control the properties of model updates if we can get both individually, hence the blog's proposal to revisit some fundamental linear algebra techniques for the computation. The methodological takeaway from the blog's thought process is that there are three components to breaking down a ML problem: (1) how to be able to compute something (power iteration), (2) how to compute it fast (cholesky decomposition), and (3) how to compute it accurately given finite floating points (repeated orthogonalization). The goal of reading inspiring blogs like this is, in Feynman's term, to be able to "reinvent" them at any time to grasp the fundamental approach of doing similar work. Original blog: kexue.fm/archives/11654
10
141
3
1,686
78,588
Xiaomi MiMo-V2.5-Pro achieves multiple breakthroughs in the latest Arena rankings (Apr 26, 2026) 🔥 🏆 Text Arena (Expert) — #6 globally | #1 open-source model Also #1 among Chinese models, with Xiaomi ranking #3 globally by lab, behind only Anthropic and OpenAI. Expert is defined by high-difficulty tasks and expert voting, measuring core model intelligence. 🏆 Text Arena (Overall) — #2 open-source globally Strong across math, coding, creative writing, and general text tasks. 🏆 Code Arena (WebDev) — #3 open-source globally Evaluated by real community blind voting on frontend code generation. 🏆 Text Arena sub-rankings — #1 open-source globally in 4 categories Hard Prompts, Hard Prompts(English), Instruction Following and Long Query. Real-world preference, real model strength.
51
73
13
743
85,145
xcjqaq retweeted
An explainer on CSA
Replying to @nrehiew_
The biggest change is 2 types of 'sparse' attention. The first is Compressed Sparse Attention (CSA) which is the successor to NSA with all 3 elements of compressed attention, selection and sliding window. It operates in a sliding window of size M(=4) with stride M. The hidden states are projected to a KV (similar to MLA) and another compressed vector S. S is softmaxed and used as the mixing coefficient for all KVs belong to tokens in this M-sized sliding window, resulting in a single composite compressed KV form from these M tokens. This is done twice over 2 M-sized sliding windows with different coefficients/projections for each. However, since the stride is M, the compression ratio for the KV is 1/m. If the current query token doesnt complete the M sliding window, its padded. This compressed KV is passed to the DSA style indexer and then MQA is performed. Its a little hard to explain purely with words but it can be thought of as 2 levels of compression: MLA style dimension based and then block wise along the sequence dimension. Lastly, we have the sparsity implemented via the indexer
1
6
117
11,429