@yan04540903
Joined October 2018
字节跳动 Seed 团队发现,DeepSeek-V4 在长文本处理中存在一种规律性弱点。研究人员让模型补全一段代码,代码本身没有变化,只延长了前面一段无关注释,模型的答案就会在正确和错误之间来回切换。V4-Flash 基础版每隔 4 个 Token 重复一轮,V4.1-Flash 也存在类似现象,周期变成 2 个 Token。 问题与 DeepSeek 节省显存的「分块 KV 缓存压缩」有关。模型会把连续几个 Token 的信息合并保存,降低处理长文本的成本。但这种压缩会让不同位置的信息得到不同程度的保留,某些位置的内容更难在之后被准确找回。字节将这种周期性检索差异称为「相位敏感性」。 在约 12.8 万 Token 的检索测试中,V4-Flash 基础版在不同位置上的准确率最高相差 40.2 个百分点。经过后训练,差距缩至 19.1 个百分点;V4.1-Flash 进一步缩至 6.1 个百分点,但周期性波动仍然存在。字节还从零训练了多组对照模型,发现波动周期与压缩步幅一致,未使用分块压缩的对照模型没有出现相同规律。
13
3
4
76
15,062
Maxwell’s equations become simpler when time is replaced by frequency. The key transformation is ∂/∂t → jω. Time derivatives become multiplication, turning differential equations into algebraic relationships between electric and magnetic fields. The result is a powerful framework for analyzing waves, antennas, and electromagnetic systems. One frequency, four equations, an entire electromagnetic world.
8
34
2
206
5,208
only 9MB (!!) of weights for a single voice, single language text-to-speech model, distilled from Kokoro-82M only 8M parameters, runs anywhere, ultra fast. no more excuses to have 'robotic' voices, except on nostalgia/artistic lanes now
Kokoro TTS, but 10× smaller 🫰 Paradee distills Kokoro-82M into a 8M-param TTS just 9 MB of weights can speak faster than real time on a single CPU thread and runs on your toaster ▶️ on Spaces hf.co/spaces/hugging-apps/pa…
10
69
5
1,079
47,947
1/ We have autoformalized the resolution of singularities in Lean. Hironaka's 1964 theorem is one of the great mathematical results of the 20th century. It says that any singular variety is the "shadow" of a smooth one living in higher dimensions. So why did we do this? 🧵
14
78
10
448
38,070
What is not covered in the report is how we scaled model training to be efficient on 768 B200 GPUs. I explain it in detail here: aleph-alpha.com/en/blog/scal… A how to guide to scale up model training without wasting your GPU hours. Even if you only have 1 or 2 GPUs and cannot scale, our optimisation philosophy can still make your GPUs go brrr. I couldn't explain ALL the details of our method in the blogpost, since the big launch came after. You can connect the dots now, the 30B MoE architecture was in fact Kolibri Origin all along.
14
78
3
740
135,145
you can still play GTA 5 on a web browser after it got taken down btw bless internet archive web.archive.org/web/20261006…
46
196
14
4,131
262,961
yan retweeted
someone ported 𝗚𝗧𝗔 𝟱 to run entirely inside a web browser. no App Store bypass. no jailbreak. no installs. just opened it on an 𝗶𝗣𝗵𝗼𝗻𝗲 𝟭𝟱 𝗣𝗿𝗼 𝗠𝗮𝘅 and it launched in 𝟰 𝘀𝗲𝗰𝗼𝗻𝗱𝘀. URL: playgta5.com
3
16
1
224
620,041
We shipped a mobile game with no game engine. Ninety-Nine is built with React Native, @expo , Skia and Reanimated. Skia handles most of what you see: • hand-drawn paths • live drawing under your finger • score animations • water, cut shapes and folding effects No Unity. No Godot. Just React Native. Every frame in this video is the real app. After almost 10 years building mobile apps, this was my first game and a fun way to push React Native beyond normal app UI. If you’re building with React Native + Skia, happy to answer questions. #ReactNative #ReactNativeSkia #Expo cc: @wcandillon
24
35
3
413
20,204
yan retweeted
苹果统一内存卖得有多贵,大家心里都有数。 今天在 Reddit 刷到一个神级项目直接看傻了: 有人用一条 10 Gb/s 的 USB-C 线,把手里的 iPhone 变成了 24GB MacBook 的“副显卡”,用来跑本地 Qwen 27B 大模型,长文本速度暴涨 40%+!
103
324
35
2,811
350,240
yan retweeted
OMG... SERIUSAN? ada meja bisa dilipat kaya gini, sebagai orang yang punya kamar kecil sih seneng banget 😭
17
923
115
13,369
713,810
Attention is fate. If your attention is compromised, no amount of willpower will save you. Defend the gates of your consciousness with your life, for those gates are, in fact, your life
36
767
41
4,993
94,587
Breath of the Wild rodando liso no PS5 em 2x de resolução
166
118
127
4,175
592,771
when you have crazily large max logits, FA kernel is fragile. (large mean key is kinda the same as large max logits, cuz mean key can shift logits) but it's nonsense that model needs logits larger than 100. exp(100) is more than 10^43. I think it's an optimization problem.
We trained an LLM with FlashAttention-3 in BF16 🔥 Training looked healthy for a long time. Then the grad norm shot up 1000× 📈 and the loss ended 0.2 nats above FP32 attention. Not a single NaN 😶 The cause is in the attention backward 🧵 📄 arxiv.org/abs/2609.34272
4
14
142
8,947
国内制度漏洞,让有钱人赚钱实在太容易了! 纳指 ETF 新批外汇额度被单一大户直接包圆,在圈内炸了。 当时二级市场相比基金净值溢价 8%-10%,申购后卖出就能套利,100 万申购大约赚 10 万。 9 月 16 日新额度落地,一位客户直接拿下全部 2400 万对应申购额度,一把套利近 200 万,其余投资者一分额度都拿不到。 事件发酵之后,各家基金迅速出台单客户申购上限,补上漏洞,回归公募普惠的初衷。 这就是典型的你越堵漏洞就越大,为啥要限制国内纳指ETF的购买?设定一个额度,搞得不但溢价严重,每人每天定投只有10元这些荒唐事情,市场供求关系就是个最好的无形的手,但偏偏要人为干预。。。
12
19
2
248
70,464
Open-source magnetic tactile sensor for $5! 🧲 Researchers from New York University introduced a magnetic tactile sensor that's low-cost, and easy to fabricate, democratizing tactile sensing for robotics. Operating in unstructured environments like homes and offices requires robots to sense forces during physical interaction. Yet the lack of a versatile, accessible tactile sensor has led to fragmented solutions and often force-unaware, sensorless approaches. Building an eFlesh sensor requires four components: a hobbyist 3D printer, off-the-shelf magnets (less than $5), a CAD model, and a magnetometer circuit board. The sensor is 3D printed with magnets embedded in the middle layer. Based on chosen mechanical properties, magnets displace in response to contact forces, measured by a magnetometer underneath. An open-source design tool converts simple OBJ/STL files into 3D-printable STLs. This enables application-specific sensors for robot hands, grippers, quadruped feet, and more. Slip detection generalizes to unseen objects with 95% accuracy. Visual-tactile control policies improve manipulation by 40% over vision-only baselines, achieving 90% success on precise tasks like plug insertion and credit card swiping. All design files, code, trained models, and conversion tools are openly available. 🔗 Project page: e-flesh.com ~~ ♻️ Join the weekly robotics newsletter, and never miss any news → ziegler.substack.com
28
193
25
1,760
119,692
We trained an LLM with FlashAttention-3 in BF16 🔥 Training looked healthy for a long time. Then the grad norm shot up 1000× 📈 and the loss ended 0.2 nats above FP32 attention. Not a single NaN 😶 The cause is in the attention backward 🧵 📄 arxiv.org/abs/2609.34272
20
63
13
547
60,961
We release Whistle: speech to text in a 16.9MB file, running on the CPU in the same engine as Needle. It mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish. cactuscompute.com/blog/whist…
106
227
57
3,629
383,888
yan retweeted
category theorists don't want you to know that category theory is dead simple if you are allowed to visualise it in 3d tear down the category cabal now!!!
83
325
31
2,822
150,458
过去一周多,可能以后被证明是AI推理部署史上天翻地覆的一周。 上个周末,我的本地机双5070ti试跑Qwen Flash,只能做到200 prefill,10decode。 今天,在Strata和我自制的PR的支持下,目前可以跑到2200/67,而且只需要一张卡。 即读写速度都提高了近十倍。 更好的消费卡比如4090/5090,当然也有幅度相当大的升级(目前5090在流式处理的支持下,速度已经开始逼近RTX6000,当然显存还是不足)。 这当然带来了消费级硬件的新一轮大涨价。如今,就连曾经无人问津的 32G B70英特尔显卡都卖断货涨价了。 2200/67的速度是什么意思? 今天除了deepseek,muse等少数模型,大部分API达不到这个速度。 而一个普通的家用游戏电脑就可以达到这个水平。 而这个水平,一周以前,只有5090,DGX spark,M3U这样的机器可以达到。 也就是说,在过去的一周多,全球的普通家用电脑带来了十倍于上周的AI产量。 这个增量,得益于deepseek- Qwen flash这种特殊设计,把智力分成了注意力,专家,和查表, 也得益于Strata这样优秀的架构,把整个处理流水线化,让不同的智力,存储在不同速度的存储设备中,在需要用时刚好抵达。 这对所有愿意拥有本地智力的人,当然是极大的利好。 也许模型层的设计,正在快速收敛到一个目前硬件条件下的最优解? 然而这对依赖出卖模型算力的生意,尤其是平庸的智力的商业来说,也可能是催命符。
买了5090和RTX 6000 pro的各位 才算是真正的慧眼,真正的成功
71
86
13
542
251,358
一位大哥分享的让普通人半年学会英语的土方法,感觉挺实用、挺接地气的。
165
900
33
4,602
320,033