flyingpetals retweeted
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR.
We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use.
Read more: anthropic.com/news/claude-di…
MiMo-V3 is getting a new architecture. The core of it, HySparse2, is out today.
Less prefill, a smaller KV cache, better long-context retrieval—and we got all three at once.
Compared with MiMo-V2.6's Hybrid SWA architecture:
• 5.02× lower prefill FLOPs at 1M tokens
• 4.5× smaller KV cache at 1M tokens
• Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL
Why build a new architecture?
Agentic inference is a very different workload. Each round, a short action can return a long observation that needs to be prefilled, while the context keeps growing. That puts prefill cost, KV-cache size, and retrieval accuracy on the critical path at the same time.
HySparse2 tackles all three with two levels of KV sharing:
• KV Bridging: Following YOCO, full-attention layers in the cross-decoder build their K/V from self-decoder hidden states.
• KV Reuse: Within each hybrid block, sparse layers reuse the preceding full-attention layer's KV cache and selection indices.
Two more changes: token-level selection replaces block-level selection, and a forced window of recent tokens replaces the separate SWA branch, so local and global tokens share one KV cache. Since all cross-decoder KV caches now come from the self-decoder, prefill can stop once the self-decoder finishes.
Paper: arxiv.org/pdf/2609.26368
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale.
We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
完了,小米mimo V2.6这波性价比直接夯爆。2.6Pro 6块钱1m输出,我靠,这个能力又把帕累托前沿往前推了。不仅是小鲸鱼,这下智谱Kimi等真得发新模型了, Glm5.3和Kimi K3已经完全不够看了, V4.1 Flash和牛来都成了牛夫人。我不等了拜拜了,mimo plan上车了哈,马上试试咸淡。
MiMo-V2.6: The Hard Road to Scaling Up RL
MiMo-V2.6 is very likely one of the largest single RL runs, by compute, that any open-source model team has undertaken to date. In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL. That takes more than research conviction. It takes a vision for AGI, respect for the unknown, and the nerve to walk straight into the hardest problems.
The result is a model whose potential was built through mid-training and unlocked through heavy RL. Today, it is the number one open-source model. I strongly recommend reading the technical report. I believe it will become one of those papers that Agent RL practitioners keep reopening and discovering something new in each time. In my view, the research innovations and engineering challenges behind it surpass those of DeepSeek R1, which I was partly involved in.
Some will ask: why MixRL instead of MOPD? First, they are not competing choices. We ran MixRL on verifiable tasks of moderate difficulty, including code and related agentic tasks, and found that the resulting models generalize remarkably well. Second, tasks that are difficult to verify, extremely long-horizon, or simply too challenging to include in a joint RL run are trained separately. Including them would substantially reduce rollout efficiency or introduce significant rollout staleness. We then merge the resulting capabilities through MOPD. Games, 3D tasks, and tasks with subjective evaluation signals all fall into this category.
There is also a third, slightly cheeky answer. Our team is flat enough and free enough of organizational silos that MixRL simply is not difficult for us. More importantly, everyone enjoys working this way. People from different domains come together every day, driven by the pursuit of AGI and intelligence that can continuously improve itself, to confront and resolve the RL bottlenecks in each field. I will always remember the RL daily update meetings from this period. They were intense and dense, with intelligence emerging in real time.
To help the open-source community focus on solving real Agentic RL problems, we have released a Qwen model distilled from MiMo RL trajectories as a stronger starting point for RL, along with 7K diverse environments and a complete RL training framework. We hope these resources will help move Agentic RL research forward.
MiMo-V2.6 is only the beginning. In an era when intelligence is easy to replicate, we still choose the hard road toward self-improvement and AGI. Much of what lies ahead remains unknown. But we are willing to keep investing the time, compute, and passion required to take on one hard problem after another and work each of them all the way through, until intelligence crosses into a new regime.
flyingpetals retweeted
MiMo V2.6 is shipping
mimo.xiaomi.com/mimo-v2-6
已经绝望了。paper挂了大导小导名字,结果大导搞雷达项目回来勃然大怒,说ai论文挂他的名字,他又看不懂负不起责任,小导更是一点知识都不懂,负责改我的论文结果连Lm head的都没听说过,我还得向他一遍遍解释,说recurrent model就是异端邪说。我的合著者也全推给我看不懂, 这论文要撤稿了,怎么办
flyingpetals retweeted
现在 paper 中不中不那么重要 因为都是抽奖
主要看 taste 和合作者的 connection
但冷启动找到知道什么是好研究 能做好研究的小圈子 很难
今天和朋友聊了聊国内 AI 论文中介的生意,听了真是触目惊心。一篇 xxxx 的中稿文章,一作 20 万人民币,共一 15 万,以此类推。
自从我本科毕业以来,想明白了一个严肃的问题,怎么评估一个人的能力是一件极其复杂的事情,而且评估成本也非常夸张。公司面试一个人,三五轮面试,最后要 C level or VP level 的人亲自下场最后一轮,耗费受试者时间的同时,对公司也是非常显著的开销。假设一个组一年要面超过 50 人,这都是接近 7 个人月单位了(可能现在的人都不知道人月是什么单位了 😂)
考虑到此,论文对于快速评估某些能力自然是有所帮助的,即便是我面试别人 inference 相关岗位,也会去搜搜面试者简历上写的论文什么来头。简历上有 paper 的人,通过我的简历关的概率会大大增加,论文对于 evaluation 的影响可见一斑。更何况除开求职,我可以列举出无数地方论文会有显著影响:
1. 评优评先评职称;
2. 大学申请(不单是 PhD,甚至是本科和高中,任何申请制项目);
3. 各类人才项目的申请;
好吧,今天和人聊到 ICLR 的投稿数,又一想到现在的 RSI 浪潮下,已经不知道多少 paper 是 AI 想 idea,AI 写稿,AI 做实验,AI 投稿,最后 AI 审稿。糊弄一圈,然后拿到了 30% 左右的中稿率。
一想到 AI 爆发以来,工业界薪资发生的实际性暴涨,而学术圈并没有出现明面上的涨幅,许许多多的 PhD、AP、AC 和这些所谓的简历提升项目合作,从中牟取不当利益,我已经不能再说什么了。
flyingpetals retweeted
这是人工智能三大会之一,放在以前能中一篇已经是人中龙凤,但现在每个 AI 会议的论文投稿数都迎来了大爆发,Auto-research 的时代显然已经到来,学校的毕业标准可能也会发生相应改变,如果还不会用 AI 去做科研,那很快就被淘汰了。
大部分论文确实和废纸一样,但只要评价体系没有变,论文在就业市场还是硬通货,只是每个人需要在数量的基础上获得影响力更高的工作才能有更强的竞争力。
我操。这模型的性价比足够让我放弃小鲸鱼吗?
Introducing Step 5 Preview: Advancing the Pareto Frontier.
Step 5 Preview is our new flagship model for agentic work, delivering frontier-level performance across software engineering and professional knowledge work, with particular strength in finance.
- 600B total / 27B active MoE, with 1M context + Vision
- Substantially lower task cost at comparable intelligence
- Broad software engineering capabilities with sustained execution over long horizons
Try Step 5 Preview: platform.stepfun.ai
Model page: stepfun.com/step-5-preview
Open weights on Oct 15.