@developerlini
iAccount based inJapan!
About this account
- Account based in
- Japan
- Connected via
- China App Store
! X says this location may be affected by a proxy or VPN.
Account-level information from X, not a live location or the device used for a specific post.
Founder @ startup · AI agents · fine-tuning & post-trained models · agentic video/aigc
Joined October 2007
- Tweets2.1K
- Following780
- Followers103
- Likes1.2K
Frank Lin retweeted
Windows now has sandboxing!
Microsoft released an open-source repo, mxc, for sandboxed code execution.
We collaborated with Windows to add mxc OS level sandboxing to Unsloth which adds just <100 ms of overhead.
GitHub: github.com/unslothai/unsloth
Guide: unsloth.ai/docs/new/studio/s…
Frank Lin retweeted
I’m obsessed with using tiny classifier models in Pipecat agents. This one, Audience, is a ~33M-param ONNX model built on Ettin that works out who each line in a scene is for. It uses names, descriptions, nearby items and where people are standing, and it can pick one character, several, the whole group, or return "unclear" if you're like me and talk jibberish.
It runs in ~6–7 ms on CPU and slots nicely into a Pipecat pipeline. Multi-character agents are really fun!
Weights here: huggingface.co/spellspeak/au…
Frank Lin retweeted
finally releasing our new OCR benchmark
tons of hard examples like this mechanical dial meter
- single value extraction
- JSON data extraction
- document transcription
- text localization and recognition
- 48 models evaluated
↓ examples and leaderboards
Frank Lin retweeted
Mobile manipulation policies can break from base pose errors of just a few cm, which are common after navigation. Can we get pose generalization without additional demos? Introducing MobileVISTA: collect demos at one base pose, get a policy that works from many 🧵👇
Website: sdwistreich.github.io/mobile…
Paper: arxiv.org/abs/2610.07511
Frank Lin retweeted
Wow Nvidia has just released a tool to turn any image into an explorable 3D world
And they made it 100% open source 🔥
- Model available on Hugging Face
- UI code available on GitHub
Lyra 2.0 creates a world you can walk through, turn around in, and even drop a robot into for simulation.
Video model can not do this
Everyone is testing Opus 5.5 to create impressive videos. The agent's ability and outcome is even more surprising than using a video‑gen model.
Just as our OneVision‑Encoder was recently accepted by NeurIPS, we would like to produce a promotional video to showcase what we consider the most elegant MLLM visual‑compression solution. Accordingly, we provided Opus with our two papers (OV‑Encoder and OV‑2).
It came back with a 3-minute launch film: every frame rendered from code, every note synthesized, every number traceable to the paper, and even the music is directly from code and with matched beats.
What a surprise to me!!! I was slumped on the sofa as if I had witnessed an atomic bomb explode😜.
The two papers👇
Frank Lin retweeted
Who wants more speech training data?
YODAS v3 is now available @huggingface
It’s 1.1M hours - the biggest audio dataset ever. And the first at this scale with stereo audio at 48kHz, with timestamped transcripts and translations. 100+ langs
hf.co/blog/espnet/yodasv3
CC-BY-3.0
Frank Lin retweeted
Impressive gift for UI/UX designers!
Ming-Image Design - 6B open-source T2I for UI, infographics, posters & text-rich designs;
- full compositions with RGBA transparent backgrounds
- pairs with Design-Layer for editable transparent layer decomp;
- 2048 x 2048
+ VLM prompt enhancer.
huggingface.co/inclusionAI/M…
We’re open-sourcing the Ming-Image-0.1-Design family:
• Ming-Image-0.1-Design, 6B
• Ming-Image-0.1-Design-Layer, 6B
• Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill
Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. 🧵
I asked Opus 5.5 to explain camera focus by building an interactive lens lab
Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost
lens.lab.sael.net
Move the focus ring and you can see the glass elements shift the sharp plane through the scene
Frank Lin retweeted
Another insane Jev use case!
Jev makes it incredibly cheap to evaluate and classify agent runs at scale.
And finally, someone open-sourced a self-improving memory layer that can put that capability to work across agent harnesses.
It turns your agent sessions into a compounding knowledge layer, where every successful run can make future agents smarter across:
- Codex
- Claude Code
- Cursor
- OpenCode and 20+ more
Beacon by @asymptotelabs continuously builds a shared history across your agent harnesses and uses Jev to identify the runs worth learning from.
It then turns the best workflows, corrections, and debugging patterns into reusable skills.
GitHub repo: github.com/Asymptote-Labs/ag….
(don’t forget to star it ⭐)
Most agent runs are messy.
They contain exploration, failed commands, dead ends, and one-off fixes that should never become permanent memory.
So Beacon preserves the full session history, while Jev helps decide what should be promoted, reviewed, or discarded.
The recording below shows this in action.
Beacon found 579 sessions across 5 coding-agent harnesses and normalized them into one consistent history.
From there, Jev surfaces the lessons worth keeping and makes them available across your agent stack.
- A pattern learned in Cursor can carry into OpenCode.
- A lesson from Claude Code can improve the next Codex run.
Every successful run adds to the shared knowledge layer, making future agents smarter.
If you want to dive deeper into Jev, I also wrote a breakdown of how it works.
The article is quoted below.
Frank Lin retweeted
卧槽,国产开源这次真把实时同传做出来了。😲
看了雨哥这条,我深度评测了网易有道的这两个模型。R2T2实时听写模型,T3PO 实时翻译。
🎉牛逼的是,两个模型一上来就把 Hugging Face 各自的 Trending 分区第一拿了:
R2T2:ASR #1
T3PO:Translation #1
🔥更骚的是,这两个都能 Streaming,而且还能直接串起来。
怎么记住这两个模型,一个是R开头,Realtime,一个是T开头,Translation。后面阅读就会顺一些。
对面的话还没说完,中文字幕已经出来了。所以我直接搞了个很抽象的 Demo:
😍如果邻居娶一个乌克兰老婆,用英语连续跟他吵 30 秒。
中间老公回嘴、老婆切俄语、再切乌克兰语,我作为隔壁老王插话。
乌克兰语,它是没训练过的,我来盘盘它 😂
ASR+Translation, 除了视频字幕,以后直接挂在电脑上,挂录音豆里,挂手机上,挂AI眼镜上:
看剧、Zoom、Meet、Teams、Discord、跨国电话……
对面老外说话,你屏幕下面直接出中文。
出国泡吧搭讪也更方便。
下面视频就是完整实测。
R2T2|实时 ASR huggingface.co/netease-youda…
T3PO|实时翻译 huggingface.co/netease-youda…
太牛逼了!Hugging Face 这周跑出一匹黑马,还是同时跑赢两个赛道的那种
语音识别 Trending 榜第一,压着 OpenAI 的 whisper-large-v3、NVIDIA 的 Nemotron、阿里的 Qwen3-ASR。翻译 Trending 榜第一,压着腾讯的 Hy-MT2、Cohere、Google、Meta 的 NLLB。两个第一都是同一家中国公司,网易有道,很多人印象里还是那个做词典的。模型叫 R2T2 和 T3PO,一个听,一个译,全开源。榜单截图我今天截的,还在第一。
一家原本在学习场景出挑的公司,在 HF 上把两个垂类做到第一,这事本身就值得点开看看。但我更想知道它到底强在哪,所以没在 GPU 上跑,把两个都塞进了一台 M1 Pro 的 MacBook,不联网。把我自己以前的口播按真实时间喂进去,中文实时上屏,英文每说完一句就跟出来。视频里的每个数字都是这台机器上的真实数字。
先说我最在意的一件事
同一段我自己的口播,8.6 秒,每 640 毫秒出一次结果。左边 Whisper 走"每次把收到的音频整段重解一遍"的伪流式,13 步里改写了 7 次,屏幕上已经显示的 93 个字被重写过,"见面了"变成"在这边了"又变回来,最后一句停在"全程媒介一找"。右边 R2T2 是 0 次改写,提交出来的字一个没动过。
更有意思的是 8.32 秒那一帧。R2T2 内部其实已经猜出"全程没接",但它只提交了"全程"两个字,把"没接"压着没放。0.6 秒后证据够了,提交的是"没截",对的。它不是猜得慢,它是知道这一刻不该猜。
为什么我在意这个。给人看字幕,闪一下无所谓。但这段文字如果是喂给下一个模型的,比如同传、比如一个要执行动作的 Agent,前面的字一改,下游就得撤销自己刚做的判断。视频第三段就是这条链路,R2T2 每提交完一句,原文直接交给 T3PO,中间没有人工,也不用等整段说完。因为 R2T2 给出来的字不会再改,T3PO 拿到就能翻,翻出来的英文也不会再改。
热词那段也值得看。不给上下文的时候,它把 Codex 写成 codet,Excalidraw 写成 escalator,无限画布写成无线画布。给一行热词之后三个全对。但我的名字"雨哥",热词给了,离线模式还是写成"宇哥",流式那一遍倒是对了。同音人名是最难扳的一类,别指望热词包治百病。
几个要说清楚的地方。
一,R2T2 是在 Qwen3-ASR 1.7B 上做的,方法在数据构造和一个叫 Longest Stable Prefix 的训练范式上,技术报告还没出。
二,Mac 上跑的是社区 audio.cpp 的移植版,9 月 19 号刚合进主仓库,brew 装的二进制里还没有,得自己编。这版每一步都重编码整句音频,句子越长越慢,M1 Pro 上 640 毫秒步长能跟上实时,官方 GPU 路线是 160 毫秒。
三,ASR 加 14B 同传同挤一块 M1 Pro 的 GPU。我一开始让 T3PO 边听边探测,结果它把 GPU 吃满,识别反而被拖到落后 3 秒。改成识别优先、每句提交完再交给 T3PO,识别回到实时,英文在每句说完后两三秒整段出来。T3PO 本身支持逐字喂、边听边译,但那需要两张卡,不是一台笔记本能同时伺候的。
四,代码 Apache 2.0,权重是网易自己的 Model License,商用前看一眼。
只增不改这条路线 R2T2 不是第一个走的,Nemotron、Voxtral 都在做。它把这条路线的准确率追到了接近伪流式和离线的水平,然后开源了,这是我觉得它值得开源社区接住的原因。
过去两年大家盯着中国开源模型的目光都在通用大模型上。这次是两个垂类,实时语音识别和同传,分别把 OpenAI、NVIDIA、Google、腾讯的模型压在榜下,而且是一家平时不怎么在这个圈子出声的公司做的。垂类这条路,中国公司是真能跑到第一的!
两个模型都开源了,代码和权重在这:
R2T2 语音识别
GitHub: github.com/netease-youdao/Co…
Hugging Face: huggingface.co/netease-youda…
T3PO 同传
GitHub: github.com/netease-youdao/Co…
Hugging Face: huggingface.co/netease-youda…
Frank Lin retweeted
Qwen-Image-2.1 in 4 steps is here ⚡
@ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model
▶️ on Spaces hf.co/spaces/Viggle/Qwen-Ima…
When several people talk at once, a transcript can get messy fast.
Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
Frank Lin retweeted
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨
A unified model for both generation and editing, delivering top-tier quality in a lightweight package.
Highlights: 👀
- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.
- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.
- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.
- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.
Start to create your next masterpiece with Qwen-Image-2.1! 🖼️
- Blog: qwen.ai/blog?id=qwen-image-2…
- GitHub: github.com/QwenLM/Qwen-Image…
- Model Scope: modelscope.cn/models/Qwen/Qw…
- Hugging Face: huggingface.co/Qwen/Qwen-Ima…
Frank Lin retweeted
🤗 MOSS-TTS-Local Transformer v1.5 is now open source.
Built with a pure autoregressive Audio Tokenizer + LLM paradigm:
>MOSS-Audio-Tokenizer-v2, 2B params
>Qwen3-4B backbone
>Native 48 kHz stereo audio
>Streaming output with theoretical sub-100 ms TTFT
>Zero-shot voice cloning
>Inline [pause] control
>🇺🇸 🇯🇵 🇰🇷 31 language synthesis
>SGLang-Omni Day0 support 🎉 @sgl_project @lmsysorg
Designed for voice agents, digital humans, game NPCs, audiobooks, and real-time speech generation.
👇
Frank Lin retweeted
Just found RelateAnything 🤯
It detects how objects in an image relate to each other, “holding,” “behind,” “sitting on,” etc., using just pixels + bounding boxes.
53M parameters, 19K+ relation words, no retraining, and ~20ms/frame.
Spatial relations are still a weak spot, but this looks really interesting. I’m planning to test it with my own detector outputs and video this weekend.
#MachineLearning #Research #Animals
Frank Lin retweeted
Is Jev is a game changer for robotics?
We gave Jev a robot body and handed it complex tasks across navigation, spatial reasoning, and world geometry
120 different real + simulated tasks and environments benchmarking performance against Dimcode, Astra, Fable, Opus, and 5.6
We graded against speed, cost, tokens, # collisions, path quality
Code, Data, and Paper dropping tomorrow. The results were surprising.
Frank Lin retweeted
Meta’s Segment Anything Model (SAM) 3.1 is now available on Meta Model API, giving developers a fast and lightweight model for detection, segmentation and tracking in a single call on inference tuned for SAM 3.1's architecture.
Use a short phrase to find objects in images and video. One API call returns detections, pixel-precise segmentation masks, and identity-preserving video tracks.
Learn more and start building: bit.ly/4rDArq9
Frank Lin retweeted
We're open-sourcing H3 HyperFlow.
A data-free flow self-distillation technology built on @MiniMax_AI H3. No external training data. Significantly reduces inference cost while preserving frontier model quality.
Full demo: videorebirth.com/lp/hyperflo…