@developerlin

Founder @ startup · AI agents · fine-tuning & post-trained models · agentic video/aigc

Joined October 2007
Frank Lin retweeted
Windows now has sandboxing! Microsoft released an open-source repo, mxc, for sandboxed code execution. We collaborated with Windows to add mxc OS level sandboxing to Unsloth which adds just <100 ms of overhead. GitHub: github.com/unslothai/unsloth Guide: unsloth.ai/docs/new/studio/s…
39
199
26
1,460
55,297
Frank Lin retweeted
I’m obsessed with using tiny classifier models in Pipecat agents. This one, Audience, is a ~33M-param ONNX model built on Ettin that works out who each line in a scene is for. It uses names, descriptions, nearby items and where people are standing, and it can pick one character, several, the whole group, or return "unclear" if you're like me and talk jibberish. It runs in ~6–7 ms on CPU and slots nicely into a Pipecat pipeline. Multi-character agents are really fun! Weights here: huggingface.co/spellspeak/au…
2
2
3
23
3,628
Frank Lin retweeted
finally releasing our new OCR benchmark tons of hard examples like this mechanical dial meter - single value extraction - JSON data extraction - document transcription - text localization and recognition - 48 models evaluated ↓ examples and leaderboards
19
10
1
169
7,422
Mobile manipulation policies can break from base pose errors of just a few cm, which are common after navigation. Can we get pose generalization without additional demos? Introducing MobileVISTA: collect demos at one base pose, get a policy that works from many 🧵👇 Website: sdwistreich.github.io/mobile… Paper: arxiv.org/abs/2610.07511
12
17
2
85
4,234
Wow Nvidia has just released a tool to turn any image into an explorable 3D world And they made it 100% open source 🔥 - Model available on Hugging Face - UI code available on GitHub Lyra 2.0 creates a world you can walk through, turn around in, and even drop a robot into for simulation.
41
277
24
2,344
150,017
Video model can not do this
Everyone is testing Opus 5.5 to create impressive videos. The agent's ability and outcome is even more surprising than using a video‑gen model. Just as our OneVision‑Encoder was recently accepted by NeurIPS, we would like to produce a promotional video to showcase what we consider the most elegant MLLM visual‑compression solution. Accordingly, we provided Opus with our two papers (OV‑Encoder and OV‑2). It came back with a 3-minute launch film: every frame rendered from code, every note synthesized, every number traceable to the paper, and even the music is directly from code and with matched beats. What a surprise to me!!! I was slumped on the sofa as if I had witnessed an atomic bomb explode😜. The two papers👇
18
Frank Lin retweeted
Who wants more speech training data? YODAS v3 is now available @huggingface It’s 1.1M hours - the biggest audio dataset ever. And the first at this scale with stereo audio at 48kHz, with timestamped transcripts and translations. 100+ langs hf.co/blog/espnet/yodasv3 CC-BY-3.0
23
73
13
459
49,106
Frank Lin retweeted
Impressive gift for UI/UX designers! Ming-Image Design - 6B open-source T2I for UI, infographics, posters & text-rich designs; - full compositions with RGBA transparent backgrounds - pairs with Design-Layer for editable transparent layer decomp; - 2048 x 2048 + VLM prompt enhancer. huggingface.co/inclusionAI/M…
We’re open-sourcing the Ming-Image-0.1-Design family: • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. 🧵
1
14
1
121
26,771
Frank Lin retweeted
I asked Opus 5.5 to explain camera focus by building an interactive lens lab Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost lens.lab.sael.net Move the focus ring and you can see the glass elements shift the sharp plane through the scene
391
1,087
291
16,293
3,517,802
Another insane Jev use case! Jev makes it incredibly cheap to evaluate and classify agent runs at scale. And finally, someone open-sourced a self-improving memory layer that can put that capability to work across agent harnesses. It turns your agent sessions into a compounding knowledge layer, where every successful run can make future agents smarter across: - Codex - Claude Code - Cursor - OpenCode and 20+ more Beacon by @asymptotelabs continuously builds a shared history across your agent harnesses and uses Jev to identify the runs worth learning from. It then turns the best workflows, corrections, and debugging patterns into reusable skills. GitHub repo: github.com/Asymptote-Labs/ag…. (don’t forget to star it ⭐) Most agent runs are messy. They contain exploration, failed commands, dead ends, and one-off fixes that should never become permanent memory. So Beacon preserves the full session history, while Jev helps decide what should be promoted, reviewed, or discarded. The recording below shows this in action. Beacon found 579 sessions across 5 coding-agent harnesses and normalized them into one consistent history. From there, Jev surfaces the lessons worth keeping and makes them available across your agent stack. - A pattern learned in Cursor can carry into OpenCode. - A lesson from Claude Code can improve the next Codex run. Every successful run adds to the shared knowledge layer, making future agents smarter. If you want to dive deeper into Jev, I also wrote a breakdown of how it works. The article is quoted below.
52
114
10
896
119,750
Frank Lin retweeted
卧槽,国产开源这次真把实时同传做出来了。😲 看了雨哥这条,我深度评测了网易有道的这两个模型。R2T2实时听写模型,T3PO 实时翻译。 🎉牛逼的是,两个模型一上来就把 Hugging Face 各自的 Trending 分区第一拿了: R2T2:ASR #1 T3PO:Translation #1 🔥更骚的是,这两个都能 Streaming,而且还能直接串起来。 怎么记住这两个模型,一个是R开头,Realtime,一个是T开头,Translation。后面阅读就会顺一些。 对面的话还没说完,中文字幕已经出来了。所以我直接搞了个很抽象的 Demo: 😍如果邻居娶一个乌克兰老婆,用英语连续跟他吵 30 秒。 中间老公回嘴、老婆切俄语、再切乌克兰语,我作为隔壁老王插话。 乌克兰语,它是没训练过的,我来盘盘它 😂 ASR+Translation, 除了视频字幕,以后直接挂在电脑上,挂录音豆里,挂手机上,挂AI眼镜上: 看剧、Zoom、Meet、Teams、Discord、跨国电话…… 对面老外说话,你屏幕下面直接出中文。 出国泡吧搭讪也更方便。 下面视频就是完整实测。 R2T2|实时 ASR huggingface.co/netease-youda… T3PO|实时翻译 huggingface.co/netease-youda…
太牛逼了!Hugging Face 这周跑出一匹黑马,还是同时跑赢两个赛道的那种 语音识别 Trending 榜第一,压着 OpenAI 的 whisper-large-v3、NVIDIA 的 Nemotron、阿里的 Qwen3-ASR。翻译 Trending 榜第一,压着腾讯的 Hy-MT2、Cohere、Google、Meta 的 NLLB。两个第一都是同一家中国公司,网易有道,很多人印象里还是那个做词典的。模型叫 R2T2 和 T3PO,一个听,一个译,全开源。榜单截图我今天截的,还在第一。 一家原本在学习场景出挑的公司,在 HF 上把两个垂类做到第一,这事本身就值得点开看看。但我更想知道它到底强在哪,所以没在 GPU 上跑,把两个都塞进了一台 M1 Pro 的 MacBook,不联网。把我自己以前的口播按真实时间喂进去,中文实时上屏,英文每说完一句就跟出来。视频里的每个数字都是这台机器上的真实数字。 先说我最在意的一件事 同一段我自己的口播,8.6 秒,每 640 毫秒出一次结果。左边 Whisper 走"每次把收到的音频整段重解一遍"的伪流式,13 步里改写了 7 次,屏幕上已经显示的 93 个字被重写过,"见面了"变成"在这边了"又变回来,最后一句停在"全程媒介一找"。右边 R2T2 是 0 次改写,提交出来的字一个没动过。 更有意思的是 8.32 秒那一帧。R2T2 内部其实已经猜出"全程没接",但它只提交了"全程"两个字,把"没接"压着没放。0.6 秒后证据够了,提交的是"没截",对的。它不是猜得慢,它是知道这一刻不该猜。 为什么我在意这个。给人看字幕,闪一下无所谓。但这段文字如果是喂给下一个模型的,比如同传、比如一个要执行动作的 Agent,前面的字一改,下游就得撤销自己刚做的判断。视频第三段就是这条链路,R2T2 每提交完一句,原文直接交给 T3PO,中间没有人工,也不用等整段说完。因为 R2T2 给出来的字不会再改,T3PO 拿到就能翻,翻出来的英文也不会再改。 热词那段也值得看。不给上下文的时候,它把 Codex 写成 codet,Excalidraw 写成 escalator,无限画布写成无线画布。给一行热词之后三个全对。但我的名字"雨哥",热词给了,离线模式还是写成"宇哥",流式那一遍倒是对了。同音人名是最难扳的一类,别指望热词包治百病。 几个要说清楚的地方。 一,R2T2 是在 Qwen3-ASR 1.7B 上做的,方法在数据构造和一个叫 Longest Stable Prefix 的训练范式上,技术报告还没出。 二,Mac 上跑的是社区 audio.cpp 的移植版,9 月 19 号刚合进主仓库,brew 装的二进制里还没有,得自己编。这版每一步都重编码整句音频,句子越长越慢,M1 Pro 上 640 毫秒步长能跟上实时,官方 GPU 路线是 160 毫秒。 三,ASR 加 14B 同传同挤一块 M1 Pro 的 GPU。我一开始让 T3PO 边听边探测,结果它把 GPU 吃满,识别反而被拖到落后 3 秒。改成识别优先、每句提交完再交给 T3PO,识别回到实时,英文在每句说完后两三秒整段出来。T3PO 本身支持逐字喂、边听边译,但那需要两张卡,不是一台笔记本能同时伺候的。 四,代码 Apache 2.0,权重是网易自己的 Model License,商用前看一眼。 只增不改这条路线 R2T2 不是第一个走的,Nemotron、Voxtral 都在做。它把这条路线的准确率追到了接近伪流式和离线的水平,然后开源了,这是我觉得它值得开源社区接住的原因。 过去两年大家盯着中国开源模型的目光都在通用大模型上。这次是两个垂类,实时语音识别和同传,分别把 OpenAI、NVIDIA、Google、腾讯的模型压在榜下,而且是一家平时不怎么在这个圈子出声的公司做的。垂类这条路,中国公司是真能跑到第一的! 两个模型都开源了,代码和权重在这: R2T2 语音识别 GitHub: github.com/netease-youdao/Co… Hugging Face: huggingface.co/netease-youda… T3PO 同传 GitHub: github.com/netease-youdao/Co… Hugging Face: huggingface.co/netease-youda…
50
142
2
753
70,970
Frank Lin retweeted
Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces hf.co/spaces/Viggle/Qwen-Ima…
12
114
6
1,359
95,970
Frank Lin retweeted
When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
207
863
240
9,731
1,062,936
Frank Lin retweeted
This WAN 3.0 feature is insane 😭 It changes the camera angle while keeping the exact same action of original video. Steal my prompt 👇
25
91
9
1,170
64,025
Frank Lin retweeted
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: qwen.ai/blog?id=qwen-image-2… - GitHub: github.com/QwenLM/Qwen-Image… - Model Scope: modelscope.cn/models/Qwen/Qw… - Hugging Face: huggingface.co/Qwen/Qwen-Ima…
301
888
452
8,304
2,392,346
Frank Lin retweeted
🤗 MOSS-TTS-Local Transformer v1.5 is now open source. Built with a pure autoregressive Audio Tokenizer + LLM paradigm: >MOSS-Audio-Tokenizer-v2, 2B params >Qwen3-4B backbone >Native 48 kHz stereo audio >Streaming output with theoretical sub-100 ms TTFT >Zero-shot voice cloning >Inline [pause] control >🇺🇸 🇯🇵 🇰🇷 31 language synthesis >SGLang-Omni Day0 support 🎉 @sgl_project @lmsysorg Designed for voice agents, digital humans, game NPCs, audiobooks, and real-time speech generation. 👇
7
22
7
118
100,508
Just found RelateAnything 🤯 It detects how objects in an image relate to each other, “holding,” “behind,” “sitting on,” etc., using just pixels + bounding boxes. 53M parameters, 19K+ relation words, no retraining, and ~20ms/frame. Spatial relations are still a weak spot, but this looks really interesting. I’m planning to test it with my own detector outputs and video this weekend. #MachineLearning #Research #Animals
5
34
3
385
15,461
Frank Lin retweeted
Is Jev is a game changer for robotics? We gave Jev a robot body and handed it complex tasks across navigation, spatial reasoning, and world geometry 120 different real + simulated tasks and environments benchmarking performance against Dimcode, Astra, Fable, Opus, and 5.6 We graded against speed, cost, tokens, # collisions, path quality Code, Data, and Paper dropping tomorrow. The results were surprising.
26
56
5
493
49,112
Meta’s Segment Anything Model (SAM) 3.1 is now available on Meta Model API, giving developers a fast and lightweight model for detection, segmentation and tracking in a single call on inference tuned for SAM 3.1's architecture. Use a short phrase to find objects in images and video. One API call returns detections, pixel-precise segmentation masks, and identity-preserving video tracks. Learn more and start building: bit.ly/4rDArq9
53
259
93
2,107
570,429
We're open-sourcing H3 HyperFlow. A data-free flow self-distillation technology built on @MiniMax_AI H3. No external training data. Significantly reduces inference cost while preserving frontier model quality. Full demo: videorebirth.com/lp/hyperflo…
31
56
11
533
935,424