Cultura & tecnología

Joined October 2021
Bernat retweeted
finally releasing our new OCR benchmark tons of hard examples like this mechanical dial meter - single value extraction - JSON data extraction - document transcription - text localization and recognition - 48 models evaluated ↓ examples and leaderboards
29
42
1
454
25,262
This is the doom I predicted a few days ago, coming for Photoshop. A clean-room open-source reimplementation. No prizes for guessing that they decompiled Photoshop to source code, processed that to some kind of non-code specification language, then fed the spec to an LLM with an instruction to generate Rust. Adobe just got nuked. And closed source is dead, dead, dead. github.com/storytold/photocr…
630
1,454
581
16,766
3,897,294
Bernat retweeted
GPT-6 and Intelligent UI, now rolling out in ChatGPT for everyone. Intelligent UI in ChatGPT delivers fast, interactive answers that make everyday questions more visual, complex topics easier to grasp, and interactive tools for your task available on the spot.
1,208
1,828
1,464
23,261
5,966,784
Bernat retweeted
one of my favorite examples from my upcoming VLM OCR benchmark. thanks a lot to everyone who contributed data.
anyone working super hard OCR problems? I’m building next version of my OCR benchmark and could really use some help I need: - handwritten text - printed documents hand corrections - technical drownings - anything you think is hard
34
26
3
602
37,717
I made a list of great startups to join. It's called the Breakout List. The list has 92 companies. These are the 20 with 25 or fewer employees: - Hone (@moritz_stephan, @CarloWillem, @oqbrady) - Normal (@ansonyuu, @hudzah) - Standard Intelligence (@G413N, @devanshpandey) - Tacit Labs (@ninklefitz, @AmDroste) - American Terawatt (@atroyn, @rslparker, @aranibatta) - Conduit (@clemvonstengel, @riopopper) - Convergent (Omkar Savant, Vivek Katara, @debnilsur) - Core Automation (@MillionInt, @_arohan_) - Engram (@dan_biderman, @EyubogluSabri, @realJessyLin) - Instinct (@noahrshinn) - Keenable (@styskin, Matthias Petri) - Lumaril (Mark Elliot, Ben Duffield) - Neion Bio (@Dimkell, Sam Levin) - Pangram Labs (@max_spero_, @bradley_emi) - Quadrillion (@echinaceous) - Re (@karnsaroya, @AnandDhillon, @thecliffwhite, @benaneesh) - Ricursive (@annadgoldie, @Azaliamirh) - Sail Research (@neilmovva, @blintzbase) - Trajectory (@rronak_, @michaelelabd, @QuantumArjun) - Watney Robotics (Sean Cheong, Ryan Gannon) Picks from Elad Gil, Charlie Songhurst, Keith Rabois, Mike Vernal, Alana Goyal, Sonya Huang, Ramtin Naimi, Marc Bhargava, Cory Levy, Aashay Sanghvi, Konstantine Buhler, John Luttig, Varun Gupta, Ray Tonsing and Avichal Garg. Disclosure: I'm a small investor in American Terawatt, Convergent, Standard Intelligence and Trajectory (in this post), and in Factory, Physical Intelligence and SF Compute (elsewhere on the list). I didn't vote. The full list is on Breakout List.
47
88
55
1,408
580,090
what in the alien architecture is this
mp.weixin.qq.com/s/qg0NU3NNU… DeepSeek V4.1 Flash 为 552B 参数的 MoE 模型,采用了全新的 Causal-Encoder-Decoder 结构,输入和输出不对称,输入激活只有 8B,输出激活 16B,成本显著低于已知的同尺寸模型。同时,V4.1 Flash 还采用了新的预训练方式、经过了更大规模的强化学习后训练,在基准测试中,成功超越了包括 DeepSeek V4 Pro 在内的一众旗舰模型的智能水平。
46
101
21
1,801
258,078
Bernat retweeted
For AI to work with us, it needs to understand us Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
88
143
122
1,416
560,671
Bernat retweeted
and here’s @mxsage’s lovely “36 points” running on this 60hz eink panel with the @Modostech glider driver board
here’s a quick demo of a macos writerdeck i’m throwing together with this 60fps eink display. you can see more of the ghosting in this demo, but i feel it works well enough for light typing and browsing.
80
457
46
5,949
396,834
Bernat retweeted
Banger paper from Apple. If you build MCP servers, this can help you turn your specification into an evaluation suite. (bookmark it) It's actually a really neat idea that's easy to implement. And it showcases the awesomeness of MCP. Agent Seer starts from a single MCP spec and synthesizes multi-turn agent test scenarios with no examples, no live tool access, and no domain-specific tuning. Function names, natural-language descriptions and typed parameter schemas already carry enough semantics to generate graded scenarios with synthetic tool outputs, which then expand into mock-data-grounded dialogues. Hand-built agent benchmarks demand deep domain expertise, do not scale across tool ecosystems, and go stale as soon as an API changes. Generating them from the live spec keeps pace with the ecosystem instead. They ran it on seven MCP specifications spanning different domains and suite sizes, with complete tool coverage on small and medium specs. Parameter schema complexity predicts quality variation far better than tool-suite size does. And argument value accuracy is the dominant failure mode, a sub-dimension that coarse name-match tool-calling metrics cannot see at all. Paper: arxiv.org/abs/2608.26133 Chat with Paper: academy.dair.ai/papers/agent…
41
56
2
507
39,204
Autoresearch, but optimizing for information gain. Very interesting that you can hard-code a harness to work on a fairly open-ended discovery problem. I guess that’s the power of a mathematically grounded harness design.
Predicting the answer to interventional "what if?" questions — the outcome of an action you never took — need a *mechanistic* model, not a curve fit. And you can only learn one by *experimenting*. Experiments are costly, so the real game is **data efficiency**. Meet the Model Discovery Agent (MDA). 🧵
19
41
597
65,741
Deep learning runs on math, but we've never had a formal language for architectures! Now we do. Our new work is out on @TmlrOrg , led by @Vincent Abbott 🎉 arxiv.org/pdf/2604.07242
36
165
16
1,237
81,895
Today, we're introducing Claude Fable for Investing While you sleep, 24/7 AI agents read filings, scrape the internet, and hunt prediction markets to hand you a swipeable deck of research-backed trades by morning Swipe: opentrade.live
36
17
12
189
54,489
Nice use of the extra tokens:
I spent some time training Glimmer to understand Human Motion. It's a fun model!
2
9
1
281
57,527
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws: • world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours, • this human data scaling law implied a scaling law on never seen robot data, • both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge 🧵
194
719
315
3,149
1,814,860
Korean artist Dakd Jung created a custom Bluetooth speaker featuring a sound-reactive ferrofluid display. The design utilizes an Arduino Nano, an electromagnet, and an audio-frequency analyzer to make a black magnetic liquid dance in real-time sync with music.
21
116
28
652
44,650
Bernat retweeted
🚀Excited to share Qwen-3D (ECCV 2026) — a generalist 3D vision–language model that does grounding, segmentation, VQA, and spatial reasoning in one place. VLMs are great at images and short clips, but long videos bury them in tokens. Our idea: put multi-view frames into a shared 3D world, so attention scales with what’s in the scene, not how many frames you feed it. 📄arxiv.org/abs/2608.02980 🌐qwen-3d.github.io
9
75
5
560
57,063
Gen UI thinking..
130
190
29
4,388
317,627
things we open-sourced lately: - train & eval in any harness (verifiers v1) - frontier evals (mazebench, pmpp-hard) - frontier harness (prime-agent) - train multi-agent systems (now) this is the open superintelligence stack.
2
13
1
259
16,729
My evaluation lecture! I walk you through different evaluation eras I've been a part of, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes (I expand on agentic more than any other topic, drawing on @xeophon's insights). This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for. 00:00 Intro: frontier evaluation is harder than ever 03:39 Part 1: The eras of post-training evaluation 17:36 Part 2: An intro to agentic evals 21:19 Part 3: Can you trust the number? 30:41 Takeaways & conclusion Thanks for watching! Just one more lecture after this :)
23
68
3
602
39,567
The final lecture of my course is an intro to character training! This is a topic that I've been quietly very invested in for ~18 months, as it: * Has potential for high real world impact * Clearly used extensively at frontier labs * Almost no empirical literature exists * More accessible on academic compute This lecture covers what character training is, reviews model specs, constitutions, the differences, the motivations in real world events, some example research papers I like, and open questions in how it relates to post-training/model use generally. Hopefully this brings more people into the field (and reach out if you have questions). It is one of the more research-y chapters in my book, but one that I felt needed the reference. There is still so little, educational content on the topic online. 0:00 Intro 6:22 Part 1: Fundamentals — character, constitutions, and model specs 19:21 Part 2: Character training in practice 23:23 Part 3: Character elicitation without gradient steps 28:03 Part 4: Open questions (and the end of the course) 32:27 The course, complete Thanks for watching. No need to like and subscribe now that the course is done, you definitely wouldn't! h/t to @_maiush for leading the technical work I got to do in the space, and @zafstojano for investing a lot of attention at this book chapter.
31
91
17
868
57,584