Bernat retweeted
finally releasing our new OCR benchmark
tons of hard examples like this mechanical dial meter
- single value extraction
- JSON data extraction
- document transcription
- text localization and recognition
- 48 models evaluated
↓ examples and leaderboards
Bernat retweeted
This is the doom I predicted a few days ago, coming for Photoshop. A clean-room open-source reimplementation.
No prizes for guessing that they decompiled Photoshop to source code, processed that to some kind of non-code specification language, then fed the spec to an LLM with an instruction to generate Rust.
Adobe just got nuked. And closed source is dead, dead, dead.
github.com/storytold/photocr…
GPT-6 and Intelligent UI, now rolling out in ChatGPT for everyone.
Intelligent UI in ChatGPT delivers fast, interactive answers that make everyday questions more visual, complex topics easier to grasp, and interactive tools for your task available on the spot.
Bernat retweeted
one of my favorite examples from my upcoming VLM OCR benchmark. thanks a lot to everyone who contributed data.
Bernat retweeted
I made a list of great startups to join. It's called the Breakout List.
The list has 92 companies. These are the 20 with 25 or fewer employees:
- Hone (@moritz_stephan, @CarloWillem, @oqbrady)
- Normal (@ansonyuu, @hudzah)
- Standard Intelligence (@G413N, @devanshpandey)
- Tacit Labs (@ninklefitz, @AmDroste)
- American Terawatt (@atroyn, @rslparker, @aranibatta)
- Conduit (@clemvonstengel, @riopopper)
- Convergent (Omkar Savant, Vivek Katara, @debnilsur)
- Core Automation (@MillionInt, @_arohan_)
- Engram (@dan_biderman, @EyubogluSabri, @realJessyLin)
- Instinct (@noahrshinn)
- Keenable (@styskin, Matthias Petri)
- Lumaril (Mark Elliot, Ben Duffield)
- Neion Bio (@Dimkell, Sam Levin)
- Pangram Labs (@max_spero_, @bradley_emi)
- Quadrillion (@echinaceous)
- Re (@karnsaroya, @AnandDhillon, @thecliffwhite, @benaneesh)
- Ricursive (@annadgoldie, @Azaliamirh)
- Sail Research (@neilmovva, @blintzbase)
- Trajectory (@rronak_, @michaelelabd, @QuantumArjun)
- Watney Robotics (Sean Cheong, Ryan Gannon)
Picks from Elad Gil, Charlie Songhurst, Keith Rabois, Mike Vernal, Alana Goyal, Sonya Huang, Ramtin Naimi, Marc Bhargava, Cory Levy, Aashay Sanghvi, Konstantine Buhler, John Luttig, Varun Gupta, Ray Tonsing and Avichal Garg.
Disclosure: I'm a small investor in American Terawatt, Convergent, Standard Intelligence and Trajectory (in this post), and in Factory, Physical Intelligence and SF Compute (elsewhere on the list). I didn't vote.
The full list is on Breakout List.
Bernat retweeted
what in the alien architecture is this
mp.weixin.qq.com/s/qg0NU3NNU…
DeepSeek V4.1 Flash 为 552B 参数的 MoE 模型,采用了全新的 Causal-Encoder-Decoder 结构,输入和输出不对称,输入激活只有 8B,输出激活 16B,成本显著低于已知的同尺寸模型。同时,V4.1 Flash 还采用了新的预训练方式、经过了更大规模的强化学习后训练,在基准测试中,成功超越了包括 DeepSeek V4 Pro 在内的一众旗舰模型的智能水平。
Bernat retweeted
For AI to work with us, it needs to understand us
Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
Bernat retweeted
and here’s @mxsage’s lovely “36 points” running on this 60hz eink panel with the @Modostech glider driver board
Banger paper from Apple.
If you build MCP servers, this can help you turn your specification into an evaluation suite.
(bookmark it)
It's actually a really neat idea that's easy to implement. And it showcases the awesomeness of MCP.
Agent Seer starts from a single MCP spec and synthesizes multi-turn agent test scenarios with no examples, no live tool access, and no domain-specific tuning.
Function names, natural-language descriptions and typed parameter schemas already carry enough semantics to generate graded scenarios with synthetic tool outputs, which then expand into mock-data-grounded dialogues.
Hand-built agent benchmarks demand deep domain expertise, do not scale across tool ecosystems, and go stale as soon as an API changes. Generating them from the live spec keeps pace with the ecosystem instead.
They ran it on seven MCP specifications spanning different domains and suite sizes, with complete tool coverage on small and medium specs.
Parameter schema complexity predicts quality variation far better than tool-suite size does. And argument value accuracy is the dominant failure mode, a sub-dimension that coarse name-match tool-calling metrics cannot see at all.
Paper: arxiv.org/abs/2608.26133
Chat with Paper: academy.dair.ai/papers/agent…
Bernat retweeted
Autoresearch, but optimizing for information gain.
Very interesting that you can hard-code a harness to work on a fairly open-ended discovery problem.
I guess that’s the power of a mathematically grounded harness design.
Predicting the answer to interventional "what if?" questions — the outcome of an action you never took — need a *mechanistic* model, not a curve fit. And you can only learn one by *experimenting*. Experiments are costly, so the real game is **data efficiency**.
Meet the Model Discovery Agent (MDA). 🧵
Bernat retweeted
Deep learning runs on math, but we've never had a formal language for architectures!
Now we do. Our new work is out on @TmlrOrg , led by @Vincent Abbott 🎉
arxiv.org/pdf/2604.07242
Bernat retweeted
Today, we're introducing Claude Fable for Investing
While you sleep, 24/7 AI agents read filings, scrape the internet, and hunt prediction markets to hand you a swipeable deck of research-backed trades by morning
Swipe: opentrade.live
Bernat retweeted
Today we are introducing Dyna-2, a world-action model pre-trained on one million hours of human video. At this scale, for the first time, we discovered several new scaling laws:
• world-action models exhibit scaling law on human data across four orders of magnitude, from 1000 to 1,000,000 hours,
• this human data scaling law implied a scaling law on never seen robot data,
• both data and objective matter; world modeling and scaling on video data are essential for cross-embodiment scaling transfer to emerge
🧵
Bernat retweeted
Korean artist Dakd Jung created a custom Bluetooth speaker featuring a sound-reactive ferrofluid display. The design utilizes an Arduino Nano, an electromagnet, and an audio-frequency analyzer to make a black magnetic liquid dance in real-time sync with music.
Bernat retweeted
🚀Excited to share Qwen-3D (ECCV 2026) — a generalist 3D vision–language model that does grounding, segmentation, VQA, and spatial reasoning in one place. VLMs are great at images and short clips, but long videos bury them in tokens. Our idea: put multi-view frames into a shared 3D world, so attention scales with what’s in the scene, not how many frames you feed it.
📄arxiv.org/abs/2608.02980
🌐qwen-3d.github.io
Bernat retweeted
things we open-sourced lately:
- train & eval in any harness (verifiers v1)
- frontier evals (mazebench, pmpp-hard)
- frontier harness (prime-agent)
- train multi-agent systems (now)
this is the open superintelligence stack.
Today, we’re extending our RL stack beyond individual agents to multi-agent systems.
You can now express arbitrary agent interactions and train them.
primeintellect.ai/blog/multi…
Bernat retweeted
My evaluation lecture! I walk you through different evaluation eras I've been a part of, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes (I expand on agentic more than any other topic, drawing on @xeophon's insights).
This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for.
00:00 Intro: frontier evaluation is harder than ever
03:39 Part 1: The eras of post-training evaluation
17:36 Part 2: An intro to agentic evals
21:19 Part 3: Can you trust the number?
30:41 Takeaways & conclusion
Thanks for watching! Just one more lecture after this :)
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate
Bernat retweeted
The final lecture of my course is an intro to character training! This is a topic that I've been quietly very invested in for ~18 months, as it:
* Has potential for high real world impact
* Clearly used extensively at frontier labs
* Almost no empirical literature exists
* More accessible on academic compute
This lecture covers what character training is, reviews model specs, constitutions, the differences, the motivations in real world events, some example research papers I like, and open questions in how it relates to post-training/model use generally.
Hopefully this brings more people into the field (and reach out if you have questions). It is one of the more research-y chapters in my book, but one that I felt needed the reference. There is still so little, educational content on the topic online.
0:00 Intro
6:22 Part 1: Fundamentals — character, constitutions, and model specs
19:21 Part 2: Character training in practice
23:23 Part 3: Character elicitation without gradient steps
28:03 Part 4: Open questions (and the end of the course)
32:27 The course, complete
Thanks for watching. No need to like and subscribe now that the course is done, you definitely wouldn't!
h/t to @_maiush for leading the technical work I got to do in the space, and @zafstojano for investing a lot of attention at this book chapter.
This video is larger than Cloudflare's 512 MB cache, so it can't be played through. More donations are needed to cover a larger cache. Donate