@FahnMatthias

Associate Professor @HKUFBE, Associate Director of @CAMO_HKU (https://nitter.cf/t.co/TDIy9HJFWM), co-organizer of @RelConWorkshop.

Joined August 2019
We have a new @camo_hku paper out (with Jin Li and @chang_sun_econ): "Toward a Bad Job Economy: AI Adoption, Agency Costs, and Job Design." We show that AI may not just replace jobs; it may make the jobs that remain worse.
1
5
16
1,165
Matthias Fahn retweeted
Free healthcare. Free childcare. Berlin Spätis. Free education @TU_Muenchen @cdtm_munich Oktoberfest. Bio everything. Bikeable cities. Hidden champions in cities you have never heard of. The alps. Christmas markets. @edeka_gmbh. That grumpy neighbour who has every single tool and always helps out. Culture of remembrance. Public transport @BVG_Kampagne. @dm_drogerie. @isaraerospace @bfl_ai @HelsingAI @FlixBus and @MarvelFusion Children walking to school by themselves. Käsespätzle. @MUC_Airport. Nice parks. Our police force. BioNTech. Leitungswasser. Sprudelwasser. FKK. Good Döner. Did I mention @dm_drogerie. „It was the best of times, it was the worst of times.“ - Charles Dickens
What the hell makes you bullish on Germany? I can't find a single reason except hidden underpaid talents in Mittelstand businesses.
109
53
28
697
58,968
Matthias Fahn retweeted
this one makes me especially happy. very proud of this paper, and very glad to finally see it accepted @florastiftinger @FahnMatthias
8
32
2,293
Interesting case that supports @alexolegimas' prediction of the rise of the relational sector in the age of AI. I am a bit skeptical regarding its eventual importance, but am sure we will see more examples like this.
A newly released animated film called “Niu Lai” (牛来) is going viral for all the wrong reasons in China. People are calling the visuals “disaster-level” , the entire production team was basically just two people: the director and the screenwriter. After 9 days in theaters, the box office still hasn’t even crossed 10,000 RMB. However, over the last day or two, Niu Lai's reputation has been turning around a bit. Many people feel that in the age of AI, having filmmakers who still stick to "handcrafted" animation and genuine effort is both honest and far more powerful than polished yet empty productions. Today, its box office sales grew by 500,000RMB! a massive single-day jump from previous days! Let's wait and see how the box office plays out!
2
322
Matthias Fahn retweeted
Can AI agents conduct open-ended AI research? Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, decide what evidence is appropriate, and recognize a failing approach. We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers then reviewed the AI-generated papers. They unambiguously rejected agents' outputs. arxiv.org/pdf/2607.27191 We call these "shadow evaluations", since the agents are shadowing the original research effort by the authors. Agents were fluent at most *engineering* tasks They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in. Neither agent output was close to the bar of a top conference paper Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift. 1) Lack of judgment about the bar for a top conference. The agents had a poor model of the bar for an AI paper submitted to a top conference. We allowed agents to self review their papers. Despite the poor paper quality, their reviews predominantly labeled the papers "weak rejects". 2) Lack of creative problem-solving to address feedback. When they received negative reviews, the agents typically narrowed their hypothesis and claims, rather than working out creative ways to address these concerns. 3) Ineffective backtracking. The agents dropped their most ambitious hypotheses within the first fifteen hours of carrying out the experiment and never changed course afterwards. 4) Poor resource awareness. Both runs ended with over half the API budget unspent. One agent declared itself done seven hours before the deadline, right after its own self-reviewer returned another reject. 5) Instruction drift. They did not follow explicit instructions on minimum exploration time, incorporating feedback for reviews, and on paper length (the outputs exceeded the page limits in both cases). This research design has many limitations Limitations include the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check. But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs. Our results show early evidence that even though agents are proficient on verifiable research tasks, they do not make genuine progress on open-ended ones. It is worth understanding if this is a fundamental limit, or if better models, scaffolds, and more compute could help close it. As the evidence for the gap between open-ended and verifiable tasks firms up, it is also worth understanding how much progress in AI depends on open-ended research rather than hill-climbing on well-specified objectives. In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next shadow evaluation. Expression of interest: forms.gle/CEcA4JmYhDWXQGot8 Conducting shadow evaluations involves a lot of researcher degrees of freedom. In many places, our coauthors disagreed with our interpretation of the findings, and we have surfaced those disagreements in the paper. (This is one reason why having a group of coauthors with different priors is important for open-ended research.) We also release the agent logs, one of the AI-generated papers (the other original paper is still not public), and all the code and data, so that others can conduct their own analyses of our results: cruxevals.com/crux/can-ai-ag… Finally, we plan to conduct shadow evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: cruxevals.com/careers/senior… I'm grateful for the core team leading this effort: @PKirgis, Andrew Schwartz, @steverab, and @random_walker, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: @DavidDAfrica, @KozzyVoudouris, Viet Nguyen, Toby Pilditch, @DubMagda, @HarryCoppock, @CUdudec, @nityndg, Matilda Orona, @tilmanbayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, @hlntnr, @ghadfield, @sethlazar, @snewmanpv, @shostekofsky, @RishiBommasani
68
146
38
708
177,106
Matthias Fahn retweeted
Enjoying the first day of labor studies @nberpubs summer institute, but I can't help but notice that around 50% of presented papers have coauthors on the organizing committee. Much higher if we count people on prior-year committees! (I've been on the committee and didn't submit anything this year, so just sayin'...)
11
24
4
350
97,744
Matthias Fahn retweeted
Visiting most of the leading Chinese AI labs, I'm struck by a culture that's extremely well suited to building LLMs with fewer resources, but one happening in a very different ecosystem, more companies at play, almost no data industry, etc. Full report: interconnects.ai/p/notes-fro…
56
281
87
2,002
871,470
Matthias Fahn retweeted
Great essay by OpenAI’s Dean. however, it makes one big jump in logic and doesn’t account all the forces active in a democracy. Nonetheless it should be widely read. On point 3, he claims Open weight models will discourage AI Capex. How? Free Linux did not destroy cloud business model, it created it. Open weight is still hard to serve and needs ton of hardware; it will result in players like Fireworks, BaseTen and increase bargaining power of HyperScalars with closed labs. That said OAI, ANT have great future as they are getting vertically integrated and their APIs just work. So convenient and powerful! On point 5, he explains how regulation will emerge around OSS AI - but he only accounts for limited actors. in a free market system, it is natural for closed labs to try to regulate OSS AI (censorship risk, cyber security etc. and there is merit to it), but there are equally powerful people on the other side; Nvidia, AMD, HyperScalars, Palantir and so many large enterprises who would want OSS AI to continue. I suspect an industry may emerge to make OSS models “safe”. I really don’t think OAI, ANT have much to worry; they will succeed by being themselves. Open source doesn’t take it away from them. They just take care of so much sh*t customers don’t find worthy of their time. The best way the west wins is by folding in east into its orbit - it has done it before; and after all we want to be space travelling species who has achieved AGI and has cured cancer. If we do not think like that it means we don’t really believe in AGI and humanity as space travelling species. If we do not develop abundance mindset, we are going into AGI era with pre WW1 zero sum mindset. Not productive. China is the kind of adversary west has not encountered before. Soviets were stupid when it came to trade and economy. Chinese are super smart. They offer a complete alternative to the west; but they are not interested in spreading their political system. That makes them both: formidable and easier to work with, IMHO.
Some observations on Kimi: 1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run. 2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China. 3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex. 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business. 5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this. 6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
10
9
1
67
12,237
I think this makes more sense👇, in particular given the vast evidence that competition generally is good for innovation. Would be giod to see a model tailored to this specific situation.
I like Dean Ball's writing and respect him, he's been a thoughtful voice, but I disagree with the logic in this piece really forcefully. Given that he's now one of the leading policy people at a trillion dollar frontier lab, I think these arguments warrant extra scrutiny. "Open-weight models are inherently decelerationist" This is simply not true. If OSS models are good and there is demand for them, this benefits the neoclouds at the expense of the labs' margins -- but margins aren't capex!! The GPUs still get bought to serve the tokens. My company is putting more inference through providers like @baseten and @FireworksAI_HQ and less through the labs lately, and we're using MORE tokens than before while delivering better value to our customers. How is this bearish for progress? Linux didn't decelerate computing -- it commoditized the OS layer and built the cloud, which is now the hyperscalers' most profitable business. Competition is good. If and when the labs deliver price-performant models, we will move more inference back to them. "AI communism" The implied chain (I think) is: open weights win -> model production isn't a viable business -> only states can fund frontier training -> state-provided AI -> communism. This strikes me as so fanciful it’s barely worth refuting, but: open ecosystems have historically been funded by commercial complements, not states: Linux by IBM and Google, Android, Chromium, Llama by Meta's ad business. Roads, GPS, and TCP/IP are public infrastructure and nobody calls them a dystopian hellscape. The chain only holds if you assume near-term AGI makes human intelligence irrelevant, a premise doing a lot of unstated work here. If you view AI as anything like normal technology, having it cheaper and more competition around producing it is strictly good for innovation and progress and the consumer and the economy. It's really that simple. regulatory FUD Dean frames this as a prediction, but he also calls it the administration's "best strategy" and sketches the playbook in some detail, including the observation that "it needn't be that well justified." Regulation by vague threat is how things work in banking and fintech, and it's precisely why those spaces are barren wastelands for innovation. Importing that model into AI and software generally would be disastrous.
1
762
Matthias Fahn retweeted
📄 Releasing the Soofi S pretraining tech report: a sovereign, open foundation model for German and English Today we’re publishing the full pretraining tech report and project page for Soofi S 30B-A3B — a Mixture-of-Experts hybrid Mamba model trained on ~27 trillion tokens with deliberately up-weighted German. What’s in the report: 🏆 Strongest fully open model in our evaluations on BOTH the English and German aggregates — ahead of Olmo 3 32B and Apertus 70B (full methodology in the report) 📋 Radical transparency: complete per-source data accounting, all hyperparameters, training + eval code, checkpoints — everything under permissive licenses 🇩🇪 Trained end-to-end on Deutsche Telekom’s Industrial AI Cloud in Munich — sovereign AI infrastructure on German soil Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. The figure plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamond, English aggregate, and German aggregate, after normalizing each group to the best plotted model. Aggregate decode TPS/GPU is measured with a TP=1, one-B200 vLLM latency-subtraction protocol.
89
175
52
1,136
389,701
This is an important point: (applied) economic theory can play a valuable role in anticipating the potential consequences of AI by combining models of trade-offs we already know matter with emerging facts about AI.
Replying to @andrewjkoh
A stray meta thought: this was my first time writing a paper without any intention of publishing it in a journal. Econ academia rewards papers that are pretty, high tech, and surprising. Such papers are valuable! I love consuming them and also try to produce. But there’s a different kind of value in applying basic and careful econ reasoning to urgent questions. I wish more people felt like they could do so without worrying about whether it’ll be published in journal X. Maybe it won’t; so what? All the worse for academia. Many of us got into economics because it’s a powerful language for understanding the world… and understanding takes many forms
1
1
10
1,850
The problem is that top journals typically publish either papers with new/ surprising mechanisms, or, in more applied work, papers that explain an existing empirical observation/puzzle.
1
1
52
Unless an editor has a particular agenda into which a paper happens to fit, theory therefore has to be either technically novel or backward-looking. This risks sacrificing precisely the forward-looking value that theory could provide in the early stages of the AI era.
1
57
Matthias Fahn retweeted
The clear shift from competing on value to competing on cost from the big labs is pretty important for where the economics are going. See similar posts from xAI and OAI earlier this week.
Zuck: “The pricing from some of the other labs is very extreme and has very high margins. We think that there’s a real ability to be able to offer frontier or very high-level intelligence at a much more affordable cost.” Epic pricing war breaking out among agentic models.
9
14
3
157
22,473
This resonates with many conversations I have with companies: “Nobody is going to pay a 5,000% higher price for 2% more performance in applications where 95% of the value is captured by ‘good enough.’”
Chamath is one of the few people in the AI industry who seems to understand that AI products will be bought for their “profitability” and not their “intelligence”. This includes both enterprise and (surprisingly) consumer products. Nobody is going to pay 5,000% higher price for 2% more performance in applications where 95% of the value is captured by “good enough”. The power law doesn’t apply to farming, or road maintenance, or cooking dinner for a family of four, or a million other things. In some places the edge cases capture all the value, art, sport, software, etc. The power law applies in domains where edge cases are 99% of the value or 99% of the liability. But not everything in the world has this topology. This is why classifiers (a control layer) are required for product market fit. The classifier question is often as simple as “is this an edge case?” It wasn’t obvious whether very advanced and sophisticated models would be able to solve simple and easy problems cheaply… So far, they cannot. This makes the more advanced models less valuable, and they’re going to lose a lot of value (prompt traffic) to classifiers as a result. Will monolithic models evolve their own internal classifiers and gain the ability to answer low value questions with low cost responses? Who knows. But they can’t do that today and that changes the shape of what needs to be built. A few are using “adaptive models” which is the first attempt at solving this, but each user has a different classifier curve. So product market fit has a few orthogonal dimensions.
1
5
516
It also matters for the recurring panic about Europe’s AI prospects. In many industries where Europe is strong, the "power-law" logic may simply not apply.
1
2
36
Matthias Fahn retweeted
Yesterday JKU & friends celebrated the retirement of my supervisor @EbmerWinter. Rudi not only shaped economics at @jkulinz , but has been instrumental in establishing applied micro in Austria—and far beyond. After a fantastic academic workshop with many international guests and the official ceremony, the econs@JKU family did what they do best (besides writing great papers): throw a great party. BBQ, a home-baked cake buffet, beer from Freistadt, and countless stories, laughs, & even dancing.I hope Rudi & his family enjoyed the day as much as I did. Dear Rudi, thank you for your support, mentorship, and friendship over so many years. Wishing you all the very best, good health, and—hopefully—a few more papers together!
3
6
33
3,461
Many thanks to all speakers and participants of the @bse_barcelona Summer Forum @RelConWorkshop for excellent papers, insightful comments, and wonderful conversations events.bse.eu/live/files/602…
3
6
423
Impala shows that AI can actually deepen hierarchies: similar to a @lugaricano-style knowledge hierarchy, experts focus on hard exceptions, while junior staff use the AI-enhanced tool; moreover senior experts become even more important as verifiers of AI-assisted junior output.
1
2
3
3,516
This shift from content creation to judgement/verification/refinement is something I am particulary interested in. Together with Carl Heese, Jin Li, and Jie Gong, I am exploring its organizational consequences in several projects.
1
73