@DocXavi

Applied scientist at @amazonscience Barcelona, Catalonia. Made at @la_upc & @columbia. Promoting @dlbcnai. Opinions my own.

Badalona, Catalonia
Joined July 2012
X and @elonmusk have failed into promoting the values of democracy and human rights. Time to leave this platform. We learned a lot here, thanks to those who made it possible. Find me on LinkedIn and Bluesky.
La #UPC deixa de publicar a X per mantenir la seva comunicació en entorns que garanteixin la qualitat i la veracitat de la informació. Una decisió que ha pres per consens el #ConsellGovernUPC, el 19 de febrer. 🔗upc.edu/ca/sala-de-premsa/no…
689
📣 Amazon Research Awards fall 2026 call for proposals is now open for submissions. Successful applicants will receive unrestricted funds, AWS promotional credits, and training resources. Research areas: • Agentic AI • AI for Information Security • Amazon Security • Automated Reasoning • AWS Cryptography • Build on Trainium • Multilingual and Contextual AI Evaluation • On-Device Multimodal Perception • Synthetic Data Generation for Multimodal Perception Deadline: November 4 #AmazonResearchAwards amazon.science/research-awar…
1
18
1
141
14,412
Antonio M. López from the Computer Vision Center (@CVC_UAB) will be an invited speaker at #DLBCN 2026. scholar.google.com/citations…
1
2
4
243
Xavi Giró retweeted
Accepted at #NeurIPS26. See you in Paris (hopefully).
Diffusion model ✅ Generated item ✅ Successful item ✅ ➡️ Now: Which training data was influential? ➡️ Which are the items that, without them, such generation would not have been there? This is the counterfactual question our latest paper tries to answer more accurately. 1/5
1
7
2
41
2,023
I’m 36. Solo founder from Catalonia. Wish all expats speak catalan and learn about our traditions. Live is better when you get out from expats bubble. With love from Premià de Mar.
I'm 29. Solo founder from Brazil, based in Barcelona. Looking to connect with more marketers & indie hackers!
32
9
2
282
37,842
Do authors understand their own papers? Our desk rejection rate at TMLR is about 53% while it was only 6% in 2023 To understand the situation and assess our process, Nihar asked authors about their own TMLR submissions Learn the findings of the pilot study & share your thought
TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission? medium.com/@TmlrOrg/asking-a…
2
21
2
187
25,935
What is the role of academic computer vision research in the age of increasingly powerful large models? Is GPT-6 Astra a step change? How can a researcher have an impact today in academia? These are the questions I ask myself as I head off to ECCV 2026, a conference I’ve attended since 1992. One of my papers this year is VIGA, a method that takes an image as input and outputs a 3D Blender scene that represents that image. This is a classical inverse-graphics task and VIGA was the first method to solve it using an agentic approach. The idea is now several years old and the first version of the paper was rejected. This delayed publication significantly. After it was accepted at ECCV, it was quickly surpassed by people using Claude Code for the same purpose. Today GPT-6 Astra blows away all previous results. But we still head off to ECCV to tell the community about our invention that is now fully out of date. The way academic work often progresses is that one reads recent papers, notices that they have limitations, comes up with a new idea, explores this, publishes it, etc. Any published paper I read today is based on ideas that are at least a year old. And those ideas were based on the literature of the time, which was also a year old. That means that any paper I see at ECCV is likely two years out of date. In AI today, two years means your work is likely irrelevant. At CVPR this summer I noticed that many authors have not gotten the message. They continue to work on “old” problems that have a long history. This history is based on assumptions about how the “vision problem” will be “solved”. The truth is that it is being solved in a very different way and many of these problems are no longer relevant. Another group of papers focuses on very niche problems where large models likely fail because of insufficient data or lack of business interest. The impactful papers were largely from industry and had long author lists and massive data+compute behind them. These papers were also out of data, describing systems that had been released months before, but at least they served to provide the community with more complete documentation and analysis of commercial systems. So what should academics do? First, we need to put aside the tools we’ve used for years and start from scratch. Every project should start by trying really hard to solve the problem with existing tools. I would like to see every paper begin with a detailed experimental analysis of how existing models perform and why they fail (if they do). This gives the kind of insight we need today. Then, assuming current models fail, the solution should provide some fundamental insight that will outlive the next release of such models. Reviewers today still focus on technical novelty. This pushes people to focus on tweaking architectures rather than clearly moving the field forward. Papers need to be judged based on their novel insight and not their novel technical contribution. This is a real shift in thinking but it focuses us on what matters - progress of the field. If we want there to be a “field” of computer vision, then it can’t become a marginal backwater, focusing on esoteric problems. If you haven’t tried using Astra (or whatever comes next) to solve your problem, then you have not done your homework. This omission should be seen as negatively as not having a previous work section. Concretely, I think papers should include a new section analogous to “Related Work” where that related work is current models and how they perform on the task. Reviewers should start asking for this and expecting authors to be able to articulate their insights about the limitations of existing large models. I'm interested in your thoughts.
91
390
67
2,366
676,562
Xavi Giró retweeted
My journey to develop AGI spans 25 yrs, including 10+ yrs thinking about technical & societal perspectives at Google DeepMind. AGI is on the horizon - we need deeper understanding of its implications. To help, we've created the DeepMind Institute.
152
527
86
3,713
521,963
BREAKING: OpenAI might have stolen another major proof. In a detailed Mastodon post, which I report in full in the comments, Andreas Thom presents several pieces of evidence suggesting that OpenAI may have trained Astra on conversations in which he and Gábor Kun were working on Gromov’s soficity conjecture, one of the ten problems OpenAI later announced Astra had solved. I know Andreas. We met several times early in our careers. He is an exceptional mathematician, a leading expert on sofic and hyperlinear groups, and one of the most respected scholars in the field. He has spent two decades working on this problem. If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself. And if the allegations raised by Levent Alpöge, Tristan Buckmaster, and now Andreas Thom are all substantiated, we are no longer looking at isolated incidents. We may be looking at one of the greatest intellectual scandals in the history of science. AI is not discovering new mathematics. AI is stealing human discovery.
671
4,832
787
19,795
1,805,937
Xavi Giró retweeted
What a paradigm shift for TTS, going end to end and breaking many inductive biases in speech synthesis: waveform level (no vocoding), categorical (no regression), dil. causal convnet (no RNN) 🤷‍♂️ And lucky me to witness the demo from @heiga_zen in the SSW9 right after unveiling
10 years today since we unveiled WaveNet! Autoregression with long context before it was cool😅 Maybe it looks a bit silly now that we didn't use a Transformer, but we had a good reason: it would take another year for that to be invented🙃 Thanks @heiga_zen for the reminder!
1
1
16
2,705
Xavi Giró retweeted
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
249
1,080
223
12,232
1,727,164
in view of OpenAI's claim about Navier Stokes, it is important that we hear alternative points of view on how it happened. From Tristan Buckmaster, Prof. at NYU. cims.nyu.edu/~tristanb/state…
3
43
3
479
73,199
Xavi Giró retweeted
Super happy to share our intention to join forces with NVIDIA in a $12,930,300,000 acquisition 💛💚 10 years after starting Hugging Face, open-source AI is at an inflection point. Thanks to the community, we’ve shown that it can be a complement, and even an alternative, to closed-source APIs. But for it to happen at larger scale, it needs more compute, more support, more collaboration and more visibility. That’s why we went to talk to Jensen, who offered to do exactly that with us. In addition to doubling down on NVIDIA’s massive contributions to open-source AI (I called them the “King of American open-source AI” earlier this year), they’ve committed to strongly supporting Hugging Face and our mission while keeping the platform open, independent and compute agnostic. The founders and the team are all staying to keep pushing this mission forward. Together, we think we can make open source the default way to build AI, with the goal of empowering 100 million AI builders to own their intelligence rather than rent it. Excited about the next 10 years! 🤗🤗🤗
984
1,196
419
13,421
1,515,855
Xavi Giró retweeted
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
776
3,495
1,534
32,429
10,831,813
When LLM judges agree, the right question is why. Shared prompts, model families, or training lineage can make a majority look stronger than it is. Dependence-aware aggregation via Ising models accounts for this, improving accuracy 9–14% over weighted majority vote. amazon.science/blog/when-llm…
5
14
2,942
Xavi Giró retweeted
New blog post: continuous diffusion for language is back! This research direction receded into the background for a while, but as of this year, it is once again a hot topic. I wrote down a historical perspective and some thoughts on the recent revival. sander.ai/2026/08/24/continu…
22
137
20
769
92,798
Joan Serrà is a Program Chair of #DLBCN 2026.
Here we advance in the unification of Version and Track ID tasks (cover songs and fingerprinting). Intuitively, all these music retrieval tasks should share a single embedding space, but top performance requires some tweaks.
2
4
371
We put the stage and audience, but we need the main actors. Submit your top AI science paper for presentation, and help building a stronger community in Barcelona:
Call for presenters is published. New in 2026: Only peer-reviewed publications will be accepted. We can no longer accommodate posters from preprints and/or ongoing doctoral research. Submission deadline: October 17, 2026. sites.google.com/view/dlbcn2…
2
7
1,246
Xavi Giró retweeted
Excellent talk by @thjashin: - What separates autoregression and diffusion? - Bridging the gap between continuous and masked discrete diffusion via loss reweighting - Insertion-based sequence generation Very clear and easy to follow! (via @mblondel_ml) youtube.com/watch?v=kMimQxIJ…
5
65
3
425
48,644
Xavi Giró retweeted
Ten years ago, I crossed paths with a guy on a remarkable journey. When I told him of my desire to understand the world, he invited me along, suggesting that I might find my answers where he was going. I soon realized I had joined a true polymath who had assembled an unparalleled vessel of minds - thinkers, builders, philosophers, and artists. As he transitions to the next phase of this important mission, I am left with nothing but gratitude for the profound adventures we've shared. Thank you, my friend @demishassabis, for what you have built and for the path ahead.
12
29
2
972
68,333