@monkeymktplacei
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Manhattan, NY
Joined March 2016
- Tweets667
- Following1K
- Followers137
- Likes1.2K
whatever Google is doing with compute, what we know is that they nerfed even the flash models in chat. you can talk about something, comparing items A and B, then next message ask for another attribute for comparison, and it would ask "what are you comparing?". Going back several messages seem to literally reset the context? it would forget everything that had been discussed thus far.
hi @Google whats the reason why you do not surface the map & location information when its in AI search mode? e.g. if i look up a restaurant, in classic search, this would bring up the maps modal to show me where it is and information such as operating hours, etc, saving it for future "want to go". but in the AI search results, this stop showing up. is this intentional design choice? if so, was this to increase friction and thereby increase user use time? because this is one of the worst ux ive ever experienced, a significant downgrade of ux compared to classic search.
hard to say how well the 3.8 model can handle long running research tasks or more ambiguous stuff. but so far it’s pretty good.. it’s obviously not SOTA from the way it approaches general context gathering or (lack of) checking itself, but for a flash model i’m quite impressed by the performance and how it’s going through my personal evals. i think it’s definitely on the pareto frontier. definitively better than terra and sonnet and opus5
one last comment on this thread. it’s remarkable that some variant of flash model is served across almost all consumer facing google services, even if it’s US only. i don’t believe 3.8 is the model they serve in workspace / gmail / search yet, but wow, if they can, they are absolutely, without a doubt, at the frontier in a way no other lab can compare.
strangely, i actually think muse spark is benchmaxxed, if you use muse code you would find that its actually not that well rounded and performance is definitely not same tier as Sol.
gemini 3.8 flash is more preferable for me.
thanks @JeffDean for everything! So much of what I learned in school were what you worked on. You inspired me to always think bigger and challenge the working paradigm, and always look for ways to do something better! Look forward to you & rest of the discovery loop team’s next big thing sir!
hi @GeminiApp, spark and regular chat differentiation is strange, why have them separated and give different capabilities if model is the same? spark is nice why not use this as default? also don’t call it a task? why specify that? so confusing
tbh, i think the problem with Uber is not just AVs, but also how AI will change people's user interfacing with their phones. Uber is an app for completing a goal. In a scenario where I just want to call a ride home, do I really care that I get Uber or something else?
But if interface change, then you can be sure there's going to be demand allocated if you just keep your cost down. This ends up being a race to the bottom in terms of margins. And margin for an AV is without a doubt going to be lower compared to human drivers.
youtube.com/watch?v=8hfpLa5w… - no comments the personnel changes being positive / negative, but 3 / 4 folks from this video recorded just 2 months ago are no longer at Google
idk, obviously everyone is selling and from latest round of earnings call it sounds like it’s about how you market the capex, but kinda seems wild that reaction to google and meta are negative? meta especially if zuck is claiming some shit like 1p roic will be > selling compute?
no idea why people are trashing google when the information coming out all seem to point to strong performance from them?
model sucks due to 1. their focus on serving across their surfaces well ( speed and cost and ux ). 2. not enough capacity it seems?
not enough capacity is due to a boom in gcp, which is positive.
frontier (coding) model development will remain challenging for folks not getting good user data, but distillation and efficiency optimizations are core to google DNA, and recent models - tho not just simply distilled, is still largely relying on distillation, so feel like they should never be that far behind as long as they are not compute constrained?
nothing against josh, but gemini product team is honestly so behind if they are asking and getting these kind of answers. it’s like they don’t use the app at all and have no idea what they have on their own property is broken. let alone advanced features like going back in chat or integration with other products through mcp / agentic capabilities. or personalization doesn’t work or why do we have notebooks / projects when they aren’t able to reference past chats properly.
the model is bad, the product is bad.
on the other hand i have heard multiple folks telling me gemini and its related teams are giving sky high offers for just swe positions.
What's something that you're surprised @GeminiApp can't do well, and we should have fixed a long time ago?
i dont rely on googles antigravity much for anything outside of coding, but wow this is such a shitty experience using it for financial modeling. 1. it generate worse off models requiring a lot more manual refinements from me. 2. it hallucinate like crazy - its as if it doesnt
write permission only read permission.
all of this feels completely unacceptable from the end user POV. @GeminiApp