daily AI productivity tips that work in real life. the apps, the shortcuts,honest reviews for people who want to use AI without becoming obsessed with it.

FL
Joined December 2017
the interesting AI story today isn't another chatbot feature. Anthropic says Claude is now leading around 26% of the R&D work involved in developing its next models, while humans still set goals and supervise the work. that's a very different productivity loop from “ask AI to write an email.” the more interesting question is what happens when AI starts contributing to the process used to build better AI. research → model → better research → better model. the feedback loop is getting much tighter.
14
one thing i keep noticing with AI products: the model gets most of the attention, but the workflow around it is where the usefulness starts showing up. tools. permissions. memory. data. testing. security. today's AI coding security story is a good reminder too. researchers used AI-assisted techniques to find vulnerabilities in OpenAI's systems during a sanctioned bug bounty. better models make people more productive. they can also make mistakes move faster.
1
867
this is such a smart use of context in advertising. the location becomes part of the message, so the campaign lands differently depending on where you see it. way more memorable than another random billboard.
whoever came up with the marketing for the new hunger games movie deserves a raise because this is genuinely genius. putting capitol propaganda in front of real places associated with power instead of just buying random billboards makes the LOCATION part of the campaign. you see the ad, then you see the building behind it, and suddenly the whole thing hits differently. they’re not just promoting a movie about propaganda and power. they’re making you experience the idea for like five seconds in the real world. that’s such a ridiculously smart use of outdoor advertising. obsessed.
2
713
effort controls make sense for computer agents. not every task needs maximum reasoning, so being able to trade depth for speed and efficiency should make them much easier to use day to day.
We’re rolling out effort controls in Computer’s model selector. Effort presets combine the orchestrator model and reasoning depth to control how deeply and efficiently Computer works through a task. Available now on web. Coming to mobile and desktop soon.
1
29
music discovery has basically become: "i'll just listen to one song" 45 minutes later i'm looking up the producer, checking the artist's older albums and wondering why i've never heard of them before. the algorithm doesn't need to give me the perfect playlist. it just needs to find one song that makes me curious enough to keep digging.
16
sometimes you need to see the progress to feel it. tracking sleep, activity and other health patterns can show you what’s working, what needs attention and where small changes could make a difference to your physical and mental wellbeing 🫶
3
1
4
271
the split between reasoning and voice is the interesting part here. google is pushing smarter agents while the live voice track is still playing catch-up. that gap is going to matter for anyone building real-time workflows.
google shipped gemini 3.8 flash this week calling it their most intelligent flash model yet, built for long-horizon agents and enterprise workflows. reasoning is on by default, you can't even fully turn thinking off, only dial it between three levels. but here's the part that's actually interesting: 3.8 flash has zero live api support. no bidirectional voice, no real-time conversation. it can take audio as a file input and respond in text, that's it. if you want an actual live voice model you're still on 3.1 flash live preview, a completely separate track that hasn't gotten the same upgrade. so google's newest, smartest flash model can't do the one thing google's been putting front and center in every gemini demo for the past year. the reasoning upgrades and the voice upgrades are shipping on different clocks, and right now voice is lagging behind
41
claude + salesforce is the kind of ai integration i'm more interested in than another chatbot demo. because the annoying part of work usually isn't generating the sentence. it's finding the customer record. checking the history. getting permission. updating the system. making sure nothing broke. if ai can handle that layer safely, that's where productivity starts feeling different.
1
33
ai is getting better fast, but the safety conversation is getting louder too. when researchers inside the labs are asking for slower development while the companies keep racing ahead, that’s probably worth paying attention to. the capability jump is exciting. the control problem is harder. reuters.com/technology/artif…
21
16b active for input and 8b for output is an interesting way to keep costs down. for agentic workloads that spend a lot of time processing context, that kind of efficiency could matter more than chasing bigger parameter counts.
DeepSeek v4.1-Flash, a new multimodal model from @deepseek_ai, is now available on Modal. v4.1-Flash uses DeepSeek's Causal Encoder-Decoder architecture, activating 16B parameters for input and 8B parameters for output to improve cost efficiency for input-heavy agentic workloads.
28
useful if you've been babysitting agent runs by tailing raw logs, having an actual session viewer instead of guessing what's happening mid-run is a real quality of life fix. the --web flag opening a local UI isthe one worh trying first honestly 👀
31
hy4 preview is worth keeping an eye on if you’re building with open models. among open models in agent arena, with a median task cost of $0.26 vs $0.80 for kimi k3 max. performance + cost is a pretty interesting combo.
Last month’s Hy4 preview launch landed @TencentHunyuan the #2 spot on the top 10 labs in Agent Arena among open-source labs! Across 14.5K+ real-world agent sessions, Hy4 preview ranks #2 among open models (#10 overall) with a +5.4% net improvement, trailing only Kimi K3 (Max) at +6.6%—a gap of just nearly 1 percentage point. This level of performance comes at roughly 68% lower cost: $0.26 per median task, compared with $0.80 for Kimi K3 (Max). Hy4 preview is now on the Agent Arena Pareto frontier. See the link below to explore its placement. By signal, Hy4 preview excels in: - Confirmed Success: 12% — explicit user feedback that the task worked - Praise vs. Complaint: 9% — implicit sentiment in user reactions - Bash Recovery: 7.5% — recovery from CLI errors It shows no issues with Tool Hallucination — calling tools that don’t exist By category among open models, it’s also #2 in Code and Work, and #3 in Chat. With this strong performance at a competitive price, Hy4 preview is a huge contribution to the open-source ecosystem. Congrats to the @TencentHunyuan team on this strong release!
86
christiano joining openai’s safety board is a strong move. you want someone in the room who’s willing to challenge the assumptions behind frontier ai safeguards, not just approve them.
Paul Christiano, founder of the Alignment Research Center, is joining the OpenAI Foundation Board and its Safety and Security Committee, which provides governance over the safety and security practices across OpenAI. As AI capabilities advance, strong safety, security, alignment, and governance matter more than ever. Paul’s work on AI alignment and his years at @NIST will strengthen the Foundation’s oversight, bringing an independent voice to challenge assumptions, assess safeguards, and reinforce accountability around critical decisions. Paul will also serve as a non-voting observer on the OpenAI Group PBC Board. openai.com/index/paul-christ…
72
researchers this week demonstrated an AI-powered worm that can spread through something as ordinary as a missed WeChat call. in plain terms: the worm uses AI to craft a convincing follow-up message after a missed call notification. it reads like a normal person following up. the target clicks, the worm spreads. the reason this matters for everyday people who use AI tools: the same AI capabilities that make writing assistants and chatbots useful are being used to make social engineering attacks more convincing. the pattern is not new. it is the phishing email, but now the message is personalised, contextually aware, and written at a quality that used to require a native speaker. three practical things to do today: one: any unexpected message referencing something you actually did recently — a missed call, a file you opened, a meeting you attended — deserves extra scrutiny before you click anything. two: turn on two-factor authentication for any accounts that do not have it. AI-powered worms still need your credentials. three: if a message feels slightly off even though it reads well, trust the feeling. the tells are getting subtler. your instinct is still useful. the tools are getting better. the attacks are getting better at the same rate. personal habits remain the first line of defence.
1
6
69
back from labor day. the playlist needed to match the return-to-work energy without being aggressive about it. went with solange's "a seat at the table" this morning. if you have not heard it: it is an album that is completely confident about being quiet. no big moments trying to announce themselves. just a consistent emotional temperature held for 50 minutes. the album came out in 2016 — which means it fits the 2026-is-the-new-2016 nostalgia wave that was going viral in the summer without trying to. the AI worm story from this morning, the google regulation piece, the CPI thursday — it is a loud week already. the right response to a loud start is not a loud playlist. "cranes in the sky" if you need one track. the whole album if you have the time. good tuesday.
30
drafts should honestly be standard in every creative tool. half the stress of experimenting is knowing one bad idea can mess up something that already works.
Introducing drafts. Create parallel versions of your project, let your team explore ideas, and tinker freely. Only apply changes when you’re happy with them.
1
87
there's a pretty important shift happening in how people are using ai. the smartest model available isn't necessarily the model you want running every task. perplexity's tokenomics approach is built around that idea: route each request to a model that can get the job done without wasting compute. here's why that matters 🧵
1
2
1
3
65
this is why “ai is too expensive” is becoming more of an infrastructure problem than a model problem. better routing can reduce wasted compute without forcing teams to give up the models they need for harder tasks.
1
14
the interesting future isn't everyone using the biggest model. it's ai systems getting better at deciding which model deserves your tokens. cheap model for simple work. strong model for difficult work. and hopefully, a much smaller bill at the end of the month.
17