@sam_commonlyi
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
Building https://nitter.cf/t.co/K2zJZjAmNC and https://nitter.cf/t.co/ZbN46Kbuvr SWE working on AI Infra (I will follow back bluemarks!) CS @UCLA
San Francisco, CA
Joined February 2026
- Tweets382
- Following152
- Followers129
- Likes215
Pinned Tweet
还记得之前发过的 Commonly 吗? 肝了几个月,2.0 版本 + 公开测试来了 🎉
这版彻底变了个样:给你的 AI 拉个群。
把你本地正在用的 Claude Code、Codex、OpenClaw 接进来,组成自己的 agent 小队 —— 你在群里说话,它们干活:读文件、写方案、跑代码,交付成果,整个群共享同一份项目记忆,换个 AI 也不用从头再解释一遍。
Commonly 是开源项目,在GitHub 已经有 1000+ star ⭐
测试期间全免费,自部署永久免费。不抽成、不收代理费,用你现有的订阅直接接入。
小团队独立开发,求试用求反馈 🙏(无需邀请码,邮箱直接注册,支持GitHub/Google账号登录)
链接在评论区👇
做题家来了
Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon!
It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback.
Here’s a look at the benchmarks:
This is going to be my new avatar now
그록봇 스타일 케릭터 만들어 주는 프롬프트 공유
grokbot-icon-studio.serio-ai…
파딱이 아니라 긴 텍스트 업로드가 안되어 아예 웹앱 형태로 배포합니다. 다음 사이트에서 복사 버튼을 누르고 사용하는 이미지 생성 Ai에 붙여넣기해서 활용해 주세요
Before committing to a coding agent, check where its context lives. Claude Code keeps auto-memory under ~/.claude/projects, skills and settings under .claude, outside the repo. None of it moves with you. What moves is what you wrote in the repo: AGENTS.md, decisions, what failed.
Searched GitHub for merged PRs that say "Generated with Claude Code". In a sample of 5,000 from the last 25 days, across 2,061 repos, only 4 repos had more than one person doing it. Yet PostHog alone: 358 from 36 people. Is team use rare, or does it just live in private repos?
This is some real useful stuff, better than a random LLM judge
We benchmarked fx auto mode (safety) classifier with @typesafeai's Jev.
tl;dr: ~5-18x faster and more accurate than 𝚐𝚙𝚝-𝟻.𝟼-𝚕𝚞𝚗𝚊, our current top choice
You can now get frontier Claude model access at cheaper rates via DeepSeek and Kimi🤣
We're publishing our most detailed threat intelligence report to date.
It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them.
We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies.
These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve.
We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop.
Read the report: anthropic.com/threat-intelli…
Another refactor
Workers now enables Node.js compatibility by default, supports applications up to 64 mebibytes, and adds a URL-based module registry with import.meta, lazy compilation, shared code caches, and clearer errors. cfl.re/4ymgiqz
Halted our agents last night. Killed six processes, watched them exit.
Ten minutes later four were back. A supervisor I set up last week restarts any seat that is down, exactly as I told it to. I had forgotten telling it.
How do you stop a system built to survive being stopped?
Just received the global reset on ChatGPT, thanks!
Never gonna give you up
Never gonna let you down
Never gonna run around and desert you
Never gonna make you cry
Never gonna say goodbye
Never gonna tell a lie and hurt you
Thanks for reading. We will do a global reset of the usage for all paid subscriptions so that you can keep enjoying Astra after burning through all of it doing fun 3D modeling in blender. The work week is about to start.
Lands around 6pm PST today.
I disagree, internal tooling team knows better on how to scale and utilize existing server resources. If everyone can just go and vibe an internal tool, who will decide which one to use in the team, and who’s gonna maintain them? Everyone can say: I just vibed a better tool, please use mine.
Not all companies have the crazy computing resources like OpenAI
Super interesting detail from inside OpenAI:
The company hired a bunch of ex-Meta people. Some of these people wanted to invest heavily into internal tooling teams (Meta did it, worked great)
OpenAI leadership resisted, saying that in an AGI-first world, there will be no internal tooling teams.
And they were right: now, with Codex, there's a massive internal tools explosion, without having any internal tooling teams!
Our reviewer agent put a BLOCK on a pull request, with a real finding in it.
The PR merged an hour later. Nothing in the branch settings required that check, so the block was advice wearing the costume of a gate.
If you have agents reviewing your code, the useful test is not whether the review is good. It is whether a merge is actually impossible while the review is unhappy.
205 pull requests merged into our repo last week. Two of us are human. 33 named agents did the rest.
The thing that broke first was not code quality. It was answering "who decided this, and what did they rule out" three days later, when the only record is a diff.
Still the open question here: where does that belong, if not the commit?