Copenhagen. I build practical tools and write about them.

Copenhagen
Joined June 2026
I'm Christian. Copenhagen. I build practical tools and write about the ones worth using. This account is mostly: • AI models, pricing, and what actually changed • Grok, Grok Bot, and the xAI / SpaceXAI stack • Agent workflows, tuning, and the parts that make them reliable • Crypto news when it moves something real If you want the launch note plus the “should I switch” note, follow along.
57
Claude Sonnet 5.5 is live. Same $2 / $10 card as Sonnet 5. Anthropic says it runs 30%+ faster and costs up to 30% less per task because it burns fewer tokens. Terminal-Bench 4.0: 70.6% vs 10.3% for Sonnet 5. Mid-tier just became the daily coding default for a lot of teams. Docs, slides, and scoped agent work too. Haiku 5.5 is next. anthropic.com/claude-sonnet-…
23
New from @claudeai: Sonnet 5.5. Anthropic says it generates 30%+ faster than Sonnet 5 and costs up to 30% less per task. Built for everyday coding, agents and polished docs; Opus 5.5 remains stronger for complex, open-ended work. Curious to try it. #Claude #AI
18
New public scoreboard for cyber defense agents. Grok 4.7 and Xiaomi MiMo-V2.6-Pro both land at 56 on Artificial Analysis’s Cyber Index. The test is find the bug, prove it, then patch it. Useful takeaway for builders: if your agent has to finish the job, not just write a polite note, measure completion, not only the refusal rate.
Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with @CollinearAI, @IBM, @nvidia, and @vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from @CollinearAI, covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from @vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from @BerkeleyRDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%.
3
99
Loving what the team is building with Hermes Agent. @Teknium @karan4d @NousResearch — any chance a dedicated mobile app is on the roadmap? Personally, I’m really hoping for Android! 🙏 Telegram is a useful way to stay connected to my agent, but I’d love a mobile experience built for Hermes itself—not limited to what fits inside a messaging platform. Imagine picking up a desktop session on your phone, viewing interactive previews, managing tasks and files, reviewing approvals, and following ongoing work in one place. Not just chatting with your agent, but having a proper workspace in your pocket. With so many AI assistants already offering dedicated mobile apps, that convenience is becoming an expectation. I’d love to see Hermes bring its own strengths to mobile, too. Huge appreciation to the founders, developers, and contributors. This is a hopeful feature request from someone who wants to use Hermes even more. Fingers crossed we’ll see it soon! 💛
2
8
757
Not bad not bad at all
grok imagine is getting AI music generation. an early version of grok music is now being tested on android, letting you describe the track you want and generate it directly inside imagine. rap, phonk, pop, country, synth and more. just prompt the sound, and grok builds the track.
6
BREAKING: SpaceXAI and X are in final-stage testing to connect Grok on X with Grok.com and bring Grok Bots to XChat. Chats will remain synced, conversation history will carry over, and Grok Bots will appear directly in XChat—including the ability to create new ones. Another step toward the Everything App.
BREAKING: SpaceXAI and X are in final testing on linking Grok on X with Grok.com and bringing Grok Bots to XChat. Chats stay in sync, history flows over, and your Grok Bots show up right in X Chat, including creating new ones. Another step toward the Everything App. Lets Goooooo!
25
SpaceXAI just launched the Grok Bot Template Rewards Pilot. Invited creators can get paid for Grok Bot templates they publish and post on X. X decides the amount. It is not a revenue share, and it is separate from Original Content Rewards. You need to be 18+ in a US state where X Money works, on X Premium, with a Grok Bot account and an X Money account that can get paid. Rewards are planned every two weeks, in dollars, to that balance.
𝕏 just launched a new Grokbot Template Rewards program Creators can now build Grokbot templates, share them on 𝕏 and may earn rewards when people actually use them • Invite-only pilot • Rewards determined every 2 weeks • Usage + repeat usage matter • Likes and impressions don’t determine payouts • Payments go directly to X Money • Completely separate from Original Content Rewards The pilot is starting in the U.S. and is expected to run for roughly 2 months This is basically 𝕏 starting to reward people for building useful AI workflows, not just content
2
1
131
Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS. You can design a voice from a prompt, steer pacing and emotion line by line, or pick from 2,000+ ready voices. Live now in the Gemini API and AI Studio. Output is SynthID-watermarked. Useful for voice agents, audiobooks, and product audio without a custom studio. blog.google/innovation-and-a…
1
40
Christian Cesar retweeted
You can now watch your agents work live as Hermes Desktop streams the screen of any bot's screen or session in real time. Watch it drive a browser, type into a terminal, or open a window. Take over anytime to interact with the Bot Screen or type credentials and seamlessly hand back off when you're done.
143
151
55
2,256
483,664
Amazing video and are stupid jealous of your black cyber truck.. displaying grok bot in tesla is impressive..... congratulations
Grok Bot just released for Tesla and I'm blown away I was lucky enough to have early access. Having your car drive you around while you talk to an army of agents is incredible In this video I take you for a ride in my Cybertruck and show you just how awesome this new release is
26
Same-day drop from two labs. Claude Opus 5.5 is live. Fable 5.1 level on most work, about 40% cheaper than Opus 5, and more than 30% faster. OpenAI followed with GPT-6 Sol and Luna. Astra-class work at half the old Sol/Luna API price. Sol is $2/$10. Luna is $0.10/$0.50. More models. Lower cost per task. Good day to build.
54
Grok 4.7 is live. Same $2 in / $6 out as 4.6. CursorBench 4.0: 46.3% vs 40.4% on 4.6. Terminal-Bench 4.0 almost doubled: 38.0% vs 20.3%. 500k context. In Cursor, Grok Build, and the API today. It stays on longer jobs and checks its own work more carefully. Fast variant is 2x output speed at 2x price.
30
Grok 4.7 just dropped. Same $2/$6 pricing as 4.6. Stronger on long coding + knowledge work. Live in Cursor, Grok Build, and the API. Time to put it to work. x.ai/news/grok-4-7
30
Christian Cesar retweeted
Grok 4.7 by @SpaceXAI is 50% off for one week in Hermes Agent via Nous Portal portal.nousresearch.com/mode…
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
45
32
23
742
90,524
Christian Cesar retweeted
SpaceXAI just released Grok 4.7 And it’s already showing a huge jump in multi-hour office work Grok 4.7 outperforms GPT-6 Astra and is already nearly matching Fable 5.1
120
232
20
1,817
367,133
Christian Cesar retweeted
Grok Build could soon control your computer from your phone Remote Control would let Grok Build keep working on your actual computer while you manage it from the web or mobile app.
16
4
67
55,024
Christian Cesar retweeted
Grok Bot can now send you voice notes This makes it feel way more natural....instead of everything coming back as text, your Bot can literally talk to you this is going to change how people actually use Grok Bot day to day....and make the entire experience so much better
15
22
1
240
11,848
Christian Cesar retweeted
Introducing Grok Voice Transcribe 2.0. It’s the world’s most accurate speech transcription model.
302
467
149
5,562
9,810,334
Christian Cesar retweeted
Grok @bot but now it lives in a widget on your phone. In two interface versions: One for keeping several Bots in sight, and one for when you only care about a single one.
67
97
16
1,966
148,002