@a_lohan

product lead, ai & platform. ex-quant. i publish llm-systems research and run a live ai trading lab (https://nitter.cf/t.co/9KBvORCKQ6).

San Francisco
Joined May 2024
i build ai products by day and run ai experiments at night. this account is the lab notebook. nightclaude: claude trades a $100k s&p 500 portfolio nightly through a real brokerage. day 75: down 5.8% while spy sits at record highs. every fill published anyway. a scorecard you can't cherry-pick. research: my latest paper certifies which cached llm answers survive a prompt edit. 225 fresh calls, 81,289 of 89,184 cells kept, error provably bounded. product: notes from shipping ai platforms. evals, agent ux, what breaks in production. receipts, not takes.
5
743
I built a late-night show that writes, voices, scores and sings itself. Type a topic, get an episode. Four @MiniMax__AI models on @gmi_cloud, made for MiniMax Week. Demo below: youtu.be/pakzV9D2tW8
1
147
I tested what a sub-agent does when it sees the hierarchy above it across 15,589 pre-registered delegations on two frontier reasoning models. Here is what the data says, and what it means for Claude Code, Codex, and Agents SDK setups. Research Paper: doi.org/10.5281/zenodo.22548…
2
1
7
77,218
Our first arbiter reported 10 departures in concurrent code runs. Every one was a sibling's edit inherited through a shared file. Diffed against what the worker last read, they vanish. Diffs against the original file report scope creep that never happened.
1
21
For Claude Code, Codex, Cursor, or Gemini CLI: write the goal and brief into the task prompt, ask for a concerns field, and keep concurrent writers off shared files. For OpenAI Agents SDK handoffs: use input_filter to trim the full transcript to the brief.
29
. @ClaudeDevs thank you for the suprising limit reset wow. Fable 5.1 is my only child no one could make another.
61
day 101: nightclaude is down 6.45% since may. still trailing the s&p by 9.25 points. every fill posts live regardless.
52
We analyzed 888 Flock Safety police surveillance portals across America. 82.5% display exactly the same data-retention period: 30 days. Not because 733 cities independently chose 30. Because that's what Flock shipped. Paper + full dataset: doi.org/10.5281/zenodo.22157…
1
1
1
16,588
[3/] The finding that surprised us most: cities that passed surveillance oversight ordinances, with mandatory hearings, impact reports, and recorded votes, still didn't deliberate the retention number. Berkeley had 156 speakers at its ALPR renewal. The word "retention" appears zero times in both the adoption and renewal agendas.
1
23
[4/] We're releasing everything: the full dataset, every analysis script, and the preprint.
13
Proposed Data Center in Atherton, CA. # I stand with this Data Center.
🤖 Made with AI
121
. @rakyll i have questions. but it is 12:24am so will pass for now.
1
68
few realize what has happened. whatever is your take away from this. great outcome for american consumers.
Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously previewed as Ox Alpha, running entirely on Chinese AI chips Blog: z.ai/blog/glm-5.3-flash Available now across all official platforms: Weights: huggingface.co/zai-org/GLM-5… API: docs.z.ai/guides/llm/glm-5.3… Coding Plan: z.ai/subscribe ZCode: zcode.z.ai/en Chat: chat.z.ai AutoClaw: autoclaw.z.ai
1
70
day 92: max conviction, 3x, same as most nights. down 9.42% total return, 11.27 points behind spy. tripling your position has never once been the same as tripling your chance of being right.
32