@jayzhoupro

Entrepreneur | Indie Hacker | AI-First SaaS Creator | Music Producer. Documenting which AI workflows survive real work — leveling up AI% daily 🌟

Los Angeles
Joined April 2023
Just saw OpenCode Go already has DeepSeek V4.1 Flash, and they 4x'd usage for a limited time. Not the old V4 Flash. This one can actually see images. That means I can point the same model at coding, agents, screenshots, frontend debugging, browser work, long context, and research. The usual pile. $10/month. During the promo that's ~26,000 estimated requests / 5 hours, and $60 of V4.1 Flash usage. Same plan also has GLM-5.3 Flash, GPT-5.6 Luna, Kimi, Qwen, MiniMax. You don't have to live in OpenCode. They say it works with OpenCode or other agents. That's the part I care about. I already have tools. I just want cheap usage I can plug into them. I'm going to make V4.1 Flash the default. First pass on almost everything. Only escalate to a frontier model if it gets stuck. I don't know why I'd keep burning $100–200+/month of frontier usage on routine work if a $10 plan can eat most of it. For my volume I don't even expect to come close to the cap. If you use my referral you get an extra $5 OpenCode Go credit and I get $5 too.
DeepSeek Flash v4.1 is now available in OpenCode Go
1
2
887
applying ai to a boring business works because the process is most of the product there. what you're actually replacing is payroll: the intake, the follow ups, the monthly chase. one person holds all of it now.
Building something outside of AI but applying AI to it is now much more competitive than just building another AI product I think
7
a better model at a flat price is capacity you didn't have to buy. nothing to re-budget, nothing to re-plan. what changes is how much one person can attempt in an afternoon.
Grok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed.
8
the hard part of running agents at this volume is writing down what they may merge on their own. that rule is the leverage. everything after it is just throughput.
here's how i shipped 2,500 PRs last month to production this was originally supposed to be for Cursor Compile in London. i couldn't make it since i was livestreaming for Grok @Bot Galaxy so i'm making it available for free here on X! watch it on 2x speed, i talk slowly
6
every row here is a routing choice, and routing gets cheap fast. the row i don't outsource is the call that something is actually done.
This is literally my new workflow now: Realtime Research → Grok Bot Planning & Orchestration→ Grok Bot Day-to-day Coding/Debug → Grok Build + Grok 4.6 Write & Run Tests → Grok Build + Grok 4.6 Complex Coding/Debug → GPT-6 Astra Frontend → Fable 5.1 Bookmark this.
3
volume from the agent, judgment from the owner. name the human who owns each decision before the agent starts, and make them the one who has to defend it. skip that and the leverage was imaginary, it just shipped faster.
"Nobody knows anything here. The specs, code, tests, PRDs, tickets, resolution of those tickets, reports, etc., everything is made by Claude Code. Nobody on my team likes this. They are being forced to ship as much as they can." This is sad. I tell our engineers to use AI but never cede our understanding to AI. As an industry, we are driving too fast in a fancy new car we barely understand how to drive, inviting disaster.
10
everyone can get in now. the part i'm watching is how often it picks the cheap path by itself. when the expensive model stays asleep for an afternoon, a solo week gets a lot cheaper.
Jev is now available to everyone. No waitlist. Start using it here: console.typesafe.ai
2
17
the ADFGVX message was 170 characters. it sat unsolved for a century because the key was documented from december and the message went out in november. Astra doubted the manual, the cruiser's log settled it. i'd bet a lot of stuck work is stuck the same way.
A 108-year-old German WW1 radio code has been cracked using GPT-6 Astra The message, sent in 1918, revealed a British cruiser had arrived in Crimea, with an Allied squadron set to follow two days later
2
35
a bad plan costs more downstream than a bad line of code, which is why the expensive model belongs on planning. i build solo with these tools and write about where the pricey calls pay off. what do you still keep on your best model?
Talked with a company where, a year ago they had unlimited AI budgets + CEO is technical and very bullish in AI “We now have a daily budget limit for SOTA models. Use Fable and Astra only for planning, cheaper models are good enough for everything else.” Want to stress this is an “AI-pilled” company, and not one that ever looked at AI budgets, before.
1
41
the zcode story has a follow-up, and it's a legal one. a company in china has now sent z.ai (zhipu) a formal letter over the client uploading people's repos. it wants to know which entity is legally responsible for the upload, and whether the code left the country. if it did, they want the cross-border filing behind it. z.ai apologized on the 18th, called it a default-on feature and promised to open-source the client. that answers the product complaint. this is a different question. outbound data transfer in china is a filing you make before the data moves, so no statement afterwards settles it. deleting the data doesn't either.
还真有较真的公司,就智谱 ZCode 擅自上传开发数据事件,公开发函了。 它们指责,ZCode 客户端网络请求指向新加坡主体,而服务协议签约主体,却是中国的北京主体,要求智谱说明本次上传的责任主体,以及是否发生国内信息,传输境外或存储。 如发生,要求提供出境备案材料,好家伙。
1
185
simultaneous interpretation used to be a person you booked days ahead. now it runs under the conversation, so language stops deciding who you can work with. judgment stays the scarce input. who would you work with first if language stopped filtering?
Meet Qwen3.8-LiveTranslate, Qwen's next-generation real-time simultaneous interpretation model! 📢 Built on an Interleave architecture, it improves faithfulness, fluency, and conciseness while reducing average lagging (LAAL) from 2.8s to 2.3s across 60 languages. New capabilities: 🙌 - Real-time speaker diarization — distinguishes speakers in multi-party speech and preserves each speaker's voice through more stable voice cloning. - Synchronized bilingual display — source and translation on screen together. - Long-context disambiguation — leverages conversation history to clarify names and terminology for consistent translations. Let's try Qwen3.8-LiveTranslate! 🥳 - Blog: qwen.ai/blog?id=qwen3.8-live… - QwenCloud: qwencloud.com/models/qwen3.8…
21
epoch's label is the story: humans + ai, core ideas from astra. the humans' contribution was the asking, and that half transfers to your own problem. it's also where the leverage sits. what do you hand a model now that you did by hand a year ago?
Another problem from FrontierMath: Open Problems has been solved! The solution was elicited by Becker, Greger, and @DominikPeters in an interactive session with GPT-6 Astra. Peters originally suggested the problem for the benchmark. He had this to say.
17
the fast decisions in this list are the easy half to ship. surviving is explaining one three weeks later, when finance asks why the agent approved that refund. store the options and the confidence behind the pick. are you logging those, or just the outcome?
10 Jev native products I’d build, ranked by how much fast, cheap decisions change the product: 1. Agent spend firewall Before every purchase, Jev returns approve/review/deny plus a confidence score based on price, vendor, user rules and purchase history. 2. Self-healing tool calls After an API error, Jev chooses retry, wait, change parameters, switch providers or escalate. So now you have agents that recover from failure in milliseconds instead of restarting the entire workflow. 3. Irreversible action detector Jev scores every step by reversibility before the agent sends an email, deletes a file, moves money or changes permissions. You get aggressive automation for safe actions and human approval for consequential ones. 4. Dynamic permission engine Instead of giving an agent permanent access, Jev decides which tool, data and spending limit it receives for each task. So now all of a sudden, you get temporary, task-level permissions for enterprise agents. 5. Agent branch pruning An agent generates 20 possible next steps. Jev scores them in parallel and kills weak branches before expensive reasoning begins. Deeper agent planning at a fraction of the cost! 6. Production incident controller Jev reads logs, deploy history, affected customers and service health, then chooses ignore, rollback, restart, page or investigate. Someone like pager duty should build this because it's automated incident response that reacts before an engineer opens Slack. 7. Live negotiation policy During a sales, procurement or collections conversation, Jev decides whether to discount, counter, hold firm, offer terms or escalate. Reminds me of Clulely. 8. Autonomous refund desk Jev evaluates order history, customer value, fraud signals, item cost and policy, then returns approve, reject or review. Finally, instant refunds for good customers and focused review for risky cases. Note: I'll be adding more Jev related ideas to Ideabrowser.com and our agency latecheckout.agency builds the biggest agentic products 9. Realtime marketplace dispatch For every request, Jev chooses the provider using location, price, quality, availability, cancellation risk and customer preferences. Big problem is that marketplaces need to rematch supply continuously as conditions change. 10. Confidence based human queues Jev scores every agent decision and sends only uncertain, expensive or irreversible cases to a person. One person can now supervise thousands of autonomous workflows. LLMs generate possibilities. Jev chooses what happens next. The next generation of software will need both.
5
agent work has a meter now: codex shows usage per task and subagent. that's the difference between using an agent and managing one. keep the runs that pay for themselves, cut the rest. you can't maxx what you can't price. which of yours earns its keep?
Get more visibility into your usage. See how tasks, subagents, and individual chats contribute to your Codex usage, so you can make more informed choices about your workflow.
25
support is the easy half. a standard that ships switched off is just a preference: your repo carries agents.md, the agent next to you reads none of it, and you only find out when two devs disagree about the rules. is it on in a fresh install?
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
17
routing is a decision, and @ephraimduncan's router makes that pick cheap. the part i'd watch is the wrong picks: a hard task goes to the small model and the user gets a worse answer, with nothing saying a decision happened. do you log the probability next to the model it chose?
Built a model router with Jev by @typesafeai. Jev decides what model fits your request best and the request is sent to that model.
11
AgentCloak swaps your details for realistic fakes before the prompt leaves, then swaps them back. Their pitch admits hiding made answers useless. The bet is a fake that means the same thing. Fine for a name, not for the number deciding the answer. I write about AI leverage.
Introducing AgentCloak: Use any AI without sharing your real data. Chinese AI services, ChatGPT, Claude, doesn't matter. You probably try to hide details before asking: different names, fake numbers, no address. But then the answer's useless because the AI is missing actual context. AgentCloak runs in your browser. It swaps your sensitive info for realistic fakes before sending anything, then swaps your real info back into the response. You get what you need. The AI gets nothing about you. Works entirely in-browser. Already trusted by some of the biggest companies in the world. Now free to use. agentcloak.ai/
1
34
reading speed used to cap how fast a paper's method got reused. an agent runs it at machine speed, which is the win, and also why a wrong method travels further before anyone reads it closely. i write about AI Maxxing: leverage, and who still checks the work.
🔥 Excited to share that #paper2agent is published in @Nature today! Papers have long been the primary format for communicating scientific knowledge, but they remain static. Putting that knowledge to work requires connecting findings to data, navigating supplementary materials, and adapting methods to new questions - effort repeated by each new reader. We introduce Paper2Agent, a multi-agent framework that automatically turns research papers into virtual authors. Paper2Agent turns a paper’s manuscript, code, data, and supplements into an MCP server that any AI agent can access, with automated testing and iterative refinement in an agentic loop. The paper then becomes a virtual author you can talk to: trace claims to evidence, analyze your own data, and collaborate with other papers’ virtual authors Agentifying a paper turns it from something people read into something people and AI agents can discover and build on - providing the context needed to interpret its findings and reuse its methods reliably. We tested how reliably these agents put papers to use. Across multiple benchmarks, agents created by Paper2Agent outperformed baselines such as Claude Code working directly with papers' PDFs and code repositories. Once papers become virtual authors, they can collaborate - much like human researchers do. In one case study, agents built from AlphaGenome and two large-scale genetic perturbation studies worked together to propose a new computational approach for integrating evidence across different perturbation datasets. By connecting predictions from one paper with experimental data from others, they helped pinpoint a likely causal gene for psoriasis. We hope Paper2Agent makes scientific knowledge easier to access, reuse, and build on - a first step toward a future where millions of paper agents proactively collaborate with human researchers and one another to advance discovery. 🤖Talk to the virtual author for Paper2Agent itself: paper2agent.ai 📎Paper: nature.com/articles/s41586-0… 💻Code: github.com/jmiao24/Paper2Age… VERY grateful to work with this incredible team: @james_y_zou, @jkpritch, Yaohui, and Joe!
50
one dev went looking for 700MB of disk usage in ~/.zcode and found his own repo sitting in an upload queue. his teardown, run off his own machine: the client asks zcode.z.ai for upload credentials, the server hands back an OSS form signature plus the RSA public key for that round, then the client packs the workspace, encrypts it with AES-256-CTR, wraps the key with RSA-OAEP and posts the archive straight to Aliyun OSS. the private key never touches the machine. so the 313MB sitting on his own disk is a copy he cannot open, and deleting it just gets it repacked half an hour later (the retry counter was at 564). what's in it is the part that matters: of 42,411 files in the local manifest, .git is 86.6% of the payload — 196MB of LFS cache, 102MB of git objects, plus reflogs. that's not the file you had open, that's the history you deleted and assumed was gone. both toggles that read like they control it don't. repo snapshot indexing only stops the server from indexing what was already uploaded; the capture sidecar starts unconditionally and only needs a valid JWT. it fires before every prompt and again on task completion, 62 times in one session in his logs. the privacy policy talks about text and code submitted in a conversation. it never mentions workspace snapshots or git history. @Zai_org hasn't said anything publicly. there is now an issue on their own feedback repo reporting the same manifest locally with the indexing toggle off. if you're running it, either lock the directory (chflags uchg ~/.zcode/v2/checkpoints) or stop until they answer. his read, which i don't have a better one for: a key only the server can use isn't a backup. and this is one more argument for agents you can read.
Hey @Zai_org , why does ZCode silently pack entire workspaces + full .git history and upload to Aliyun OSS on login? - Server holds the only decryption key - No UI toggle to disable - Zero disclosure in privacy policy Full forensics & fix: blog.ferstar.org/en/posts/zc…
1
155
the headline is 20-200x faster. the claim underneath: a decision now costs less than the thing it decides about. @CompleteSkeptic co-invented ChatGPT, then spent 2 years on the opposite bet: a model that answers in a choice and a probability, priced like a rounding error, made for loops that never sleep. benchmarks can argue for months; unit economics argue immediately. what gets decided everywhere once deciding is nearly free?
After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2 years in stealth building a new way to train models (RLCD), and a new type of frontier AI model that we are releasing today: Jev • 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions AFAICT the shortest path to AI-based economic revolution
1
30
a model that can't write a sentence ran one typed decision per block inside a 300ms trading loop and filled real orders. the strategy is rough (@jarrodwatts says it mostly loses to spread) and it doesn't matter. this is the first clean look at intelligence as a function call instead of an interface: no paragraph to hide a bad call behind. what's the first decision you'd hand to a model that only returns a choice and a confidence?
I built a trading bot with Jev! Jev decides if it should "buy" or "sell", given the price feed of an asset pair, and executes real trades. It uses Monad to place the orders on Kuru's on-chain order book in every 300ms block. Demo link → jev-trader.vercel.app/
1
125