@jonperl

ceo @QAWolfHQ. views are mine only

Seattle, WA
Joined July 2009
not just tests, even our ai chats get a computer. here is why:
A new model launches every other week. GPT-6.1 Sol, Gemini 4 Argon, Claude Opus 5.5, each new release beating the last on some benchmark. But on its own, even the smartest model can only tell you how to fix a broken test. To actually fix it, it needs somewhere to work: your code, a browser to open your app, and a way to run the test and see it pass, that’s the sandbox. QA Wolf AI’s is warm in milliseconds, with a real file system, a git clone of your test suite and a live browser you can watch. When the task is big, it spins up 50 more sandboxes, each with its own browser and branch. Here’s how we built it: qawolf.com/blog/every-ai-age…
3
3
157
Jon Perl retweeted
A new model launches every other week. GPT-6.1 Sol, Gemini 4 Argon, Claude Opus 5.5, each new release beating the last on some benchmark. But on its own, even the smartest model can only tell you how to fix a broken test. To actually fix it, it needs somewhere to work: your code, a browser to open your app, and a way to run the test and see it pass, that’s the sandbox. QA Wolf AI’s is warm in milliseconds, with a real file system, a git clone of your test suite and a live browser you can watch. When the task is big, it spins up 50 more sandboxes, each with its own browser and branch. Here’s how we built it: qawolf.com/blog/every-ai-age…
1
3
1
7
3,725
Jon Perl retweeted
Welcome to self-driving QA 🚗 Your agent builds a feature ⬇️ Calls QA Wolf to test it ⬇️ QA Wolf bulk automates Playwright & Appium tests ⬇️ Tests run in parallel ⬇️ Verified bugs are shared with your agent ⬇️ Agent fixes ⬇️ Re-test & ship Start here: docs.qawolf.com/
1
4
2
6
286
web and real phones streamed to the browser. you don't need to watch the tests live. but if you want to you can. it's little things like this that make a product experience magical
Now! Watch test runs live and in living color. When test runs get stuck they hold up your release. Before today, if you wanted to know why a test was hanging, you had to wait for the recording and logs. Now you (or our agent) can jump in and intervene. qawolf.com/changelog/live-st…
3
4
161
it's insane how fast we are shipping. we completely redesigned and shipped a new UX in <2 weeks but there is no way we could have without our test coverage
1
47
so glad we shipped MCP before the US government
1
5
197
@shl is IRS MCP on the roadmap?
44
RT @_theonly1me: It’s here! the QA Wolf MCP server is live 🎉 You focus on building features. Claude and Codex can request test coverage fr…
This quoted post is unavailable.
1
4
Everything you can do in the UI works via MCP and always will be (via a shared contract with the API)
Our MCP server is live for Codex, Claude Code, or the coding agent of your choice. Now when your coding agent says the PR is done and bug free, you can get a second set of eyes to run through your end-to-end test suite. docs.qawolf.com/quick-start
3
3
172
Jon Perl retweeted
Still figuring out what all we can do with Jev. One of the first places we’re putting it to work: giving Playwright tests a little judgment. ai.expect checks meaning. ai.act can take a different route through the UI. The rest of the test stays deterministic. Wrote it up qawolf.com/blog/semantic-ass…
4
1
9
363
Jon Perl retweeted
One good use case for Jev from @typesafeai for e2e testing is that agent testing is now cheap at scale. Playwright does most of the heavy lifting, and Jev verifies that the agent did the right thing.
2
4
7
167
do not mess with our test review agent
1
3
96
Jon Perl retweeted
Your test suite is out of date the moment you ship a new feature. 📉 Before AI coding tools, a focused team *might* have kept up. Now it’s impossible. Our product-mapping AI explores your app, spots coverage gaps, and outlines new tests as you ship. qawolf.com/changelog/mapping…
30
26
12
59
26,819
Jon Perl retweeted
1/ Ship fast and your customers find your bugs. Ship safe and your competitors outrun you. Agents destroyed the balance between verification and velocity. We had to rethink how we test our own releases — here's exactly what that looks like. 🧵
22
27
11
61
18,961
Jon Perl retweeted
Everyone's talking about “lights out” code factories: AI agents that plan, code, review, and ship autonomously. Here's the part no one shares: it doesn't work. You just end up shipping untested or undertested code. And bugs. Lots of them. We wrote about it. 🧵
3
2
2
9
1,652
I really need today’s models but 10x cheaper
1
60
was worrying to see the frontier labs keep pushing more expensive models when we really needed better cost performance. very grateful the work @OpenAI did for 5.6. massive progress in this direction
14
AI;DR codebases 📈
The most dangerous trend in software isn't a new exploit. It's a behavior. With autonomous agents, developers are shipping massive PRs that are too big to actually read. That's led to more unverified or underverified code in prod. We wrote about it.🧵
3
1
6
286
crazy how smart the sand is these days
4
136
Never thought I would be coding from my phone a year ago
3
130