@jonperli
iAccount based inUnited States
About this account
- Account based in
- United States
- Connected via
- United States App Store
Account-level information from X, not a live location or the device used for a specific post.
ceo @QAWolfHQ. views are mine only
Seattle, WA
Joined July 2009
- Tweets1.3K
- Following25
- Followers680
- Likes2.8K
not just tests, even our ai chats get a computer. here is why:
A new model launches every other week. GPT-6.1 Sol, Gemini 4 Argon, Claude Opus 5.5, each new release beating the last on some benchmark. But on its own, even the smartest model can only tell you how to fix a broken test. To actually fix it, it needs somewhere to work: your code, a browser to open your app, and a way to run the test and see it pass, that’s the sandbox.
QA Wolf AI’s is warm in milliseconds, with a real file system, a git clone of your test suite and a live browser you can watch. When the task is big, it spins up 50 more sandboxes, each with its own browser and branch.
Here’s how we built it: qawolf.com/blog/every-ai-age…
Jon Perl retweeted
A new model launches every other week. GPT-6.1 Sol, Gemini 4 Argon, Claude Opus 5.5, each new release beating the last on some benchmark. But on its own, even the smartest model can only tell you how to fix a broken test. To actually fix it, it needs somewhere to work: your code, a browser to open your app, and a way to run the test and see it pass, that’s the sandbox.
QA Wolf AI’s is warm in milliseconds, with a real file system, a git clone of your test suite and a live browser you can watch. When the task is big, it spins up 50 more sandboxes, each with its own browser and branch.
Here’s how we built it: qawolf.com/blog/every-ai-age…
Welcome to self-driving QA 🚗
Your agent builds a feature
⬇️
Calls QA Wolf to test it
⬇️
QA Wolf bulk automates Playwright & Appium tests
⬇️
Tests run in parallel
⬇️
Verified bugs are shared with your agent
⬇️
Agent fixes
⬇️
Re-test & ship
Start here: docs.qawolf.com/
web and real phones streamed to the browser. you don't need to watch the tests live. but if you want to you can. it's little things like this that make a product experience magical
Now! Watch test runs live and in living color.
When test runs get stuck they hold up your release. Before today, if you wanted to know why a test was hanging, you had to wait for the recording and logs.
Now you (or our agent) can jump in and intervene.
qawolf.com/changelog/live-st…
RT @_theonly1me: It’s here! the QA Wolf MCP server is live 🎉
You focus on building features. Claude and Codex can request test coverage fr…
This quoted post is unavailable.
Everything you can do in the UI works via MCP and always will be (via a shared contract with the API)
Our MCP server is live for Codex, Claude Code, or the coding agent of your choice.
Now when your coding agent says the PR is done and bug free, you can get a second set of eyes to run through your end-to-end test suite.
docs.qawolf.com/quick-start
Jon Perl retweeted
Still figuring out what all we can do with Jev. One of the first places we’re putting it to work: giving Playwright tests a little judgment.
ai.expect checks meaning. ai.act can take a different route through the UI. The rest of the test stays deterministic.
Wrote it up qawolf.com/blog/semantic-ass…
Jon Perl retweeted
One good use case for Jev from @typesafeai for e2e testing is that agent testing is now cheap at scale.
Playwright does most of the heavy lifting, and Jev verifies that the agent did the right thing.
Your test suite is out of date the moment you ship a new feature. 📉
Before AI coding tools, a focused team *might* have kept up. Now it’s impossible. Our product-mapping AI explores your app, spots coverage gaps, and outlines new tests as you ship.
qawolf.com/changelog/mapping…
1/ Ship fast and your customers find your bugs. Ship safe and your competitors outrun you. Agents destroyed the balance between verification and velocity. We had to rethink how we test our own releases — here's exactly what that looks like. 🧵
Everyone's talking about “lights out” code factories: AI agents that plan, code, review, and ship autonomously.
Here's the part no one shares: it doesn't work. You just end up shipping untested or undertested code.
And bugs. Lots of them.
We wrote about it. 🧵