Practical guide · Dale Spoonemore
The four QA questions I kept asking
Source post: October 7, 2026. Adapted here: .

A shortened excerpt from my October 7 post.
Every time an agent told me a fix was ready, I’d watch the video and ask basically the same four questions: Are there any risks with this? Any unintended side effects? Do we have tests that cover this everywhere? Is there anywhere else in the app where this same problem could show up?
So I had Devin go back through our sessions, look at what I’d been doing manually, and turn that into an actual QA process. Now when a fix is finished, it goes to a separate QA agent first. That agent can read everything, but it can’t change anything.
It checks that every part of the fix has a test that would fail if the fix broke, looks for the same bug anywhere else in the app, and then watches the entire final video to make sure everything the narration claims actually happens on screen.
Only after it passes all of that does it come to me.
The handoff is the useful part
I often start and steer work in voice chat. I'll talk through what I want or what needs checking. The checklist exercise below uses independently written, fictional voice starters and optional written aids you can ask an agent to prepare. These examples haven't been tested and don't reproduce my sessions, internal prompts, or the company's setup.
I've also had a coordinating agent turn my shorter input into a detailed brief and handoff instructions, with me steering and correcting the proposal. The written instructions don't all have to start with me.
The published workflow has three roles: an implementation agent makes the change, a separate QA agent reviews it, and I decide whether it is ready. The reviewer checks the result against the requested behavior, tests, related code, and the demonstration.
To try this yourself, use a small project and give the reviewer the original task, the exact change being reviewed, test results, and a short demonstration. Without the task, it can tell you whether code looks reasonable without answering whether it solves the right problem.
Try it on a checklist app
In the checklist app, checked items should stay checked after a reload, and Reset should clear the saved state.
- Give an implementation agent the checklist persistence plan linked below. Work in a disposable branch or copy of the app.
- Once the change is ready, record the revision and collect the changed files, actual test output, and a short demonstration of check → reload → Reset → reload.
- Start a separate review with the task and that evidence. Where your tools support it, restrict source access to read-only and keep credentials and production access out of the review environment. A prompt alone does not enforce permissions.
- Ask the reviewer to check the behavior and report reproducible failures. If a check cannot be performed, record it as not run rather than treating it as passed.
- Return concrete findings to the implementation agent. Review the resulting revision again, then make the final release decision yourself.
A conversation you can try
Keep credentials and production systems outside this exercise. Enforce access limits in the tool or environment; prompt wording alone is not a security boundary.
For the optional written plan, capture stable item IDs, malformed or unavailable storage, and the two reload checks. For the review packet, capture the exact revision, diff, actual results, unrun checks, and demonstration or manual steps. Use read-only source access where supported; any execution belongs in a disposable local copy without credentials or external writes.
1. Talk through the change
Context to supply: Provide a disposable copy of your checklist app and its existing development instructions.
I've got a little checklist app, and I want the things I've checked to stay checked after I reload. Reset should clear them for good. Can you look at how it works now and tell me what you find before changing anything? Keep this to the disposable copy. If something about the behavior isn't clear, ask me.Expected output: Optional agent-produced structure to request and check: a short map of the relevant files and current behavior.
Before continuing: Check that the agent found the actual checklist and test setup. Resolve unanswered behavior questions before asking for a plan.
2. Agree on the scope
Context to supply: Provide the inspection report and your answers to its questions.
That sounds like the right part of the app. What's the smallest change that gets us there? I don't want accounts or sync or a bunch of cleanup. Think through what happens if saving doesn't work, and make sure Reset still works after another reload. Let's agree on how we'll check it before you build it.Expected output: Optional agent-produced structure to request and check: a scoped plan with observable acceptance checks.
Before continuing: Confirm the plan covers a second reload after Reset and names decisions rather than silently inventing requirements.
3. Ask for implementation and evidence
Context to supply: Provide the approved plan. Give write access only to the disposable project.
Go ahead with the plan we agreed on in this disposable copy. Stay with the existing tools. Run the checks we picked and show me check, reload, Reset, and reload again. Tell me exactly what version you tested and anything you couldn't check. Don't deploy it or touch other systems.Expected output: Optional agent-produced structure to request and check: the change plus an evidence packet tied to a specific revision.
Before continuing: Inspect the diff for unrelated changes and verify that the evidence refers to the same revision. Pass this packet to a fresh reviewer.
4. Ask for a separate review
Context to supply: Provide the original task, approved plan, exact revision/diff, results, and demonstration to a fresh review session.
Take a separate look at this change against what I asked for. Don't edit it. Does it actually keep checked items after a reload, and does Reset stay cleared? Look for side effects and whether the tests would catch the problem. Compare the demonstration with what the handoff says. Give me concrete findings and tell me what you couldn't verify.Expected output: Optional agent-produced structure to request and check: a review report with evidence, reproducible findings, and explicit gaps.
Before continuing: Return concrete findings to the implementation session. After a fix, produce a new revision and repeat this review. A person makes the final release decision; a review report does not authorize deployment.
What a useful finding looks like
Illustrative finding, not an observed bug: “Reset clears the checkmarks on screen, but reloading restores them. Reproduction: check one item, reload, choose Reset, reload again. Expected: no items checked. Actual: the first item is checked. The demonstration ends before the final reload, so it does not establish that saved state was cleared.”
That gives the implementation agent a behavior to fix and gives you a way to verify it. A vague request to “improve persistence” leaves both decisions unresolved.
Keep the first attempt small
Start with one change and one separate review. Keep the original task visible, ask for evidence, and look at what the reviewer could not check. Add steps when they address a problem you have actually seen.