Check the agent's work
An agent stops when the work looks done to it. Ask it to test its changes and show you proof, then try the result yourself before you move on.
Claude Code’s guide describes the problem: “Claude stops when the work looks done. Without a check it can run, ‘looks done’ is the only signal available, and you become the verification loop.” The fix is to give the agent a way to check itself, and to check the result yourself.
Ask the agent to test and show proof
- Say how you’ll judge it. End requests with what should be true when it’s done. See Describe what you want.
- Ask it to try the result in the browser. Agents can open your site in Enjoy’s browser beside the conversation, then read the page, click, fill in forms, scroll, and move between pages. On a Mac, they can also take screenshots of it. They can’t upload files there. If you’ve installed the Claude in Chrome extension and have a paid Claude plan, Claude Code can also use your Chrome browser, which helps with sites where you’re already signed in. See Let Claude Code use your Chrome browser.
- Ask for evidence. Claude Code’s guide recommends having the agent “show evidence rather than asserting success: the test output, the command it ran and what it returned, or a screenshot of the result.” OpenAI’s Codex guide gives the same advice: “run the relevant checks, confirm the result, and review the work before you accept it.”
Check it yourself
- Try it like a visitor. In the computer app, the site appears beside the conversation. If it’s hidden, select the globe button (Show product browser). Select Expand product for more room, Use mobile width to see it at phone size, and Refresh product after changes. The web app shows the site beside the conversation too, or as a card above the message box in a narrow window. In the iPhone app, tap the row under the conversation’s title. See See your site from the web app or your iPhone.
- See what the agent did. Select the status under the conversation’s title to open the agent activity panel. It lists the commands the agent ran, the files it read and edited, and what happened.
- See what changed. When the agent commits its work, its reply shows View changes, with each changed file. See Review the changes an agent made.
- Get a second opinion for important changes. Start a new conversation and ask another agent to review the work. Claude Code’s guide notes that a fresh conversation reviews better “since Claude won’t be biased toward code it just wrote.”
An example
You asked for a booking form. Before saying thanks, you send:
“Before we call this done: open the booking page in the preview, book a test class with a made-up name, try it at phone width, and check that the confirmation shows. Tell me exactly what you tried and, if you can, send a screenshot of the confirmation.”
The agent reports that the form worked but the confirmation was hidden behind the footer on phones. It fixes that and checks again.
Common pitfalls
- Taking “Done” at face value. Claude Code’s guide calls this “the trust-then-verify gap”: “a plausible-looking implementation that doesn’t handle edge cases.”
- Testing only the happy path. Try a wrong email, an empty field, a second booking.
- Checking only at desktop width. Many visitors use phones.
- Asking the agent to grade its own work in the same conversation. For anything that matters, use a fresh conversation for the review.