A useful coding-agent bug fix starts with a failure you can reproduce and ends with a change you can verify. Give Claude Code or Codex the expected behavior, a failing example, and a boundary for the patch. Then inspect the tests and diff yourself.
Here is a complete small exercise. We ran the fixture locally with Bun 1.3.14 on October 8, 2026: the original implementation passed one case and failed three; the corrected implementation passed all four. These are results from the code below, not a benchmark of Claude Code or Codex. No model generated the measured attempts.
The bug: a search link loses its meaning
Imagine a search-link helper that adds a query to an absolute URL. Its contract is:
- set the
qparameter to the supplied string, replacing an existingq; - preserve other query parameters and the fragment;
- encode the query as a parameter value.
The original implementation in searchLink.ts is too simple:
https://example.test/search?q=coffee looks fine. But an existing ?sort=date, a #results fragment, or a query containing & changes the outcome.
Reproduce it with four checks
Create a disposable directory outside your real project. Save the helper above and this file as searchLink.test.ts:
Run:
The expected original result is 1 pass, 3 fail. Keep that output. If your result differs, check the copied files and runtime before asking an agent to repair them.
Give the agent a bounded assignment
Use the failing fixture with this prompt:
Claude Code's official best practices emphasize giving the agent a way to verify its work. The important part here is the observable contract, not a particular phrase in the prompt.
For a real repository, use its existing test runner and conventions. The existing-codebase workflow helps find those before editing.
Inspect the correction
One implementation that satisfies this fixture is:
Run the same command again. Our local result was 4 pass, 0 fail.
The change uses the existing URL structure instead of appending a second ? or placing the query inside a fragment. searchParams.set replaces the old q, while preserving the other parameters. The test checks decoded values because equivalent URL serializations can encode spaces differently.
This helper deliberately accepts absolute URLs. It does not define relative URL resolution, which would require a base URL and another requirement. It may also normalize serialization, so do not apply it blindly where a signed URL must retain its original bytes.
What these four passing tests do—and do not—establish
They show that these four cases satisfy the stated helper contract in the tested runtime. They do not prove that your UI calls this helper, that the server interprets q as intended, or that every input is supported.
In a real search feature, continue with a browser check: submit the form, inspect the resulting URL, and verify that the displayed results match it. Add checks for relative URLs or empty values only if your feature needs those behaviors.
If an agent reports success without running the command, ask for the missing evidence. If the command fails before loading the tests, resolve the environment error separately. Neither response should be converted into a claim that the model can or cannot debug.
Try the exercise in one Workspace
In Agent.Space, use a disposable project with the two fixture files, select an available Claude Code or Codex configuration, and send the bounded assignment. Review the saved helper, the test output, and the diff before using the approach on production code. See the Claude Code Workspace guide for setup.
If you compare two configurations, give each an independent copy of the original failing files. Running the second agent against the first agent's corrected files is not a comparison. For a broader evaluation, use the coding-agent regression guide.
