9. Telling a Real Improvement From Noise
We ran one prompt five times without changing a character. Output ranged from 80 to 100 words. If you compare one run to one run, you measured the dice.
The WJS Desk
Sep 7, 2026 ยท 7 min read

You change a prompt, the output looks better, you keep the change. That is how everybody does it and it is unreliable, because the output would have looked different anyway.
This lesson is a way of telling a real improvement from noise, and it takes about five minutes.
First, see the noise
We took one prompt and ran it five times without changing a single character:
Our team ships a release every two weeks. Last quarter we
shipped 6 releases, 2 slipped by a week, and we had 3
rollbacks. Write a one paragraph status update for our VP.
| Run | Length | Opened with |
|---|---|---|
| 1 | 96 words | Over the last quarter, we maintained our cadence |
| 2 | 100 words | Our team maintained our biweekly release cadence |
| 3 | 94 words | Over the past quarter, our team maintained |
| 4 | 80 words | Our team maintained our bi-weekly release cadence |
| 5 | 89 words | Our team executed on the bi-weekly release cadence |
Identical prompt, five runs, and the length varies from 80 to 100 words. That is a 20 word spread, roughly a quarter of the output, on a prompt that did not change.
This is why single-run comparison does not work. If you tweak a prompt and the new version is 15 words shorter, you have learned nothing. The baseline moves by 20 words on its own. You measured the dice.
The five run method
The fix is not complicated, it is just something almost nobody bothers with.
- Decide what you are actually measuring before you look. This is the step people skip and it is the one that matters. Not "is it better" but "does it lead with the number", "is it under 60 words", "does it mention the deadline".
- Run the current prompt 5 times. Fresh conversation each time, so no history bleeds across.
- Score each one against your criterion. A tally, not a feeling. 3 out of 5.
- Change one thing. One. Change two and you will not know which worked.
- Run 5 more and score again. 5 out of 5 beats 3 out of 5. 4 versus 3 is probably still noise.
What it looks like when you do it
We added two constraints to the prompt above: lead with the number the VP will care about most, under 60 words, no preamble. Then ran it five more times.
| Original | With constraints | |
|---|---|---|
| Length range | 80 to 100 words | 32 to 48 words |
| Spread | 20 words | 16 words |
| Led with cadence | 5 of 5 | 0 of 5 |
| Led with the rollback rate | 0 of 5 | 5 of 5 |
Look at which row carries the finding. The spread barely moved, 20 down to 16, and as a proportion of a much shorter output it actually got wider. If length variance had been our measure we would have concluded the change did nothing.
The bottom two rows are unambiguous. Every unconstrained run opened by saying the team maintained its release cadence, which is the flattering framing. Every constrained run opened with the rollback rate, three of six, fifty percent. That is a 0 to 5 flip on the thing that actually mattered, and we only saw it because we decided what to measure before running anything.
Measure the thing you care about, not the thing that is easy to count. Length is easy to count and it was the wrong measure.
Why the obvious measure was the wrong one
It is worth sitting with that result, because the mistake in it is the one you are most likely to repeat.
Length was the obvious thing to measure. It is a number, it requires no judgement, and it is right there. By that measure our change looked marginal: spread went from 20 words to 16, and relative to a much shorter output it arguably got worse.
But nobody actually cared about length variance. What we cared about was whether the update led with the bad news or buried it, and on that measure the change was total. Every single unconstrained run opened by noting the team had maintained its cadence. Every single constrained run opened with the rollback rate.
The lesson generalizes past prompting: the easy measurement and the important one are rarely the same, and the easy one will quietly become your definition of better if you let it. Deciding what counts before you look is what stops that happening.
What a tally actually buys you
Three things, and the second is the one people do not expect.
It settles arguments with yourself. A score of 4 of 5 against 3 of 5 is not an improvement, and knowing that stops you carrying a change you cannot justify.
It finds prompts that were already fine. Perhaps a third of the time we run this, the current prompt scores 5 of 5 and the thing that annoyed us was one bad sample. That is a genuinely useful outcome: you were about to spend twenty minutes fixing something that was not broken.
It gives you a floor to defend. Once a prompt scores 5 of 5 on a criterion, any future edit has to keep scoring 5 of 5. That is how a prompt stops drifting as you tinker with it over months.
Criteria that work
A good criterion is answerable yes or no by someone who is not you. Vague ones give you the same guessing you were trying to escape.
| Useless | Scoreable |
|---|---|
| Is it better | Does it lead with the risk rather than the good news |
| Is the tone right | Does it avoid opening with thanks for your time |
| Is it accurate | Does every number appear in the source I gave it |
| Is it complete | Does it name a date and an owner |
Two or three criteria is plenty. You are not building a benchmark, you are deciding whether to keep an edit.
When five runs is too much effort
Most of the time it is, and we do not do this for a one-off email. Reserve it for prompts you will run many times, because that is where a small reliability gain compounds and where you will otherwise carry a superstition for months.
For everything else, two cheaper checks:
Run it twice, not once. Even two runs tell you whether a difference is stable. If your improvement shows up in one of two, it is noise.
Test the hard case, not the easy one. Prompts do not fail on typical input, they fail on the awkward one. If you are summarizing tickets, test the ticket that is three words long and the one that contains two unrelated complaints. A prompt that survives the edges works everywhere in between.
Testing a prompt that has no right answer
Most of this is easier when there is a fact to check. Writing has no correct output, so people conclude it cannot be tested. It can, you just score properties rather than correctness.
For a piece of writing, useful scoreable properties:
- Does it open with the point rather than with context
- Is every claim in it traceable to something I supplied
- Would a stranger know what they are being asked to do
- Does it avoid the specific phrase I banned
- Is it inside the length I asked for
All five are yes or no, none of them requires you to judge quality, and together they catch most of what makes a draft unusable. The one about a stranger is the strongest: paste the output into a fresh chat and ask "what is this asking me to do", then compare the answer with what you intended.
Common mistakes
- Comparing one run to one run. The whole reason this lesson exists.
- Changing two things at once. Then you keep both forever, including the one that did nothing.
- Testing in the same conversation. The earlier attempts are in the context and pull the later ones toward them. Fresh chat every run.
- Deciding what counts as better after seeing the outputs. You will find a reason to prefer whichever one you already liked. Write the criterion down first.
- Judging on the easy case. Everything works on the easy case.
One measurement worth running once
Before you tune anything, run your single most-used prompt five times and just look at the spread. Not to improve it, only to see it.
Most people have never watched the same prompt produce five different answers side by side, and it recalibrates everything. You stop reading one output as the tool's opinion and start reading it as one sample. That alone changes how much you trust a single answer, and it is the mental shift the rest of this lesson depends on.
Try it now
Take a prompt you use repeatedly. Write down one thing it should always do, in a form someone else could score.
Run it five times and tally. If you score 5 of 5, the prompt is fine and your instinct that it needed work was wrong, which is a useful thing to learn for free. If you score 2 of 5, you have found a real problem and you now have a baseline to improve against.
Next
Lesson 10 puts the whole course together into a setup you keep using, and covers the one habit that separates people who get value from this from people who drift away.


