Compare Test Versions
Success rates from your own trials, each with a 95% confidence range, so you can see whether two versions are actually different or whether you have not run enough attempts yet.
Inputs
What you tested
Planning the test
These four questions decide whether the numbers above mean anything. They are here to help you run the test, and they are saved with your evidence so you can see later what the conditions were. Nothing you type here is written up for you.
Results
Highest measured rate
—
Its success rate
—%
Its 95% confidence range
—%
Versions compared
—
Attempts counted
—
How this is calculated
A success rate on its own is not a measurement. It is a measurement plus an unknown amount of luck. Twenty-one successes out of thirty is 70%, but run the same thirty attempts again and you would not get twenty-one a second time. The confidence range says how far the true rate could reasonably sit from the number you happened to measure.
Where the range comes from
It is a 95% Wilson score interval. The simple textbook interval breaks down at small samples and at rates near 0% or 100%, where it can hand back a range extending below zero. Wilson does not, which is why it is the one used here and in the reliability tools.
Why overlap is the thing to look at
If two versions have overlapping ranges, your data has not separated them. That is not a failed test, it is a test that needs more attempts. The gap between 70% and 87% looks decisive until you notice that thirty attempts puts about fifteen points of uncertainty either side of both numbers.
Making the attempts worth counting
- Change one thing at a time. Two changes at once and you learn nothing about either.
- Alternate between versions. All of A and then all of B lets a draining battery masquerade as a design difference.
- Decide what counts as a success before you start.Deciding afterwards, attempt by attempt, is how a test comes to tell you what you were hoping for.
- Write down what you saw, not just the count. The counts say what happened. The observation is where the reason lives.
Questions worth asking about your own numbers
These have no answers here, on purpose. They are yours to answer, and the answers are the part that is worth anything.
- Which version has the widest range, and what does that say about how much you tested it?
- If two ranges overlap, how many more attempts would it take to separate them?
- Did anything in the observation column happen often enough to be a pattern?
- What would you have to see to decide a change was not worth keeping?
Save this as evidence
Collects what you entered, what came out, how it was worked out, and anything the tool flagged, with a timestamp and a version so someone else can reproduce it.
This is evidence, not a notebook entry. It deliberately does not write your problem statement, your reasoning, or your conclusion, because under RECF rules an Engineering Notebook has to be the students' own work and no tool may generate or organise its content. Take the numbers, decide what matters, and write it yourself.
Sources & assumptions
No VEX data is used. Every number comes from attempts you counted yourself.
The interval is the Wilson score interval at 95%, the same method used by the reliability tools, chosen because it stays sensible at small samples and at rates close to 0% or 100%.