Symptom
Watching the log during a batch, failures were only 8. "Almost done."
Cause
That was the front of the batch only. Sampling across offsets gave a completely different distribution.
| offset | success |
|---|---|
| 0-4,000 | 92-95% |
| 6,000 | 44% |
| 8,000 | 0% |
| 9,400 | 84% |
The real figure was 68.3%.
★★ Sort order manufactures bias. Sorted by creation date, the front is the newest data and therefore the best maintained. Estimating the whole from the front is always optimistic.
Fix
"8 failures" without a denominator says nothingThe same thing happened with sample size. The survey was run three times and all three were wrong in the same direction.
| Sample | Defect rate |
|---|---|
| 500 | 36% |
| 1,900 | 48% |
| 4,900 | 69% |
★★ If an estimate keeps getting revised in one direction only, it has not converged yet.
If it worsened every time you widened the sample and never once went the other way, the current figure is still a lower bound.
Small samples and front-loaded samples both err optimistic.