Your inspection pass rate is not your defect rate
A pre-dispatch inspection tells you whether a lot conformed on the day it was checked. It cannot tell you what happens after six weeks in a customer's hands. Here is how to build the second number from signals you already collect, and how to clean them enough to use.
- An inspection pass rate and a field defect rate measure different things, so a high pass rate alongside steady complaints is not a contradiction.
- Choose the denominator first: units shipped from a single batch, compared at a fixed age, in one channel at a time.
- Replacement and warranty requests are the least polluted signal, while return reason codes and review text are best used to name the fault mode rather than to size it.
- Stop a batch on two conditions together, a rate materially above the last comparable batch and one fault mode accounting for most of the events.
Two different questions, two different answers
A pre-dispatch inspection asks whether a lot conforms to specification on the day it is checked. Field data asks whether the product survives contact with a real customer for six weeks. Those are different questions, so a high pass rate sitting alongside a stream of complaints is not a contradiction. Inspection works, and this post assumes you already run one. The problem is everything that happens after the container leaves.
The faults that reach customers are usually the ones a static check cannot see. A seal that holds at the factory and fails after three temperature swings in transit. A cell that passes a bench test and dies in week five. An adhesive that lets go in coastal humidity but not in an air conditioned QC room. You will only ever observe these in the field, which means you need a field number sitting next to the pass rate.
Pick the denominator before you pick the numerator
Most teams start by counting complaints. Start at the other end. What you divide by decides whether the number means anything at all.
- Units shipped from a batch, not units sold in a month. A monthly rate blends three production runs and tells you nothing reliable about any of them.
- A fixed maturity window. A batch that shipped last week has not had time to fail yet. Compare batches at the same age, say thirty days from first delivery, or the comparison quietly favours whatever is newest.
- One channel at a time. Marketplace, quick commerce and your own site have different return windows and different complaint friction, so they produce different rates on identical stock.
Suppose, purely as an illustration, you shipped four thousand units of one batch and after cleaning the data you are left with forty-eight confirmed defect events at day thirty. The arithmetic is trivial. The result is also meaningless on its own. It becomes useful only when you have the same figure, computed the same way, for the batch before it.
Four signals, ranked by how much they lie
Replacement and warranty requests are the cleanest. The customer is asking for the product rather than for their money back, so remorse is largely filtered out already. Support contact reasons come next, but only if agents tag by cause and not by outcome. A queue full of tickets tagged refund tells you what you did, not what broke. Return reason codes are coarse because they are dropdowns, and the quality option is a bin for everything the customer could not classify. Review text carries the richest description of a fault and the worst denominator, because reviewers self-select and one articulate complaint can outrun a hundred quiet ones.
Use all four, but never add them together. Build the rate from the cleanest signal you have and use the rest to name the fault mode.
Netting out the pollution
Every one of those signals is contaminated, and the contamination is not random.
- Wrong size, shade or variant. This is a catalogue and size chart failure. Track it, act on it, keep it out of the defect numerator.
- The customer did not read the listing. Also a real problem, also not the factory’s. If your copy implies a capacity the product does not have, that is a listing defect.
- Return fraud. A separate discipline with its own evidence process. For this purpose you only need to exclude it, not to fight it.
- Transit damage. A packaging or handling failure. It clusters by lane, courier and pack format rather than by production batch, which is exactly how you tell the two apart.
The working test is convergence. When the same fault description shows up across several pincodes, more than one courier and one batch code, you are looking at the product. When it clusters on one lane and one courier across many batches, you are looking at the box.
Cohort by batch or you learn nothing
None of this survives without the batch code reaching the person handling the complaint. That means capturing batch or serial at the point of contact rather than hunting for it afterwards, and it means the code has to still be legible after a month in someone’s kitchen or bathroom. Without that link you have a trend line, which is enough to worry about and not enough to act on.
The rule that decides whether you stop a batch
Write the trigger down before the batch ships, because you will not write a fair one while the phone is ringing. Two conditions matter together. The first is a rate materially above the last comparable batch at the same maturity. The second is concentration, meaning one fault mode accounts for most of the events. A raised rate with scattered causes is usually noise or a seasonal handling issue. A raised rate with one dominant cause is a production problem.
Define the mechanics in advance too. Who can call a hold, what the hold covers, whether you quarantine the unsold remainder, whether you pull the retained samples for that batch, and at what point the platform is told rather than left to find out.
What you actually put in front of the supplier
A number without a method invites an argument about the number. Send both. The pack should carry the batch code and manufacture date, units shipped, the observation window you used, the event count with your reclassification rules stated plainly, the dominant fault mode with photographs and a failed unit available for return, and the identical calculation for the previous two batches.
That last item is what changes the conversation. A supplier can dispute a rate. It is much harder to dispute a rate when the same method, applied by the same team, produced a lower one three months ago on their own output.