The Research Corpus Your Customers Already Wrote
- Each source answers a different question, which is why reading one of them is not the same as doing the work.
- The instinct is to read everything or to read only the one star reviews.
- Reading produces impressions. Coding produces findings.
Most consumer brands in India say they cannot afford customer research. Meanwhile their customers write to them every single day, for free, in volume, unprompted, and nobody reads it systematically.
Product reviews. Marketplace question and answer threads. Support tickets and chat transcripts. Return reason fields. Courier remarks on failed deliveries. Comments under your own ads. Replies in your WhatsApp broadcast. This is a standing research corpus. The only cost is the time to read it properly.
The corpus you already commissioned
Each source answers a different question, which is why reading one of them is not the same as doing the work.
Reviews tell you what mattered enough for someone to write about after the fact, filtered heavily towards delight and anger. Marketplace question threads tell you what a buyer needed to know before purchase and could not find on your listing, which makes them a direct audit of your content. Support tickets tell you where the operation broke. Return reason fields tell you what failed badly enough to be sent back, though the fields are coarse and buyers pick whichever option gets the refund fastest. Comments on your own advertising tell you how the promise reads to people who are not already customers, which is the only place in the corpus where non buyers speak.
Competitor reviews belong in the corpus too. They are the cheapest product research available, because your competitor paid for the launch and their customers wrote the debrief in public.
Sample it rather than reading all of it
The instinct is to read everything or to read only the one star reviews. Both are wrong. Reading everything does not scale past a few hundred items, and reading only the angry ones tells you what breaks without telling you what holds.
Design a sample instead. Stratify across three axes: rating band, so you cover the top, the middle and the bottom rather than the extremes only; SKU, so a single problem product does not colonise the findings; and time, so you can separate a permanent defect from a bad batch or a monsoon window. Take a fixed recent window and hold that window constant every time you repeat the exercise, because comparability across quarters is worth more than volume in any single quarter.
The three star reviews deserve special attention. They are usually written by people who had a real experience and no strong emotion about it, which makes them the most literal and least performative text in the whole set.
Coding is the work
Reading produces impressions. Coding produces findings. The difference is a codebook.
Do a first pass with no categories at all, tagging each item in the customer’s own language. Then collapse those open tags into a stable set of themes and write a one line definition for each, including what does not belong in it. Freeze the codebook before you count anything, because a category that keeps expanding while you code will always look like the biggest theme.
Have a second person code a slice of the same sample independently and compare. Where the two of you disagree, the definition is unclear and needs rewriting. Keep an explicit unclassified bucket and read it at the end, because that bucket is where genuinely new problems hide before they have a name.
Separate two levels while coding: what happened, and why it mattered. “Bottle leaked” is an event. “Leaked into a bag with a laptop in it” is the consequence, and the consequence is what determines whether the customer ever orders again.
One India specific caution. A large share of this text arrives in mixed languages and in transliterated Hindi, Tamil, Marathi and more. Automatic sentiment scoring handles this badly and will quietly drop or misclassify exactly the items you most need. Keep the raw text, code by hand, and use tooling to organise rather than to judge.
Frequency and severity are different axes
This is where most theme analysis goes wrong. The team counts codes, sorts descending, and fixes the top of the list.
The most frequent complaint is very often the least damaging. Slightly slow delivery generates enormous volume and almost no churn. Meanwhile a rare complaint about a product arriving spoiled, or a claim on the pack that did not hold, may end the relationship permanently and take a public one star rating with it.
So score every theme twice. Frequency is the count in the sample. Severity is the consequence, and you should define the ladder in advance: refund issued, order abandoned, customer never returned, public rating damage, safety or regulatory exposure. Plot the two together. High frequency and low severity is a cost problem to be engineered down. Low frequency and high severity is a risk problem to be fixed regardless of the count. Anything high on both should already have an owner.
Route each finding to the function that can act
A theme with no owner dies in the deck. Before the analysis is circulated, map every theme to exactly one function.
Missing information before purchase goes to whoever owns listing content, not to support. Pack failures in transit go to packaging and to the fulfilment partner together. Repeated confusion about how to use the product goes to product and to on pack instructions. A gap between what the advertising promised and what arrived goes to marketing, and it is the finding brands are slowest to accept. Genuine product defects go to sourcing or manufacturing with the sample attached.
Note that this is a different artefact from the reason codes your support team uses to route tickets day to day. That taxonomy exists to close cases efficiently. This one exists to change what you sell and how you describe it, which is why it is reviewed quarterly by brand and product rather than weekly by operations.
Why almost nobody does this properly
It is unglamorous. There is no invoice, no vendor, no fieldwork report, and nothing that feels like a purchase. So it gets handed to whoever is free, with no codebook, no sampling frame and no deadline, and it produces a list of grievances that everyone nods at and nobody acts on.
Fix it with ritual, not with tooling. One named owner. One fixed window each quarter. The same sampling design every time. A frozen codebook that carries forward so themes are comparable across quarters. A written output with themes, frequency, severity, owner and review date. Nothing else in research costs so little and returns so much.