What it means
Two inspectors examine the same returned jacket: one calls it new, the other notes wear and calls it open-box. If the business does not resolve the disagreement, a customer may later receive a product described incorrectly.
A product-specific rubric covering tags, packaging, visible wear, function, accessories and safety condition reduces that room for interpretation, because a vague "looks fine" grade invites guesswork. Shopify's returns-management guidance describes inspection, grading and disposition as separate steps, so a consistent grade supports a credible decision about resale, repair or another route.
Shopify's reverse-logistics guide also discusses testing returned inventory and ways to make grading more consistent, although an automated suggestion still needs validation against actual condition standards. Preserve the original item ID for each unit, because returned units from one order may be in different conditions and a single parcel-level grade can hide that.
To measure consistency, have a sample of returns graded independently by two reviewers, blind to each other's grade, and record both initial grades rather than only the consensus outcome. An illustrative exact-agreement rate is the sample items receiving the same grade from both reviewers divided by all double-graded items, so 90 agreements out of 100 gives 90%.
Also show the material disagreements, for example in a confusion matrix, because a split between new and unsellable matters far more than a split between adjacent clearance categories. Agreement is not accuracy, since two reviewers may both apply a flawed rule.
Periodic expert audits, customer complaints and repeat returns test whether the grades are appropriate, and a high share of restocked items re-returned for condition issues can show the rubric misses defects even when reviewers agree. Sample across product types and note whether each item was physically inspected or judged from photos, because a single blended number can hide a weak category such as electronics that need function tests.
Calibrate before peak periods by reviewing examples of each grade together, including borderline cases, and investigate disagreement causes such as unclear terms, lighting, test equipment or training. A specialist can settle a disputed item and update future examples, while the initial disagreement stays in the metric.
Targeted samples and mandatory checks for high-value or regulated goods are often more practical than double-grading every low-risk item. Watch for commercial pressure, because a warehouse target to restock quickly can encourage overly generous grades, and keep a disputed safety-related item out of available stock until it is resolved.
Treat the resale grade and the customer's refund rights as related but separate decisions, and record the rubric version used for each item so old and new periods can be compared cautiously. For an owner, grading consistency tests whether return-condition decisions are repeatable and defensible without pretending that agreement alone proves the product is sound.
In practice
Real-world examples.
Example
A fashion retailer asks two reviewers to grade the same 100 returned jackets without seeing each other's labels. Ninety receive the same category, giving a 90% exact-agreement rate. The ten disagreements are listed so trainers can see which tag and wear issues cause the confusion.
Example
An electronics reseller finds that one reviewer grades a returned tablet as new from photos, while another finds a faulty charging port in a function test. Because new versus unsellable is the most serious type of disagreement, the item goes to a specialist before it can be listed. The reseller also records whether each grade came from photos or physical testing.
Example
A homeware seller notices that lamps graded as like-new keep coming back for scratched bases. The repeat returns suggest the rubric misses a defect that reviewers consistently overlook. The team reviews the rubric with an expert and adds a base-surface check.
Formula
Calculation
Exact agreement rate = double-graded items with matching independent grades / all double-graded items x 100
Worked example. A fictional returns team double-grades 100 returned items in a month, and the two reviewers give the same grade on 90 of them.
- Exact agreement rate = 90 / 100 x 100 = 90%.
- The 10 disagreements are then split by severity: 7 differ by one adjacent category, such as open-box versus clearance, and 3 differ between new and unsellable.
- Those 3 items are 3 / 100 x 100 = 3% of the sample and deserve the closest review.
A high rate shows that reviewers apply the rubric in the same way; it does not prove the grades are right.Case study
Seen in the real world.
In this entirely fictional example, Cedar Wear tests its return grades and finds disagreement on missing tags. It clarifies the rubric with examples and repeats a blinded sample review. It holds disputed items from new-condition resale until an authorised decision is made.
In the first blinded sample, the two reviewers agreed on 78 of 100 coats, mostly because the old rubric did not say whether a missing spare button counts as an accessory defect. After photographed examples were added and staff retrained, a second sample of 100 coats reached 92 agreements. Cedar Wear still sends a small sample to an expert each month, because agreement between reviewers shows repeatability but not that the grades are correct.
Watch out
Common mistakes.
- Treating two agreeing inspectors as proof the grade is correct.
- Sampling only easy product categories and claiming company-wide consistency.
- Allowing a disputed safety-related grade to release stock before review.
Questions
People also ask.
Is agreement the same as accuracy?
No. Reviewers can consistently apply a mistaken rule.
Should every return be double-graded?
Not necessarily. Use risk-based sampling and stronger checks where needed.
Does a condition grade decide the refund?
Not alone; customer rights and transaction terms also matter.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%