Automated product coding

Every product in the file, coded to your schema, every cycle

The attributes and their allowed values are agreed once, then applied to every product as the file changes. Coverage, completeness and accuracy are measured and reported rather than asserted.

One file, coded end to end

Fortune 100 CPGAutomated product coding
Products coded every cycle
80,076

The whole file, every cycle, on a schedule. Coding by hand had topped out near 68% of it, and the products left behind were missing from every analysis that followed.

Delivered weekly · four times the contracted cadence
The state of the fileI · III
Measure · where it stands · how it is checked
ICoverage100%
Reached by hand 68%26,185 products recovered

About a third of the file had never been tagged at all. Those products are now in every count, every cut and every report.

Blind spot closed
IICompleteness99.95%

On core attributes, dollar weighted, so completeness is measured against what the business actually sells rather than against a flat product count.

Core attributes
IIIAccuracy98%

Scored against an independent test set rather than self reported, which is the difference between a quality claim and a quality measurement.

Independently validated

The number that changed the work was not the accuracy. It was that nobody has to choose which products to leave out any more.

Figures from one published engagement with a Fortune 100 CPG. The catalog is maintained to the client's own schema, so the coverage holds as the file changes.

Why the coverage holds

Automated coding is only worth having if it stays correct as the file moves. Three things keep it there.

Defined once
your schema, not ours

The attributes and their allowed values are agreed up front and applied to every product the same way. Nothing is re-interpreted batch to batch, so one cycle is comparable to the next and to the one before it.

Run on a schedule
weekly, not on request

New products are coded as they arrive rather than in an annual push. The file is current when someone queries it, instead of current as of whenever the last project finished.

Measured against a test set
independent, not self reported

Accuracy is scored against a held-out set the coding pass never saw. That is the difference between a quality claim and a quality measurement, and it is the number that survives a stakeholder asking how you know.

One retailer catalog, three measures

A different engagement, on a retailer file where the product label was often the only description a product carried.

One retailer catalogThree measures
01

Non-informative labels removed

71%

of labels that described nothing useful, in a catalog where the label was often the only description a product carried.

Retailer attribution →
02

Improvement in categorization

22%

in the same pass, because a product that is described properly is a product that can be filed properly.

Retailer attribution →
03

Hours to attribute a new product

72

from detection, so a new item is described and filed within three days rather than waiting for the next coding cycle.

Retailer attribution →
Figures as reported in the linked case study

Harmonya Helps You Make Faster, More Confident Decisions

A coded file is infrastructure. What it is worth is whatever the systems downstream can finally do once every product is described the same way.

Product Data Enrichment

Take maintained attributes into search, audiences, space planning and reporting, on your own schema.

See the use case

Request a Demo

Thirty minutes on your categories. How the record gets built, what it holds, and the questions your team could put to it.