Skip to main content
Materials & Procurement · 30 August 2026 · 7 min read

Where Duplicate Material Codes Come From, and How Plants Clean Them Up

The same seal, stocked under two codes, bought at two prices. Duplicates are not a discipline problem. They are what free-text descriptions do to a material master over twenty years.

Priyansh Srivastava

Co-founder & CEO, Raven

Where Duplicate Material Codes Come From, and How Plants Clean Them Up

A planner is preparing a pump overhaul and needs a mechanical seal. SAP shows nothing in stock, so a purchase request goes out, the vendor quotes a long lead time, and the job moves to next month. The seal was on the shelf the whole time. It had been catalogued years earlier by a different team as SEAL, MECHANICAL; 45MM instead of MECH SEAL 45 MM T-21, under a code nobody thought to search for.

Ask anyone who has worked in stores or maintenance planning at a plant more than a decade old, and they will recognize the scene. Nobody decided to catalogue the same part twice. The material master grew the way the plant grew: a project handover here, a system migration there, three generations of cataloguing conventions layered on top of each other. This piece is about where those duplicates come from, why they resist the obvious fixes, and what a cleanup that lasts actually involves.

Where duplicate codes come from

In most material masters, the description field is the only real handle on a part, and the description is free text. One cataloguer writes the noun first, another writes the maker first, a third abbreviates. A migration truncates long descriptions. A capital project hands over its spares list with its own numbering, and the codes get created in bulk against a deadline. A sister site catalogues the same pump's spares independently, in its own style.

Then the loop closes. A storekeeper searches for a part, phrases it differently than the person who catalogued it, finds nothing, and raises a new code. From that day the plant buys against both. Every failed search is a chance for a new duplicate, and each duplicate starts accumulating its own stock, its own purchase history, and its own vendors.

What a duplicate actually costs

The visible cost is inventory. The same seal held under two codes is working capital parked twice for one insurance policy. The less visible cost is the stockout that happens with stock on the shelf, which is the version that delays a job. And because the purchase history is split across codes, the same part can be bought from two vendors at two prices for years without anyone seeing the spread. The spread is only visible to someone who knows the two codes are one part, and that is precisely what the system does not know.

Why text matching does not clean it up

The tempting fix is to run matching over the descriptions. It disappoints in both directions. Two descriptions of the same part can share almost no words, so real pairs get missed. Two descriptions of different parts can share nearly all of them, so the review drowns in false candidates. Free text is the reason the problem exists; it cannot also be the evidence that resolves it.

The reliable join is the maker part number and the specification, and those mostly do not live in the material master. They live in the vendor documents: the spare parts list, the IOM manual, the GA drawing, the datasheet. A material master cannot be cleaned from inside the material master. The cleanup that holds up is document work, which is why the category names its stages the way it does: de-duplication, standardisation, enrichment, obsolescence review. Each stage is someone reading a document and carrying a value back to a code.

Why cleanups do not stick

Plenty of plants have paid for a cleanse and watched the catalogue drift back within a few years. The reason is rarely the quality of the cleanse. It is that the cleanse was a photograph. Projects keep handing over spares, storekeepers keep raising codes under deadline, OEMs keep superseding part numbers. Without a check at the door, matching each new code against the evidence before it enters, the same loop that created the duplicates starts refilling the catalogue the day the project closes.

What changes when agents do the reading

The bottleneck in all of this was never judgment. A materials engineer shown two codes, one maker part number, and the page it came from can confirm a duplicate in seconds. The bottleneck was the reading: thousands of codes, and a shelf of vendor documents behind each equipment. That reading is work AI agents are now genuinely good at, and it is the work Raven does. The agents read the spares lists, manuals, and drawings, propose duplicate pairs with the evidence attached, and the materials team confirms. Confirmed pairs are valued at the stock held against them, so the cleanup produces a number a plant can act on, not just a tidier catalogue. The corrections go back into SAP or Maximo; nothing migrates.

Where to start

Not with a project plan. Start with an extract. A material master export and a stock report are enough to count candidate duplicates and put a value on them, and that number tells you whether you have an annoyance or a working-capital problem. A plant whose candidates are worth little can stop there, better informed. A plant that finds real money sitting under paired codes has its scope, its business case, and its priority list in the same table.

Either way, the honest sequence is the same: count them, value them, fix the ones the evidence proves, and put a check at the door so the count stays down. The catalogue does not need to be perfect. It needs to stop lying about what is on the shelf.

About the author

Priyansh Srivastava is co-founder and CEO of Raven (YC S22). Raven reads a plant's own documents, builds a verified model of the plant from them, and checks the material master, BOMs, and inventory against it.

Material MasterSpare PartsSAPDuplicate CodesData Cleansing