Skip to content

An industry taxonomy mapping table beats forcing every show to renumber

Data platformUpdated 2026-08-188 min read

In short

An industry taxonomy mapping table holds one row per local category value per show, linking it to a value in a conformed master sector list. Local labels stay on show screens and contracts, portfolio reports group on the master key, and the unmapped rows form a visible backlog instead of silent data loss.

The portfolio project lands on a show director's desk as a request to renumber. From next edition, please classify exhibitors using the group master list of 28 sectors instead of your own 41 categories.

She says yes in the meeting. Six months later her team is still using the 41, because her exhibitors recognise those labels, her rate card is built around them and hall 3 is named after one of them. An industry taxonomy mapping table is the design that gets the portfolio what it needs without asking her to do any of that, and it is a smaller piece of engineering than the renumbering it replaces.

What the mapping table actually holds

One row per local category value, per source.

The columns are unglamorous. Source system or show, local code, local label, master sector key, a review flag, the date the mapping was made and the person who made it. That is seven columns and it is enough.

The master sector list is a separate small dimension: one row per conformed sector, with a stable key that never gets reused. Both tables are hand-curated, both are small, and both belong in version control so a change to a mapping arrives as a diff somebody approved.

What makes the design work is that the local label survives. It appears on the exhibitor's profile, on the floorplan, in the show directory and on the contract, exactly as it does today. The master key is invisible to everyone except the portfolio report.

Three shows, 132 local values, 28 master sectors

Take three shows in one portfolio.

Show A has 41 local categories, built by its sales team over a decade. Show B has 62, inherited from an acquisition and structured around a printed directory. Show C, the smallest, has 29. That is 132 local values in total, and the mapping table therefore has 132 rows.

Map them to a master list of 28 sectors. On average 132 divided by 28 is 4.7 local labels per master sector, which is the compression the portfolio report has been missing.

Nine local values map to nothing. In practice they are Show B's directory headings that describe services rather than sectors, plus two of Show A's categories that turned out to be feature areas. So 123 of 132 values are mapped, which is 93.2 per cent coverage by value.

Coverage by value is the wrong number to publish on its own, because values are not equally used. Weight it. Across the three shows those 9 unmapped values carry 214 exhibitor records out of 4,540, which is 4.7 per cent of exhibitors. That figure is what a portfolio report should carry as an unclassified bucket, with the count visible.

The two numbers together are the backlog and its priority. Nine rows to resolve, worth 214 exhibitor records, and a person can work through nine rows in an hour once somebody decides what a feature area is.

Why does renumbering fail when everyone agreed to it?

Because the cost and the benefit sit with different people, and the cost is recurring while the benefit is a report.

A renumbering asks the show team to change the vocabulary their customers use. Exhibitors search the directory by category. Sponsors buy category-level packages. The sales pipeline is segmented by it. Renumbering breaks all of that for one edition and gains the show nothing, so it competes with everything else on that team's list and loses.

There is a second reason that is less about politics. Local lists carry real information about their market. A machine tools show with 41 categories has 41 because that is how buyers in that market think, and a 28-value master list built to span eight verticals cannot preserve that granularity without becoming a 300-value list nobody maintains. Collapsing to the master for portfolio reporting is a deliberate loss of detail, and it should be reversible, which a mapping table makes it and a renumbering does not.

The statistical agencies settled this argument decades ago and their answer is a mapping table. Eurostat and its partners in the European Statistical System publish a correspondence table between NACE Rev. 2 and NACE Rev. 2.1, and Eurostat's own description of it is that the table "describes, for each NACE Rev. 2 heading, how activities currently classified therein should be classified in NACE Rev. 2.1". Nobody asked every national statistical office to retype its history. The UN Statistics Division does the same for ISIC, whose Revision 5 was endorsed by the UN Statistical Commission at its 54th session in 2023, and publishes correspondence tables alongside it so series built on earlier revisions remain usable.

If the institutions that maintain the world's industrial classifications treat mapping as the normal mechanism for reconciling two code lists, an exhibitions group with three of them can stop treating it as a compromise.

Where the mapping gets hard

Two cases break the one-to-one link, and both need a decision rather than a cleverer schema.

A local value that spans two master sectors. Show B has a category covering both commercial refrigeration and food service equipment, and the master list separates them. You can attach an allocation weight, splitting that category 60 to 40 across the two masters, or you can pick the dominant one and record the choice. I would pick one and record it, because allocation weights propagate into every measure and nobody downstream will remember they are there. Where the split genuinely matters commercially, the honest fix is at source: ask that show to split the category next edition.

A master sector with no local value anywhere. This is the reverse gap and it is worth watching, because a master list with sectors nothing maps to is a list somebody designed from a market map rather than from the portfolio. Prune those. A master value that has never been used is a row that will eventually be used by mistake.

Then there is the drift problem. Every edition brings new local categories, and each one arrives unmapped by default. That is the correct behaviour, because a new category silently absorbed into an existing master is a change to historical grouping that nobody sees. Load it unmapped, count it, and put it on the list.

What does this do to the portfolio report?

The report changes shape in three ways worth expecting.

Sector totals stop double counting, because two local labels for the same market now land on one master key. The mechanism and the damage it does are the general conformity argument, and the mapping table is the specific instrument for the category axis.

An unclassified bucket appears with a real number in it, 214 records or 4.7 per cent in the example above. Some people will treat that as a defect in the report. It is a measurement of how much of the portfolio nobody has classified, and it belongs on the page.

Drill-down changes behaviour. A portfolio user clicking into a master sector should see the local labels underneath it, with the show they came from. That single interface decision is what keeps show teams trusting the roll up, because the first thing any show director does with a portfolio number is check whether her show is represented correctly in it.

Who owns the table

Somebody has to, and this is where these projects die quietly.

The mapping table needs an owner with a monthly slot: map the new local values from the last edition loaded, review anything flagged, and publish the unmapped count. Call it two hours a month plus an hour per edition. Without a named owner the coverage rate falls edition by edition and nobody notices until a portfolio report shows an unclassified bucket at 19 per cent.

The master list needs a stricter gate, because changing it changes historical grouping. Merging two master sectors or splitting one is a versioned change with a date attached, and a five-year comparison that spans a version change needs a note on the chart. Treat it like any other schema change in the data platform, with a diff and a named approver.

Where mapping is done by a model rather than a person, which is reasonable at scale, the review queue and the threshold behind it are their own design problem and the same governance still applies to the output.

Where this stops

A mapping table makes categories comparable and does not make them right.

If Show A's sales team assigns categories by whichever one the exhibitor picked from a dropdown three years ago, the mapping propagates that with perfect fidelity. Garbage maps cleanly. The coverage rate will look excellent and the sector totals will still be wrong, and the only way to find out is to sample fifty exhibitor records and check their categories by hand against what those firms actually sell.

The master list also encodes a point of view about the market that will age. A vertical splits, a new category becomes material, and 28 sectors turn out to be 31. That is normal and it is why the version and the date matter more than getting the first cut perfect.

And a mapping is one-to-one at best. Companies that genuinely operate across four sectors get classified into one, and any analysis of sector concentration inherits that flattening. Where the company dimension is itself changing underneath, with firms merging and rebranding between editions, type 2 history on the company row is what stops the categories being attached to the wrong entity.

Export the distinct category values from your two largest shows this week, put them in a spreadsheet with a blank master column, and fill in as many as you can in one sitting. Count the blanks left at the end. That count is your backlog, and it is almost always smaller than the renumbering project somebody has proposed instead.

Questions people ask about industry taxonomy mapping table

Why not just make every show use the same category list?
Because the show teams stop using it. A sales team whose exhibitors recognise a category name, whose rate card is structured around it and whose floorplan zones are named after it gains nothing from renumbering. They agree in the meeting and keep the old list in the operational system, which leaves the master unpopulated.
What columns does an industry taxonomy mapping table need?
Show or source system, local code, local label, master sector key, a confidence or review flag, the date the mapping was made and who made it. A one-to-one link is enough for most rows, and rows that genuinely split across two master sectors need an allocation weight or an explicit decision to pick one.
What do you do with local categories that map to nothing?
Leave them unmapped and publish the count. Nine unmapped values covering 214 exhibitor records out of 4,540 is 4.7 per cent of the portfolio, which is a number a report can carry honestly in an unclassified bucket. Distributing those rows across known sectors in proportion invents data with a plausible shape.

Related reading

All data platform articles