Skip to content

Address matching for exhibitor and attendee records across five years of files

Unified dataUpdated 2026-08-238 min read

In short

Address matching should parse each address into components before any comparison runs, then compare postcode and building number exactly and street name approximately. Whole string comparison rewards the city and country tokens every record in a national file shares, so it scores two tenants of one building above two spellings of one address.

A commercial director asks how many distinct companies have exhibited with the portfolio since 2021. The exhibitor tables hold 6,400 stand bookings across five editions of four shows, the company name field is a mess everybody already knows about, and somebody suggests the address will settle it, since firms move less often than they rebrand.

They are right about the premise and wrong about the difficulty. Address matching is the part of the record where the data looks structured and behaves like free text, and the usual approach, comparing the two address strings as strings, fails in a specific and predictable direction.

What five years of address files actually contain

One exhibitor, four bookings, four addresses typed by four people.

Unit 4, Kingsway House, 15 Kingsway, London WC2B 6UN, United Kingdom 15 Kingsway, Holborn, WC2B 6UN Unit 4 Kingsway Hse, 15 Kingsway, London, WC2B 6UN Kingsway House, Kingsway, London

Alongside them, in the same file, sits another exhibitor at Unit 7 of the same building, a separate company with its own stand and its own contract.

The four rows above are one firm. The Unit 7 row is a different firm. Any method you use has to get both of those calls right, and the second one is where the string comparison approach comes apart.

Why does whole string comparison reward the wrong agreement?

Score the pairs by counting how many tokens two addresses share against the total number of distinct tokens between them, the measure that jaccard token overlap in J13 treats in full.

The first row holds ten distinct tokens: unit, 4, kingsway, house, 15, london, wc2b, 6un, united, kingdom. The second holds five: 15, kingsway, holborn, wc2b, 6un. They share four. Between them they use eleven distinct tokens, so the overlap is 4 divided by 11, which is 0.36.

Now the Unit 7 record, which is a different company: unit, 7, kingsway, house, 15, london, wc2b, 6un, united, kingdom. Ten tokens, of which nine are shared with the first row, and eleven distinct tokens between them. That is 9 divided by 11, or 0.82.

The wrong pair scores more than twice the right pair. Nothing about the measure is broken. The long tail that every record in a British file repeats, the city and the country and the building name, is doing the work, while the single token that actually decides the question, the unit number, contributes one token of disagreement out of eleven.

This is why address standardisation and parsing sit ahead of comparison in the record linkage literature rather than beside it. Winkler, writing for the US Census Bureau in 2006, describes the sequence directly: standardisation replaces commonly occurring words with a single spelling, so Street becomes St, and the string is then parsed into components that can be compared. He notes that the address software developed by the Census Bureau's Geography Division produces approximately fifty components, of which the matching software outputs only the most important.

Parsing first, and what you get for it

You do not have to write the parser. The openvenues libpostal documentation reports that its parser "achieves very high accuracy on held-out data, currently 99.45% correct full parses", where a full parse counts only if every token in the address is labelled correctly. It was trained on over a billion addresses from OpenStreetMap and OpenAddresses, and it labels the components you need for comparison, among them house number, road, unit, postcode, city, state and country.

Run the four rows through it and the components separate cleanly. House number 15 in three of the four. Road kingsway in all four. Postcode wc2b 6un in three. Unit 4 in two of them, written once as Unit 4 and once as Unit 4 with the building name abbreviated to Hse. The fourth row, the one with no number and no postcode, keeps road and city and nothing else.

A 99.45 per cent full parse rate is a headline figure on held out data, and your exhibitor file is not the held out data. Non domestic addresses, free text in the wrong field and building names that look like street names all knock it down. Measure it on a sample of your own rows before you trust it, which for a thousand rows is an afternoon of eyeballing.

What each component actually deserves

The comparison plan follows from what each part does.

Postcode, compared exactly. It is short, it is checkable against a published file, and two records with different postcodes are almost never the same premises. Where the postcode is present on both sides and disagrees, that is close to decisive on its own.

Building number, compared exactly. Same argument, and it is the field the whole string method drowns. 15 against 15 is agreement. 15 against 17 is a different building, and 15 against 15A usually is too.

Unit or suite, compared exactly where present. This is what separates two tenants. Treat a missing unit on one side as unknown rather than as agreement, because most rows will not carry one.

Street name, compared approximately. Here the typing errors live, and here an approximate comparison earns its keep, which is the case jaro winkler similarity in J12 was built for.

City, county and country, given very little weight. In a file where 80 per cent of registrations are domestic, agreement on country tells you almost nothing, and agreement on city tells you slightly more than nothing.

Applied to the same two pairs, the picture inverts. The true pair agrees on house number, road and postcode, and disagrees on nothing, since the sparser row simply lacks a unit. The Unit 7 pair agrees on everything and disagrees on the one component that identifies the tenant, which is the component the plan weights heavily.

Does any of this survive a non domestic file?

An international trade show breaks the tidy version of the plan, because the components stop appearing in the order a British or American parser expects. German and Dutch addresses put the number after the street name. Japanese addresses work from the larger administrative unit down. Postcode formats vary from four digits to the alphanumeric British form, so a validity check tuned to one country rejects perfectly good rows from another.

Two things make this tractable. The first is that the parser already knows: libpostal's documentation describes training on addresses from every inhabited country and normalisation support in more than sixty languages, so the component labels come back correctly even when the surface order differs from what you expected. The second is that the comparison plan works on labels, so once house number is house number, it does not matter whether it was written before or after the road.

What does need per country handling is the exactness rule. Comparing postcodes exactly is right where the postcode identifies a small area, and it is too strict where a single code covers a whole town, which is the case in several plans that use four digits. In those countries the postcode behaves more like the city field, agreement is common and cheap, and the building number and street have to carry the decision. Set the weight per country if your file has a serious international tail, and check the distribution before you assume the domestic settings transfer.

What happens when the postcode is missing or wrong

Some share of the rows in any hand keyed file will have no postcode at all, and a smaller share will carry one that belongs to a different address entirely, usually because somebody pasted a head office postcode under a site address. Count both on your own file before designing around either.

Two rules keep this from poisoning the run. Treat missing as missing, never as disagreement, and let the score fall by less than it would for a genuine mismatch. And check the postcode against the city where you can, because a postcode from one region sitting under a city in another is a data entry error rather than evidence, and it should be neutralised rather than trusted.

The related trap is the address that is real, correct and useless. Winkler's 2006 overview gives the case for business lists: the same firm can appear under its trading premises, the owner's home address and the accountant's address, all three being genuine. Company records collect these. An exhibitor booking from a stand contractor, a billing address at a finance bureau and a registered office at a company formation agent are three addresses for one exhibitor, and no address comparison will ever link them, because they have nothing in common. That case needs the company name, which is company name normalisation in J17, or the domain on the contact email.

Where this stops

Address agreement is evidence about a location. It becomes evidence about a company or a person only through an assumption, and the assumption fails in three ways that event data produces constantly.

Serviced offices and coworking buildings put hundreds of unrelated firms behind one street address and one postcode. A frequency count catches them: any address appearing on more than a handful of records from unrelated companies is a shared building, and it should stop contributing positive evidence while staying on the record. The same treatment applies to accountancy and formation agent addresses, which cluster even harder.

Attendee addresses are worse than exhibitor addresses, because a delegate who changes employer changes address while remaining the same person, and a delegate who gives a home address on one registration and a work address on the next produces two correct addresses with nothing in common. Address disagreement is weak evidence against a match for people and moderately strong evidence against for companies, and a single scoring configuration used for both will be wrong for one of them.

The third limit is coverage. Postcode is the strongest component and it is also the most often absent, so match rates computed on rows that have one will look better than the file deserves. Report the parse rate and the postcode presence rate next to any match rate you publish, or the number will be read as a quality measure when it is partly a completeness measure.

Take 200 rows from your oldest exhibitor file this week, run them through a parser, and count how many produce a house number and a postcode. That single number tells you whether component level unified data work on addresses is worth starting, and it is the same test the phone number normalisation work in J21 needs before anyone commits to a plan.

Questions people ask about address matching

Why parse an address before matching it?
Because the parts carry different weight. A building number that disagrees is close to decisive, while a city name that agrees is worth almost nothing in a file where most records share it. A whole string comparison averages those signals together and loses both, and it cannot tell a unit number from a house number.
How accurate is automatic address parsing?
The openvenues libpostal documentation reports 99.45 per cent correct full parses on held out data, where credit is given only when every token in the address is labelled correctly. The parser was trained on over a billion addresses and labels components including house number, road, unit, postcode and city.
Which address components should be compared exactly?
Postcode and building number. Both are short, both are checkable, and a disagreement in either usually means a different premises. Street name deserves an approximate comparison because it absorbs typing errors and abbreviations, and city and country deserve very little weight in a file where nearly every record repeats them.

Related reading

All identity resolution articles