Skip to content

Phone number normalisation to a single international format before matching

Unified dataUpdated 2026-08-238 min read

In short

Phone number normalisation converts every stored number to the ITU E.164 form, a plus sign, a country code and the national significant number with no spaces or trunk prefix. The country is inferred from the registration country when the number lacks one. Four written variants of one mobile then reduce to one comparable key.

The reminder SMS for a February edition went to 14,206 numbers. The badge file behind it held 12,900 people, and the messaging vendor billed for every send. Nobody could explain the gap in the room, and the explanation was dull: phone number normalisation had never been done, so one mobile written four ways counted as four contacts.

That is the whole of it. The phone field is the one identifier on a registration record that people write in a genuinely free format, and it is also one of the strongest pieces of evidence you have that two rows are the same person, so the cost of leaving it unparsed is high in both directions.

What the same mobile looks like in four systems

Take one delegate registering across your portfolio. The web form in 2023 captured 020 7946 0958. The exhibitor's lead capture app in 2024 stored 0044 20 7946 0958, because the rep typed the international prefix out of habit. The conference platform in 2025 wrote +44 20 7946 0958. The badge collection desk in 2026 wrote 44 (0) 20 7946 0958, brackets and all.

Four strings. Every character differs somewhere. One telephone number.

Compared as text, no two of those agree. Compared after parsing, all four reduce to +442079460958, twelve digits carrying a country code of 44 and a national significant number of 2079460958. The trunk 0 in the national form drops out, because it exists to tell a domestic switch that a national number follows, and it has no place in the international form.

Why does the plus sign carry so much weight?

The target format is not a house convention. The International Telecommunication Union defines it in Recommendation E.164, the international public telecommunication numbering plan, and the 1997 text is direct about the ceiling: "The ITU-T recommends that the maximum number of digits for the international geographic, global services, and network applications should be 15 (excluding the international prefix)." The same document gives the country code one to three digits and leaves the national significant number whatever remains of the fifteen. The recommendation has been through further editions since, including one in 2010.

Two properties make this the right normal form for matching. It is unique, because a country code plus a national significant number identifies one line worldwide with no context needed. And it is context free, which the national form is not. The string 020 7946 0958 only means something once you know it was typed in the United Kingdom. Once you store +442079460958, the row carries its own country and a matcher in another system can compare it without knowing where the registration came from.

The plus sign matters as a marker, since it says the digits that follow start with a country code and are not a local dialling sequence. Storing the digits without it invites somebody downstream to guess wrongly.

Inferring the country when the number does not say

Most rows in a domestic registration file carry no country code at all, because the form never asked for one, and that is where the real work sits. Google's libphonenumber, the parsing library behind the phone fields in Android and a good deal of web tooling, makes the requirement explicit: its parse call takes the number string and a default region, so parsing 020 7946 0958 as a British number and as an Irish one are two different operations with two different results.

The region has to come from somewhere on the row. In order of reliability, the registration country the delegate selected, the country on the billing or company address, the country implied by the top level domain of a corporate email, and last the country the event is held in. The last one is a guess, and it is right often enough to be worth using and wrong often enough to be worth recording. Store the region source alongside the parsed number so that a later review can tell an inferred +44 from a typed one.

libphonenumber also separates two questions that teams tend to collapse. Its possible check looks at length information alone. Its validity check uses length and prefix, so it knows that a given national destination code has been allocated. A number that is possible but not valid is usually a typo or a placeholder, and those rows want flagging rather than deleting, because the rest of the record is often fine.

Should you compare the last nine digits or the whole number?

A common shortcut is to strip everything to digits and compare the last nine, on the reasoning that the tail of the number survives whatever mess precedes it. It handles the trunk zero, the international prefix, the brackets and a missing country code in one move, and it takes about four lines of code.

Work out what it costs. Nine digits give you a billion distinct suffixes. A file of 200,000 distinct numbers contains 19,999,900,000 pairs, just under twenty billion, and if suffixes were spread evenly across that billion you would expect around twenty accidental agreements. Twenty false pairs in a portfolio file sounds survivable until you remember that real numbers are nowhere near evenly spread, because area codes and mobile prefixes concentrate them heavily, so twenty is a floor. There is a structural failure too. Where a numbering plan uses eight national digits, a nine digit window reaches back into the country code, and the key you are comparing is part country and part subscriber.

My view is that the suffix is a decent blocking key and a bad merge rule. Use it to propose candidate pairs cheaply, hand the pair to the scoring stage, and let the comparison run on the full parsed number where the country code either agrees or does not. That keeps the cheap recall and gives up none of the precision. The design of the key itself belongs with blocking key design in J9, which treats the recall and cost trade properly.

Extensions, pauses and the noise around the number

Desk numbers arrive with tails: x2214, ext 2214, a comma for a dialling pause, a second number after a slash because the person gave both. Parse the tail off and keep it in its own column. An extension is real information about the person and useless for matching, since two colleagues on the same switchboard share every digit up to the extension and differ only in it.

The standards treat it as a separate thing too, which is a useful check on the instinct to concatenate. RFC 3966, published by the IETF in 2004, says that global numbers "MUST be composed with the country (CC) and national (NSN) numbers as specified in E.123 and E.164", and puts extensions outside that number in their own parameter, on the grounds that they "identify stations behind a non-ISDN PBX". Two columns, and the same split the standard already makes.

That produces an ordering worth being strict about. Parse first, then split the extension, then compare. Stripping non digits before parsing destroys the plus sign and the brackets that told you what kind of number you had, and once a leading 00 has been flattened into the digit run you cannot tell it from a national prefix. Parsing into parts before any comparison runs is the same discipline address matching in J22 depends on, for the same reason: the structure carries the meaning, and a flattened string has thrown it away.

What to store, and what never to overwrite

Five columns cover it. The raw string exactly as the registrant typed it. The parsed E.164 value. The region used for parsing and where that region came from. A validity flag from the parser. The version of the normalisation rules that produced the row.

The raw string is the one that must never be overwritten. Normalisation is lossy on purpose, parsers change behaviour between releases, and a country inference you made in 2024 may look wrong in 2026 when the same delegate turns up with an address in another country. Keeping the input means you can recompute every key from scratch and know which rows moved, which is the same discipline that email address normalisation in J20 needs for the same reason.

The number you dial is always the raw one. The number you match on is always the derived one.

Where this stops

Phone agreement is strong evidence and it has a failure mode that no amount of parsing fixes. A switchboard number identifies a building. If half your exhibitor contacts give the main company line, every one of them agrees on the strongest field in the record, and a matcher that weights phone heavily will merge an entire sales team into one contact. The defence is a frequency check on the parsed value: any number appearing on more than a handful of records is a shared line, and it should be excluded from the match key while staying on the record. This is the same treatment role addresses need on the email side.

The second limit is that numbers get reassigned. A mobile released by one subscriber can be reissued, and the interval varies by regulator and operator. Over a five year archive that is rare and it is not zero, so an agreement on phone alone, with a disagreement on name and employer, deserves the review queue rather than an automatic merge.

The third is that a well formed number can be confidently wrong. If the region was inferred from the event country and the delegate flew in, you have produced a valid number in the wrong country, and it will fail to match the same person's correctly parsed row from another show. That failure is silent. The region source column is what lets you find it later, and it costs one column to keep.

This week, take one registration export and produce three counts: distinct raw phone strings, distinct strings after removing spaces and punctuation, and distinct values after parsing to E.164 with the registration country as the default region. In the file behind that 14,206 send, the third count is the number of people you can actually reach, and the difference between the first and the third is what you have been paying for. Getting that far needs no unified data programme, only a parser and an afternoon, and it tells you whether the comparison you use per field in J16 is even seeing the phone evidence it should be.

Questions people ask about phone number normalisation

What is E.164 format for a phone number?
E.164 is the ITU numbering plan for international public telecommunication numbers. A number in that form carries a country code of one to three digits followed by the national significant number, with no trunk prefix, spaces or punctuation. The 1997 text of the recommendation caps the digits at fifteen, excluding the international prefix.
How do you normalise a phone number with no country code?
Infer the region before parsing. Google's libphonenumber requires a default region for exactly this case, and the registration country on the same row is usually the best guess, with the event country as a fallback. Record which source supplied the region, because a wrong guess produces a well formed number for the wrong country.
Is matching on the last nine digits of a phone number safe?
It is a reasonable blocking key and a poor merge rule. Nine digits allow a billion suffixes, so a file of 200,000 distinct numbers yields around twenty accidental agreements even under an even spread, and real numbers cluster by area code. Use the suffix to propose pairs, then compare the full parsed number.

Related reading

All identity resolution articles