What to do with the unattributed registrations sitting in your file
Unattributed registrations are rows whose source column is empty or unusable, and they come from four causes: campaigns that were never tagged, tags lost between the click and the record, genuinely unprompted arrivals, and records keyed in offline. Sizing each cause separately is what turns the bucket into work.
The post show deck has a pie chart on it. Email, paid search, paid social, partner, referral, and a slice bigger than any of them labelled unknown. Somebody in the room asks what is in the unknown slice and the answer, delivered with a shrug, is that it is mostly direct.
Unattributed registrations are the largest single category on many shows, and treating them as one thing is what keeps them large. There are four distinct causes underneath that slice. They have different sizes, different fixes, and two of them are not defects at all.
Four causes inside one bucket
Never tagged. A campaign ran and nobody put parameters on the links. The agency's retargeting, the exhibitor toolkit, the speaker's post, the association's member mailing. The click arrived clean and the row has an empty source column because there was never anything to lose.
Lost in transit. The link was tagged and the value died somewhere between the click and the saved record. A redirect that dropped the query string, a form step that did not carry the hidden field, a cookie that expired. These are the ones that make paid channels look weak.
Genuinely direct. Somebody typed your url, or used a bookmark from last year, or searched your show by name and clicked the result they already knew. Nothing is broken. This is what a strong brand looks like in a registration file.
Offline entry. Records created by a person rather than a browser: the registration desk on site, the call handler who took a booking, a bulk upload of an association member list, an exhibitor's guest allocation typed in by an account manager.
The four behave differently between editions, and lumping them together means a bucket that shrinks for a good reason looks the same as one that shrinks because you stopped taking phone bookings.
Sizing each cause over one edition
Take one complete edition, 11,400 registrations, of which 3,192 have no usable source. That is 28.0 per cent, and here is how the decomposition is built.
Offline entry is easiest, because those rows have a fingerprint. Filter on the created_by field, the registration type, or the source system flag. Say that returns 432 rows: on site desk registrations, phone bookings and two bulk uploads.
If your platform has no such field, the timestamps will give you most of it anyway. Bulk uploads arrive as dozens of rows sharing a created timestamp to the second, which no human traffic produces. Desk registrations cluster inside show hours on show days. Phone bookings cluster inside office hours in one timezone. Each of those patterns is a query rather than a project, and between them they usually recover the offline slice to within a few dozen rows.
Genuinely direct needs a proxy. Cross tab the remaining untagged rows against whether the email address appears in a prior edition's file. Returning registrants who arrive with no campaign value are overwhelmingly people who already know the show. Say that is 720 rows.
Lost in transit comes from a measurement rather than a filter. Run a hop test on one live link, as B2 sets out, and you get a survival rate. If tagged clicks that converted numbered about 2,650 and end to end survival measured 66 per cent, roughly 890 tags were lost in the relay.
Never tagged is the remainder, and it is worth checking rather than assuming. 3,192 minus 432, minus 720, minus 890 leaves 1,150. Confirm it against the campaign calendar: which activity ran without tags, and does the timing of those 1,150 rows match. If the residual does not match anything on the calendar, one of the other three estimates is wrong.
That gives 1,150 never tagged, 890 lost in transit, 720 genuinely direct, and 432 offline. The bucket is now four numbers that different people can act on.
Write the method down alongside the numbers, including which cause was estimated as a residual, because next edition somebody will rerun this and get a different answer, and the only way to tell a real improvement from a changed method is to have the method on paper.
Why not spread the unknown across channels?
Because it changes the decisions and adds no information.
Take the same edition. Known registrations are 8,208, and paid social holds 22 per cent of them, which is 1,806 registrations against a spend of 48,000. Cost per registration is 26.58.
Now allocate the 3,192 unknown rows proportionally. Paid social's share of the unknown is 22 per cent, which is 702 registrations, taking it to 2,508. Cost per registration falls to 19.14. Nothing was measured. A media buyer looking at 19.14 against a target of 22.00 renews the budget, and a media buyer looking at 26.58 does not.
The assumption underneath proportional allocation is that missing attribution is random with respect to channel, and it is the opposite of random. Tags are lost at hops that particular channels use, browsers cap cookies in ways that hit long consideration windows hardest, and offline entry has no digital channel behind it at all. Spreading the unknown proportionally takes the channels that are best measured and hands them credit for the failures of the ones that are worst measured.
Report unknown as its own line, with its four components underneath it, and let the reader see how much of the picture is missing. That is also the number that makes the case for fixing anything.
Which cause is cheapest to fix?
Rank by registrations recovered per unit of effort rather than by size.
Never tagged is 1,150 rows and costs nothing but process: a link builder, a rule that no campaign goes live without tags, and somebody who checks. That is 10.1 percentage points of the file recovered for the price of a habit.
Lost in transit is 890 rows and costs engineering time, sometimes on a platform you do not control. Worth doing, slower to land, and the hop arithmetic tells you which single fix returns most.
Genuinely direct at 720 rows is not recoverable and should not be a target. Google's default channel definitions, current in 2026, put the source at exactly (direct) and the medium at (not set) or (none), which is a definition by absence and cannot be improved by better tagging. You can learn more about it, because the direct channel splits four ways itself as B8 works through, and some of what sits in there is peer forwarding you could instrument.
Offline entry at 432 rows is already attributed if somebody writes it down, and the cost of writing it down is one dropdown on the internal registration screen with six options and a default of not asked. The registration desk knows whether a walk up came from a flyer. The call handler knows what the caller mentioned. Mediahawk's State of call tracking 2025 survey found 76 per cent of marketers naming an understanding of channel and campaign effectiveness as a top reason for tracking calls, and 67 per cent expecting call volumes to rise that year, which is a reminder that the phone is still a live acquisition route and that its records are usually the least structured thing in the file.
Setting a reduction target that survives the next edition
A target of zero is wrong, because two of the four causes are not defects. The floor is genuinely direct plus offline entry, which in this file is 1,152 rows, or 10.1 per cent.
So a defensible target for the next edition is 28.0 per cent down to 18.0 per cent, made of the never tagged bucket in full and roughly a third of the transit losses. Write it as a number of registrations rather than a percentage, because the denominator will move: 1,150 plus 290 is 1,440 rows recovered.
Then track it as a metric rather than a project. Measuring the coverage rate each edition as a data quality number in its own right is B11's subject, and it is the difference between a fix that holds and one that decays the moment the person who cared moves to another show.
Where the decomposition stops
The four causes are not perfectly separable, and the seams are real. A returning registrant who clicked an untagged email is both never tagged and genuinely direct, and the method above will put them in one bucket by rule rather than by truth.
The estimates also lean on one hop test taken on one path on one day. If the registration platform changes mid campaign, the transit loss estimate is wrong for part of the edition and there is no way to know which part after the fact.
And none of this recovers a single registration retrospectively. The decomposition tells you what the next edition should look like, which is the only thing it can honestly do. For the current edition, the choice is between reporting an unknown share and inventing an attribution, and the first one is the only defensible option. Where the unknown share is large enough to block a decision, asking registrants directly as B10 describes gives you a second, independent read to set beside the tracked one.
Start by counting rather than by fixing. Run one query against your current edition that counts rows with no usable source value, grouped by week and by registration type, and put the resulting percentage on the next campaign call. The number is usually larger than the room expects, and once it is on a slide the argument about whether it is worth the engineering time takes about five minutes. Everything else in acquisition and attribution sits downstream of that one figure, including the channel rules B5 covers, which cannot classify a row that carries nothing.
Questions people ask about unattributed registrations
- What percentage of registrations are usually unattributed?
- There is no industry figure worth quoting, because the number depends entirely on how a show is tagged and where its form is hosted. Measure your own, over one complete edition, and treat the first reading as a baseline rather than a benchmark. What matters is the direction between editions and which of the four causes is largest.
- Should you spread unattributed registrations across channels proportionally?
- No. Proportional allocation assumes the unknown registrations came from the same mix as the known ones, and the causes of missing attribution are channel specific, so that assumption is almost always false. Reporting the unknown share as its own line is honest, and it keeps cost per registration figures from being quietly inflated.
- How do you reduce unattributed registrations?
- Fix the cheapest cause first, which is usually campaigns that were never tagged at all, since it costs process rather than engineering. Then repair the hop that loses the most tags in transit. Genuinely unprompted arrivals and offline entries are not defects, so set the target against the first two causes rather than against the whole bucket.
Related reading
- Direct traffic registrations are four different problems wearing one label
- Build a channel grouping taxonomy before you argue about attribution models
- Adding self reported attribution to a registration form without wrecking conversion