Skip to content

Attribution coverage rate is the number to fix before you trust a model

Acquisition and attributionUpdated 2026-08-188 min read

In short

Attribution coverage rate is the share of registrations that carry a usable acquisition source, counted as registrations with a source divided by every registration in the file. Measure it weekly against the show calendar, publish it next to any channel chart, and treat a model built on low coverage as an estimate with a stated gap.

The channel chart goes up in the Monday review eleven weeks out. Paid social 24 per cent, email 31, organic search 18, partner referrals 12, and a slice called other. Someone asks how many registrations the chart is drawn on. Nobody in the room knows, and when a person goes and counts that afternoon, the chart turns out to describe about half the show.

Attribution coverage rate is that answer written as a fraction. Registrations carrying a usable acquisition source, divided by every registration in the file. It is arithmetic anybody can do in ten minutes, and until it is on the slide next to the chart, every percentage on the chart is a share of an unstated population.

Wang and Strong made the general version of this argument in the Journal of Management Information Systems in 1996, in a paper built from what data consumers themselves said quality meant. Their framework sorts quality into four groups, intrinsic, contextual, representational and accessibility, and the contextual group is the one that bites here. Data can be perfectly accurate and still unfit for the task, because the task is a budget decision about the whole audience and the data describes part of it.

What counts as a usable source?

Take a real edition file of 18,600 registrations. 11,900 of them have something in the source column, so the naive coverage rate is 11,900 divided by 18,600, which is 63.98 per cent. Call it 64.

Now look at what those 11,900 values actually are. 2,300 of them say direct. 640 say other, unknown, or a single space that survived a form migration. Those rows record the absence of a source rather than a source, and counting them as covered defeats the purpose of counting at all. Take them out and 8,960 registrations carry a value you would act on. Coverage against the full file is 8,960 over 18,600, which is 48.17 per cent.

That is the number worth arguing about. The channel chart in the Monday review was drawn on 8,960 rows and presented as the show.

Keep both figures, because they answer different questions. The first tells you whether the capture mechanism fired at all, which is an engineering question and belongs with the people who own the handoff between the ad click and the registration record. The second tells you how much of the audience your channel reporting can speak for.

A third figure is worth having for the same file: how many rows carry a source that is specific enough to price. A row tagged paid social is better than nothing and worse than a row that names the campaign, the ad set and the placement. In the 8,960 there were 1,240 rows whose source was the bare channel with no campaign detail, which means 7,720 rows, 41.5 per cent of the file, can support a decision at the level a media buyer actually works.

Should comped registrations sit in the denominator?

This is where most teams quietly move the goalposts, so decide it once and write it down.

The argument for a smaller denominator is that some registrations cannot have a marketing source by construction. Exhibitor guest codes come from the exhibitor's own list. Staff badges, speakers and contractors were never acquired. Including them makes the coverage rate look worse than the capture system deserves.

The argument for the full denominator is that the chart is going to be read as a description of the show, and the show includes those people.

Publish both, with the same file and the same cut-off. Of the 18,600, 3,400 arrived on exhibitor guest codes and 260 were staff, speakers or contractors. The eligible denominator is 14,940, and 8,960 over 14,940 is 59.97 per cent. So the honest pair is 48 per cent of the file and 60 per cent of the registrations your own programme could plausibly have sourced. The gap between those two numbers is a fact about your guest code programme, and it is usually the largest single acquisition channel nobody has a cost for.

What you cannot do is compute the numerator on one population and the denominator on another. If guest code registrations are excluded from the denominator, any guest code rows sitting in the numerator have to come out too, or the rate goes above what the data supports.

Measuring it weekly against the show calendar

Coverage is not a constant, and the single post-show figure hides the part you can act on.

Run the count every Monday from campaign launch and plot it against weeks to show open. A typical shape for a show that sells for nine months: coverage holds in the low seventies from launch until about four weeks out, then falls hard. In this file the last fortnight brought 4,100 registrations at 38 per cent coverage, against 71 per cent for everything before it.

The late collapse has causes you can name. Onsite and walk-up registrations skip the tagged journey entirely. Deadline traffic arrives through forwarded links, printed QR codes and word of mouth. Exhibitor guest codes are pushed hardest in the final fortnight because that is when exhibitors panic about their own attendance.

That pattern matters because the final fortnight is when the channel chart is most often used to justify next year's plan. A report built at close mixes a well-measured nine months with a badly measured two weeks and presents the average as one number.

Weekly measurement also turns a tagging failure into an incident rather than a post mortem. A broken redirect that strips query strings shows up as a coverage drop the following Monday. Found at week minus eleven it costs you a morning. Found in the post-show pack it costs you the edition.

How high does coverage have to be before a model is worth running?

Coverage sets a bound on how wrong any channel share can be, and the bound is easy to compute, which makes it the most useful thing about the metric.

Work it on the eligible population of 14,940 with 8,960 covered rows. Paid social holds 22 per cent of the covered rows, which is 1,971 registrations. The 5,980 uncovered rows contain some unknown number of paid social registrations, between none and all of them. So paid social's true share of the eligible population sits somewhere between 1,971 over 14,940, which is 13.2 per cent, and 7,951 over 14,940, which is 53.2 per cent.

A forty point interval is not a finding. It is a shrug with a decimal point.

Push coverage to 85 per cent of the eligible population, so 12,699 covered rows, and hold the channel at 22 per cent of them, 2,794 registrations. The interval becomes 18.7 per cent to 33.7 per cent. Better, and still wide enough that two channels fifteen points apart on the chart could be the same size in reality.

Those are worst-case bounds and they assume missingness is adversarial. It is not, but it is also not random, which is the assumption the narrow answer needs. The rows without a source are disproportionately mobile, late, forwarded and guest coded, so they carry a different channel mix from the rows that have one. Anderl and colleagues built their graph based attribution framework in the International Journal of Research in Marketing in 2016 on four large customer level data sets, each with at least seven distinct online channels, and the model can only ever speak about channels the data recorded. Everything else is outside the model rather than at zero inside it.

My own rule is 80 per cent of the eligible population before an attribution model output goes anywhere near a budget, and the coverage figure printed on the same page as the model output every time. Below that, publish channel counts as counts, say what they are counts of, and spend the effort on capture instead. Somebody has to decide what to do with the uncovered rows in the meantime, and the options for the unattributed pile are a separate question with better and worse answers.

Where this stops

Coverage measures presence. A file can be 100 per cent covered and comprehensively wrong, because a value in the source column says nothing about whether it is the right value.

Every row carrying last touch as the only source is fully covered and systematically credits the closing message rather than the channel that found the person. Self-reported source questions produce a value for nearly every row and a known bias towards whatever sits at the top of the list. A default source applied at the registration platform level, which several platforms do when a parameter is missing, produces perfect coverage and no information whatsoever. Check for that one first, because a suspiciously round coverage rate near 100 per cent is usually a default rather than an achievement.

The metric also says nothing about the size of the gap between what platforms report and what the file holds. Two platforms can both claim a covered row, which inflates nothing in the coverage rate and everything in the channel chart, and walking a platform number down to the rows in the file is its own worksheet.

Coverage is a floor, then, and floors are worth having. It is the one attribution number that requires no model, no assumption about credit and no argument about windows, which is why it is the one to publish first and the one to defend when somebody wants to skip straight to the attribution work the pillar page covers.

Take last edition's registration file this week and produce three counts: rows with any non-empty source, rows with a source you would act on, and total rows. Put the two ratios on the same slide as the channel chart you already publish, and keep them there.

Questions people ask about attribution coverage rate

How do you calculate attribution coverage rate?
Count the registrations in the file that carry a source value you would act on, then divide by the total number of registrations in the same file for the same edition. Both counts come from the registration file rather than an advertising platform, and both use the same cut-off date, so the ratio describes one population at one moment.
Does a direct or unknown source count as covered?
No. A row labelled direct or unknown records that no source was captured, so counting it as covered hides the gap you are trying to size. Keep two figures: the share of rows with any non-empty source, and the share with a source specific enough to inform a budget decision. The second one is the working number.
What coverage rate do you need before an attribution model is worth running?
There is no published threshold, so set one and state it. Below roughly 80 per cent the bounds on any channel share are wider than the differences most budget arguments turn on. Above that, the model output is still an estimate, and the honest presentation carries the coverage figure next to it every time.

Related reading

All attribution articles