Scan to pipeline attribution for exhibitors without pretending the show closed the deal
Scan to pipeline attribution joins badge scans to CRM contacts, then those contacts to opportunities created after the show. Both hops lose records, so publish the match rate against all leads beside the pipeline figure. Report the result as influenced pipeline with a floor attached, and resist extrapolating across the unmatched population.
An exhibitor's finance director sat in a renewal meeting last spring with a printed report showing 980 leads from the previous edition, and asked the only question that mattered to him. How much of our pipeline came from that?
The account manager had no answer, because scan to pipeline attribution was not something the report attempted. It stopped at the lead count. The exhibitor's own answer, produced by their sales operations analyst a week later, was a spreadsheet claiming 2.9 million in pipeline attributed to the show, built by filtering opportunities on a campaign field that somebody had been setting by hand since 2019.
Both numbers were bad in different ways. One was silent. The other was confident and unauditable.
The join is three tables and one honest fraction
Scan to pipeline attribution is a two hop join with a reported loss at each hop.
The first hop takes badge scans to CRM contacts. The second takes CRM contacts to opportunities. Neither hop is complete, and the entire credibility of the exercise rests on saying by how much.
The first hop is a record linkage problem. Email is the usual key, it is far from sufficient, and the fallback tiers that recover part of the loss are E16's subject. What matters here is that the hop produces a match rate, and that the match rate must appear next to every number derived downstream of it.
The second hop is a modelling decision dressed as a lookup. A CRM contact is connected to an opportunity through a contact role, and whether that connection exists depends entirely on whether somebody in the exhibitor's sales team filled it in.
What does a CRM actually model when it credits a campaign?
It is worth reading how a mainstream CRM does this, because most exhibitors are working inside one of them and inherit its assumptions without noticing.
Salesforce's documentation for Customizable Campaign Influence describes the mechanism plainly: influence models scan active campaigns to identify members who are also assigned a contact role on an open opportunity, and create an influence record from that relationship. Three things have to be true for a scan to show up in that model. The contact must exist. The contact must be a campaign member. The contact must be attached to the opportunity as a contact role.
Two details from the same documentation are worth holding onto. Campaign Influence considers every campaign member regardless of member status, so a contact who was scanned and never followed up counts identically to one who booked a demo. And once an opportunity's stage is closed, won or lost, influence records are no longer created, which means late-arriving scan data cannot attach itself to a deal that has already landed.
There is also an auto-association time frame in setup that limits when a member-contact relationship is treated as influential, and choosing that window against the exhibitor's real sales cycle is E15's subject.
The practical consequence for an organiser is that the second hop's completeness is a property of the exhibitor's sales hygiene. Nothing you do to your own data improves it. Contact roles are the single most commonly skipped field in B2B CRM practice, and an exhibitor whose reps attach one contact role per opportunity will systematically under-report show influence compared to one whose reps attach four.
Working the join on 980 leads
One exhibitor, 980 unique leads from a single edition.
Start with the key. Of the 980, an email address is present on 812. The remaining 168 were scanned from badges where the registration record had no email or the exhibitor's export dropped it. Those 168 cannot enter the join at all, and that is the first loss.
Of the 812 with an email, exact lowercased match to a CRM contact returns 612. Match rate against all leads is 612 divided by 980, which is 62.4 per cent. Match rate against leads with a usable key is 612 divided by 812, or 75.4 per cent. Publish the first, because the second flatters the method by excluding the records it failed on.
Second hop. Of the 612 matched contacts, 74 appear as a contact role on an opportunity created after show open and inside the agreed window. Opportunity-level conversion among matched contacts is 74 divided by 612, or 12.1 per cent.
Suppose those 74 opportunities carry a total open value of 2.96 million, which averages 40,000 each. That figure is the influenced pipeline, and it is the number to put on the report, with the two fractions beside it. Reporting closed won revenue against a scan needs its own written rule, which is E17's subject.
Now the tempting arithmetic, which I would compute and refuse to publish. If the 368 unmatched leads converted at the same 12.1 per cent, the show would have influenced roughly 118 opportunities and about 4.7 million. That extrapolation assumes the unmatched population behaves like the matched one, and there is good reason to think it does not: people who registered with a personal address skew junior, and people whose email was missing entirely skew towards rushed aisle scans. The honest treatment is to state the 2.96 million as a floor, note that 368 leads could not be evaluated, and leave the reader to draw their own conclusion about the size of what is hidden.
The match rate belongs on the same line as the number
A pipeline figure without its match rate is a claim about the world dressed as a measurement.
Put them together in one sentence on the report: influenced pipeline 2.96 million, from 74 opportunities, across 612 of 980 leads matched to CRM contacts, a 62 per cent match rate. Anyone reading that sentence knows immediately that roughly two fifths of the leads are invisible to the analysis, and nobody can later quote the 2.96 million as though it were the total.
CEIR's 2015 study on Exhibitor ROI and Performance Metric Practices identified exactly this gap: exhibitors struggle with closing the loop on what becomes of their leads and with tracing exhibition leads back to post-show behavioural change or sales conversion. A reported match rate is the smallest useful step towards closing it, because it turns an unanswerable question into a measurable one with a known coverage.
Who should actually run the join?
Here is where I would take a position that tends to start an argument.
The organiser should not be ingesting exhibitor CRM data. It is a data protection problem, a commercial trust problem, and an integration burden that scales with the number of exhibitors on your floor rather than with the number of shows you run. Six hundred stands means six hundred CRM connections, each with its own field mapping and its own owner who leaves.
What the organiser should ship is the key and the method. That means an export where every lead carries a stable person identifier, a normalised email, a normalised company domain, the scan timestamp, and the show day, plus a one page methodology note describing the join exactly as above, including which match tiers to run and what to report. The exhibitor runs it inside their own CRM, where the opportunity data already lives and where the data protection position is clean.
The organiser then asks for one number back: how many opportunities and how much pipeline, with the match rate the exhibitor obtained. Some exhibitors will give it to you and some will not. The ones who do become the only genuine evidence you have about what your floor produces, and it comes without your exhibitor analytics ever holding their pipeline data.
Where this stops
The whole construction measures association and calls it influence, and that is a weaker claim than most people reading the report will take from it.
A contact who was scanned at your show and later appears on an opportunity may have been in an active buying cycle before the exhibition opened. Your scan is then a footprint of a process already running, and attributing pipeline to it credits the show with something it observed. There is no way to separate those cases from the scan data alone, and the fix is a control group, which almost no exhibitor will build.
The second limit is that the match rate itself is biased. Exact email matching preferentially finds people who registered with a corporate address, and corporate addresses correlate with company size and seniority. So the 612 matched leads are systematically more senior than the 368 unmatched ones, which means the 12.1 per cent opportunity rate is measured on the favourable half of the file. That is another reason to treat the 2.96 million as a floor and to resist the extrapolation entirely.
Ask one friendly exhibitor for a count of opportunities where a contact role email appears in the lead export you sent them, and for the count of leads in that export that matched a contact at all. Two numbers, one afternoon of their time, and you will know whether your export is even usable as a join key before you build anything around it.
Questions people ask about scan to pipeline attribution
- How do you attribute pipeline to trade show badge scans?
- Join the lead export to CRM contacts on a normalised email, then join those contacts to opportunities created after show open and inside the published window. Count the opportunities where a scanned contact holds a contact role, sum their value, and print both the contact match rate and the opportunity conversion rate next to the total.
- Should an organiser hold exhibitor CRM data?
- No. Six hundred stands means six hundred CRM connections, each with its own field mapping and its own owner who leaves, and it puts the organiser in a controller relationship over records it never collected. Ship a clean key and a one page method instead, and ask the exhibitor for the two resulting numbers.
- Why is a 62 per cent match rate worth publishing?
- Because a pipeline number without it reads as though it covers every lead. Stating that 612 of 980 leads matched a CRM contact tells the reader immediately that roughly two fifths of the file is invisible to the analysis, and it stops anyone quoting the pipeline total later as if it were complete.
Related reading
- Choosing an exhibitor attribution window that matches the real sales cycle
- Matching booth scans to CRM records when email is the only usable key
- Reporting closed won revenue from scans without overclaiming what the show did