Skip to content

The types of badge scans at a show and what each one proves

Attendee analyticsUpdated 2026-08-188 min read

In short

A show produces four distinct types of badge scan: entry scans at a hall boundary, session scans at a room door, exhibitor lead scans at a stand, and access denies. Each proves a different fact and each comes from a different system, so they cannot be added together into one engagement total.

The slide said 41,300 scans. The show director asked whether that was good, and the room went quiet, because nobody could say what a scan was. Two people in that meeting thought it meant people through the door. One thought it meant exhibitor leads. The person who had built the extract knew it was all four types of badge scans added together, and had not expected anybody to read the number as attendance.

That confusion is cheap to prevent and expensive to unpick after the fact. It starts with a taxonomy, and the taxonomy has to be written down before the show, because after the show you are guessing from timestamps.

What are the four types of badge scan, and what does each one prove?

An entry scan happens when somebody presents a badge at a hall boundary and a reader accepts it. It proves that a credential was read at that reader at that second. It says nothing about who was holding the badge, since badges get handed back through fences, and it cannot tell a first arrival from the same person walking back in after lunch.

A session scan happens at a room door, usually on a handheld or a kiosk owned by the conference team. It proves a credential was read within a few metres of a room at a time that maps to a programme slot. Nothing in it says the person stayed to the end, or went in at all.

A lead scan happens when exhibitor staff choose to capture a badge on a rented handheld or their own phone. It proves that a member of stand staff pressed a button while a badge was in front of them. The exhibitor controls whether it happens at all, which makes it a measure of stand behaviour before it becomes a measure of visitor behaviour. That is the whole subject of what a lead retrieval data model has to hold, and it is why lead scans belong nowhere near an attendance figure.

A deny happens when a credential is presented and refused. It proves that somebody tried. Denies carry a reason code, they are the most operationally useful of the four during show hours, and almost nobody reads them until the debrief. The access control event stream is where that argument belongs.

Four columns that settle every later argument

The taxonomy is four columns on every scan row. They are cheap to add at ingest and nearly impossible to reconstruct a year later.

Source system. Registration, access control, lead retrieval, session capture. Record the function the system performs, since you will change vendors and you want the meaning to survive the change.

Reader class. Fixed portal, pedestal turnstile, handheld, kiosk, phone camera. Reader class predicts the failure mode. A fixed portal double-fires when somebody hesitates in the field. A handheld under-reads because a human has to decide to use it, and humans stop deciding at about four in the afternoon.

Device owner. Organiser, exhibitor, venue, contractor. Ownership decides who can be asked to fix a broken device on show morning, and it tells you whose incentive shaped the data. An exhibitor with a lead target scans differently from a steward on an hourly rate.

Subject class. Attendee, exhibitor staff, contractor, media, organiser staff. This column governs whether the row is allowed anywhere near a headline number, and it is the one most often missing, because exhibitor staff badges are frequently issued through a different workflow that never writes back to the attendee table.

Add a fifth if you have it: the credential technology printed on the badge. A QR symbol produced to ISO/IEC 18004, first published in 2000 and now in its 2024 edition, behaves differently at a door from a contactless chip, because a camera needs the badge held flat and still while a proximity reader does not care about orientation. Reader class and credential technology together explain most of the variation in read failures across your door estate.

Declaring the grain before anybody counts anything

Kimball and Ross set out the four-step dimensional design process in the third edition of The Data Warehouse Toolkit in 2013: select the business process, declare the grain, choose the dimensions, identify the facts. Step two is where scan data goes wrong, and their warning about it is blunt. Every measurement inside a fact table has to sit at the same level of detail.

They name three fundamental grains. Transaction grain, one row per event captured at a single instant. Periodic snapshot, one row summarising many events over a standard period. Accumulating snapshot, one row per process instance, revisited and updated as the process moves through its stages.

Scans belong at the transaction grain. One row per read event, carrying the reader, the credential, the timestamp and the four taxonomy columns. Resist every request to store a derived attendance flag on that row. The moment you do, two people will disagree about the rule behind the flag, and you will be maintaining two versions of the same fact table with no way to tell which one a given report used.

The derived layers sit above it. Attendances per person per day is a different grain, produced by a rule, and the rule deserves its own argument because reasonable people set the window differently. That is the deduplication question, and it starts from a clean transaction table or it starts from nothing.

There is a second reason to keep the raw table raw. Rules change. If you tighten the debounce window between editions, a derived table has to be rebuilt and the old figures restated, which is manageable when the raw reads are still there and impossible when somebody has been writing aggregates and discarding the source.

What the 41,300 actually contained

Here is the split behind the slide, from one three-day show with eight halls and a 900-seat conference programme.

Hall entry scans, 12,700. Session scans at conference room doors, 8,900. Exhibitor lead scans pulled from the lead retrieval vendor, 18,400. Access denies across every door, 1,300. Those four add to 41,300.

Now watch what each denominator does to the story. Read 41,300 as engagement and the exhibitor lead scans are 44.6 per cent of it, and those 18,400 rows came from roughly 480 exhibiting companies pressing buttons on their own devices. Read it as attendance and the figure is 3.25 times the hall entry count, while the hall entry count is itself not yet attendance, because the raw reads have not been deduplicated.

The session number carries a different trap. 8,900 session scans across a programme of 74 sessions averages 120 scans per session, which sounds healthy until you find that 31 of those sessions had no scanning at all, because their rooms had nobody on the door. The mean is being computed across a denominator that includes 31 sessions where the measurement did not exist. Over the 43 sessions that were actually staffed, the average is 207.

None of those three readings is bad arithmetic. All three come from adding rows that sit at different grains, drawn from different populations, produced by different people with different reasons to press or withhold a button.

Why does one badge identifier mean four different things?

The badge carries a single identifier and every reader in the building can read it. That is the design goal, and it is also the source of the confusion, because the identifier stays stable while the meaning of reading it does not.

Read at a hall portal by an organiser-owned reader, it means a boundary was crossed. Read at a room door by a conference steward, it means somebody was standing outside that room. Read on an exhibitor's phone at a stand, it means a salesperson wanted the contact details. Read at a door that refuses it, it means the credential and the door disagree about what the badge is entitled to.

The claim you can make from the read depends entirely on facts that live outside the scan row: who owned the device, where the device was standing, and what the device was for. Store those facts with the row and the four claims stay separate for as long as the data survives. Leave them out and every downstream query mixes them silently, which is worse than mixing them loudly.

There is a practical test for whether your taxonomy is real. Pick a scan row at random and ask what fact it establishes about the world. If the answer needs somebody to open a floor plan, phone the lead retrieval vendor, or remember which halls had turnstiles that year, then the taxonomy lives in people's heads and it will leave when they do.

Where this stops

A clean taxonomy tells you what each row means. Truth is a separate problem, and two failures survive perfect labelling.

The first is badge sharing. An entry scan proves a credential crossed a boundary, and the person carrying it might be the registrant, a colleague using a spare, or somebody who picked it up off a table in the coffee area. No column fixes this, and shows with high comp-badge volumes carry more of it than shows that charge at the door.

The second is the missing row. A handheld that ran flat at 14:00 produces nothing at all, and an absence of rows looks identical to a genuine zero. Labelling makes the data you have interpretable. Data that never arrived stays invisible, which is why device-level health telemetry deserves to sit beside the scan table as its own feed rather than as a column inside it.

Both are reasons to publish the scan mix alongside any total, and to treat a single headline scan figure as something that has to show its components. The wider attendee analytics picture rests on this one table being honest about what it holds.

Start this week by counting your last show's scan rows grouped by source system and device owner. If either column does not exist, that is the finding, and it takes one schema change before the next edition to fix it.

Questions people ask about types of badge scans

How many types of badge scan does a trade show actually generate?
Four, in most builds. Entry scans at hall doors, session scans at room doors, lead scans taken by exhibitors at stands, and denies recorded when a credential is refused. Some shows add a fifth for catering or cloakroom redemption. The count matters less than whether every row carries a label saying which type it is.
Can entry scans and lead scans be stored in the same table?
Yes, provided every row carries its source system, reader class and device owner, and provided nobody queries the table without filtering on them. Storing them together with no type column is what produces a total scan figure that means nothing. Label first, then decide whether one table or four is easier to query.
What does a badge scan prove about attendance?
An entry scan proves that a credential was presented at a specific reader at a specific time. It says nothing about who was holding the badge, and it cannot distinguish a first arrival from a return after lunch. Attendance is a derived figure that needs deduplication rules applied to the raw reads first.

Related reading

All on-site analytics articles