Chart selection for event metrics using what the eye reads accurately
Chart selection for event metrics should follow the accuracy ordering Cleveland and McGill established in 1984. Position along a common scale is read most accurately, then position on unaligned scales, then length and angle, with area and colour last. Use position for anything being compared and keep colour for identity.
The post-show deck had a bubble chart of exhibitor performance by product category. Twenty-four bubbles, one per category, sized by total leads, coloured by hall, floating in a rectangle. It looked considered. Somebody on the call asked which of two named categories had done better and the room spent forty seconds squinting before the analyst opened the underlying table.
Chart selection for event metrics is mostly settled by a question that has a measured answer, which is how accurately a person reads a quantity out of a given visual channel. Cleveland and McGill published the ordering in 1984 and it has held up under replication. The bubble chart lost the argument before anyone drew it.
What did Cleveland and McGill measure?
Their paper, "Graphical Perception: Theory, Experimentation, and Application to the Development of Graphical Methods", ran in the Journal of the American Statistical Association in 1984, volume 79, number 387, pages 531 to 554. They identified the elementary perceptual tasks a person performs when pulling a number out of a graph, then ordered those tasks by how accurately people do them.
The ordering, best first: position along a common scale, then positions along unaligned scales, then length, direction and angle together, then area, then volume and curvature, and finally shading and colour saturation.
Heer and Bostock reran the spatial encoding experiments on Mechanical Turk and reported the results at CHI in 2010, pages 203 to 212. Their abstract says they "replicate previous studies of spatial encoding and luminance contrast and compare our results", and they extended the work to rectangular area perception of the kind a treemap produces. The 1984 ordering survived the move to a crowdsourced platform, which matters because it means the finding is about eyes rather than about a lab in New Jersey.
Two practical rules fall straight out. Anything a reader has to compare goes on a shared axis. Anything a reader only has to identify can use colour.
The bubble chart problem, in pixels
Take the deck's chart and do the arithmetic. Twenty-four categories, the largest at 720 leads, drawn so the largest bubble is 64 pixels across.
If the tool sizes by area, which is the correct implementation, then 720 leads maps to an area of pi times 32 squared, about 3,217 square pixels, so each lead is worth 4.47 square pixels. The category ranked eleventh has 380 leads. Its area is 1,698 square pixels, giving a radius of 23.3 and a diameter of 46.5 pixels. The category ranked twelfth has 355 leads, an area of 1,586, a radius of 22.5 and a diameter of 44.9 pixels.
Those two categories differ by 25 leads, which is 7 per cent, and on the chart they differ by 1.6 pixels of diameter. On a laptop screen at arm's length that is invisible.
Now the compounding error. A number of charting defaults scale the radius rather than the area. Under that mistake, 720 leads gives a radius of 32 and 180 leads gives a radius of 8, so a category with a quarter of the leads is drawn with a sixteenth of the ink. A four to one difference reads as sixteen to one. Check which one your tool does before you argue about anything else, because the two implementations produce charts that disagree with each other by a square.
There is a packing cost as well. Twenty-four bubbles with diameters between 20 and 64 pixels need roughly 33,000 square pixels of ink and, once you leave enough gap that the outlines read as separate objects, something like three times that in canvas. On a 900 by 500 slide area the layout algorithm starts overlapping them, and an overlapping bubble is being judged on the visible crescent instead of the full circle. The chart is now encoding magnitude in a channel the reader cannot see all of.
Redrawing it as a dot plot
Same 24 categories, same data, one axis running from zero to 720 across 800 pixels, categories sorted descending, one dot per category with the label on the left.
The eleventh category at 380 leads sits at 422 pixels along the axis. The twelfth at 355 leads sits at 394. They are 28 pixels apart, against 1.6 pixels of diameter difference in the bubble chart. That is a little over seventeen times the separation, for the same two numbers, using an encoding the 1984 ordering puts at the top rather than fifth.
The dot plot also answers the question the meeting actually asked, which was which of two named categories did better, because sorting puts them in rank order and the labels are readable text rather than a legend lookup. Nobody needs the analyst to open the table.
The dot plot also extends to the comparison an exhibitions team asks for next, which is this edition against last. Put two dots on each category row, joined by a light connector, and the change becomes a length on a shared axis. The eleventh category moving from 355 to 380 shows as a 28 pixel segment pointing right. A bubble chart cannot express that at all without doubling the number of circles, at which point nobody can tell which pairs belong together.
A bar chart would do the same job. The reason to prefer dots for this particular case is that 24 bars anchored at zero spend most of their ink on the range between zero and 300, where no category sits, while dots put the ink where the variation is. That preference reverses the moment somebody wants to read a value as a proportion of the whole, which needs the zero anchor and therefore the bar. Where the baseline can honestly be cut and where it cannot is its own argument, and V24 has it.
Which encodings do event metrics need?
Most of what an organiser reports falls into four shapes, and each has a settled answer.
Comparing categories at one point in time. Exhibitor count by hall, leads by category, registrations by channel. Position on a common scale, so a dot plot or a bar chart, sorted by value.
One series over time. Registration pacing, daily scans, rebooking against the sales book. A line, with time on the horizontal axis and position carrying the value.
Several series over time. Eight shows in a portfolio, or six channels. This is where a single axis starts failing, and the answer is panels on a matched scale rather than eight lines in one frame. That is V25's subject and it has its own arithmetic.
Part of a whole. Registration mix by source, floor space by category. A stacked bar read left to right, or a single bar per period. A pie chart encodes the value as angle, which sits below length in the ordering, and it forces the reader to compare non-adjacent wedges. Two pies side by side to show change over time compounds both problems.
A distribution. Leads per exhibitor, dwell time per visitor, spend per booking. This is the shape most often reported as a single mean, and the mean is usually the least informative number available, because exhibitor lead counts are heavily skewed. A histogram or a strip plot of all 640 exhibitors puts position on a common scale and shows the long tail that the mean of 214 leads was hiding. If the audience will not accept a distribution, give them the median and the quartiles as three numbers in a sentence, which costs one line and carries most of the shape.
Colour is for identity
The last rung of the 1984 ordering is shading and colour saturation, which tells you what colour is bad at. It is a poor channel for magnitude and a good one for saying which thing is which.
That distinction cleans up a surprising number of event reports. A choropleth of registrations by country encodes magnitude as colour and is therefore the least accurate presentation available, though it earns its place when the geographic pattern is the finding and the exact values are in a table underneath. A category chart coloured by hall uses colour for identity, which is fine. A category chart with a red to green gradient carrying lead volume uses colour for magnitude on top of an axis that already carries it, which adds nothing and costs the reader a legend.
Choosing which specific hues to use is a separate question with its own constraints, because a portfolio report gets printed, forwarded and read by somebody with a colour vision deficiency. That belongs to V26 and the palette question, and it sits under the same BI and reporting work.
Where this stops
The accuracy ordering measures how well a reader extracts a value. It says nothing about whether the reader looks at the chart, and those two things pull in different directions more often than the literature admits.
A dot plot of 24 categories is accurate and slightly dull. A treemap of the same data is less accurate and communicates, in one glance and without any reading, that two categories hold most of the floor. If the finding is the concentration, the treemap may do a better job of getting the finding into somebody's head even though every individual value is harder to read. The honest position is that accuracy is the default and a deliberate trade against it needs a reason you can say out loud.
There is also a floor below which accuracy stops mattering. If two categories differ by 7 per cent and your lead attribution has a coverage gap of 15 per cent, the dot plot is faithfully rendering a difference that your data cannot support. Better encoding does not improve measurement.
Take the last post-show deck you sent out. For each chart, write down the one comparison the reader was meant to make and the visual channel carrying it. Every chart where the channel is area, angle or colour and the comparison is a magnitude is a candidate to redraw, and the redraw is usually ten minutes.
Questions people ask about chart selection for event metrics
- What is the most accurate chart for comparing event categories?
- A dot plot or a bar chart on a common axis, because both encode the value as position along a shared scale, which Cleveland and McGill ranked as the most accurately read encoding in 1984. Sort the categories by value rather than alphabetically, and the comparison a reader came for takes about a second.
- Why are bubble charts bad for comparing exhibitor categories?
- Because they encode magnitude as area, which sits near the bottom of the accuracy ordering, and because a small difference in value becomes a difference of one or two pixels in diameter. Some tools compound this by scaling the radius instead of the area, which squares the apparent ratio between two values.
- When is a bubble chart the right choice?
- When both position channels are already carrying variables you need and size is a rough third dimension nobody will read precisely. Plotting exhibitor lead volume against stand area with bubble size showing years exhibiting is a fair use, because the two things being compared carefully are on the axes.