Truncated y axis charts and when cutting the baseline misleads a reader
A truncated y axis is one whose baseline sits above zero, which exaggerates change because readers judge a bar by its length. Anchor bars at zero. Truncation is defensible on line charts where slope carries the meaning and the quantity has no natural zero, provided the range is stated in the axis label.
Two bars on a slide, last edition and this edition, verified attendance. The second bar came up to about a quarter the height of the first. The show director's face changed before anyone had spoken, and the meeting spent eleven minutes on a recovery plan.
Attendance had fallen by 2 per cent. The chart had a truncated y axis, the baseline sitting at 14,000 because the tool had auto-scaled to the data, and the exaggeration was entirely an artefact of that one setting. Nobody had done anything dishonest. Somebody had accepted a default.
How much does cutting the baseline actually change?
Work it with the real numbers. Prior edition verified attendance 14,382. This edition 14,094. The fall is 288 people, which is 2.0 per cent.
The tool auto-scaled the axis from 14,000 to 14,500. Bar one rises 382 units above the baseline, bar two rises 94. Map that 500 unit range onto 300 pixels of plot height and bar one is 229 pixels tall while bar two is 56. The second bar is 24.6 per cent of the first, so the drawn fall is 75 per cent for a real fall of 2 per cent.
The exaggeration factor is the interesting part, because it is a free parameter. Move the baseline and watch it move.
At a baseline of 14,000, the drawn fall is 75 per cent. At 13,900 it is 60 per cent. At 13,500 it is 33 per cent. At 13,000, 21 per cent. At zero it is 2 per cent, which is the truth. Whoever sets the baseline is choosing, from a continuous range, how large the change should look, and in most reporting tools that person is an auto-scaling algorithm nobody has met.
Put another way, the chart's sensitivity is set by how much of the quantity it hides. With the baseline at 14,000, the visible range is 500 people across 300 pixels, so one pixel is 1.67 attendees and the 288 person fall is worth 173 pixels. With the baseline at zero, the visible range is 14,500 across the same 300 pixels, one pixel is 48 attendees, and the same fall is worth 6 pixels. The data did not change. The magnification did, by a factor of 29.
That is also why the default is so hard to notice. An auto-scaling algorithm is optimising for the data filling the frame, which is a sensible thing to want on an exploratory chart where you are hunting for movement, and a terrible thing to want on a chart whose job is to convey how big a movement was. The same setting is right in one context and wrong in the other, and reporting tools do not ask which one you are in.
Pandey, Rall, Satterthwaite, Nov and Bertini put numbers on the consequence at CHI in 2015, in a paper on common distortion techniques. They tested four ways a chart can mislead, truncated axis among them, and measured how much each shifted a reader's reading of the message against an undistorted control. Truncation belonged in the category they described as exaggerating or understating the message rather than reversing it, which matches what happened in that meeting: nobody concluded attendance had risen, everyone concluded the fall was severe.
Does an axis break marker fix it?
This is where the received advice fails, and it fails in a way that has been measured.
The convention is that truncation is acceptable if it is signposted, usually with a zigzag glyph through the axis or a visible gap in the bars. Correll, Bertini and Franconeri tested precisely that at CHI in 2020, running crowdsourced experiments on how truncation changes the subjective size of an effect, and explored alternative designs meant to alert the reader. Their finding, in their own words: "We find that the subjective impact of axis truncation is persistent across visualizations designs, even for designs with explicit visual cues that indicate truncation has taken place."
The break glyph documents your choice. It does not undo the impression the chart makes. A reader can know, in the sense of having read the label, that the axis starts at 14,000, and still walk out of the room with a picture of a collapse.
I would therefore stop treating the break marker as a licence. Put it on where you have truncated, because a reader checking your work deserves to find it, and do not let its presence be the reason you truncated.
Where truncation is the honest choice
Bars carry their value as length from the baseline, so cutting the baseline corrupts the encoding. Lines carry their value as position, and the reader's job is usually to judge direction and rate. That difference decides the rule.
Registration pacing makes the case. An index running across 26 weeks moves between 88 and 104, sixteen index points of genuine, actionable variation. On a zero-based axis from 0 to 110 rendered across 300 pixels, the entire range of variation occupies 44 pixels, and the line looks flat. On an axis from 85 to 105 the same variation occupies 240 pixels and you can see the week the paid campaign landed.
The zero-based version buys nothing in truthfulness here and loses the legibility of the thing being measured, because it implies that the distance from zero means something. For an index it does not. The same holds for any quantity with no natural zero, such as a satisfaction score or a year on year ratio.
The choice of encoding comes first, though. If the comparison a reader needs is between categories at one moment, the accuracy ordering settles the chart type before the axis question arises, and V23 works through that. Axis decisions are what remain once the encoding is right.
So the rule I use has two lines. Bars start at zero, without exception, and if the resulting chart is boring then the change is boring and the chart is telling the truth. Lines may start wherever the data lives, with the range written into the axis label rather than left to the reader to infer from tick marks.
There is a third case worth naming. A line chart whose vertical extent is a rounding error, drawn tall and narrow, exaggerates slope just as effectively as a cut baseline does, and no truncation is involved. Aspect ratio is the same class of decision and it is even less often examined.
Truncation gets more dangerous once a forecast is on the same chart, because a cut baseline widens the drawn interval in the same proportion it steepens the line, and an 80 per cent band that already gets misread as a hard floor now looks alarming as well. How to shade and label an interval so a reader does not treat its edge as a limit is V33's territory. Settle the baseline before you draw the band, because the band inherits whatever the axis does.
What to do with a 2 per cent change
The deeper problem in that meeting was that a two bar chart was drawn at all.
Two numbers 2 per cent apart carry almost no information about whether anything happened. The useful question is whether 2 per cent is unusual for this show, and the only thing that answers it is history. Take twelve editions of verified attendance, draw them as one line, mark the current point, and the reader can see the normal edition to edition wobble. If the last decade has moved by 1 to 4 per cent in either direction most years, this year's 288 is weather. If the series has been flat within half a per cent and then dropped 2, that is a finding.
Correll and colleagues land in a similar place. Their closing recommendation, verbatim, is that "designers consider the scale of the meaningful effect sizes and variation they intend to communicate, regardless of the visual encoding". Deciding what size of change matters, before choosing an axis, removes most of these arguments. If 2 per cent is inside your normal band, no axis setting will make the chart useful, and the honest output is a sentence.
That principle carries into panelled portfolio views too, where every panel has to share a scale before any comparison between shows means anything. Matching the scales across eight shows is V25's problem, and it is the same decision applied eight times at once.
Where this stops
None of this makes the chart correct. It makes the axis honest, which is a smaller claim.
A zero-based bar chart of verified attendance built on badge scan coverage of 60 per cent will faithfully render a number that is wrong by 40 per cent, and it will do so with an axis nobody can object to. Getting the encoding right protects a reader from a distortion introduced in the drawing. It does nothing about a distortion introduced in the measurement, and a well-drawn chart is more persuasive, which makes a measurement error more expensive rather than less.
The other limit is that people will keep truncating, because a truncated chart looks like it contains a finding and finding-shaped charts get into board packs. That is a review problem rather than a design problem. The place to catch it is whoever signs off the pack, and the check takes two seconds per chart. Making that check routine across every page in the BI and reporting set is cheaper than arguing about one slide a quarter.
Go through last quarter's board pack and read the bottom of every vertical axis. Every bar chart whose axis starts above zero gets redrawn from zero before it goes out again, and if the redraw makes the chart look like nothing happened, that is the result you were meant to have.
Questions people ask about truncated y axis
- Is it ever acceptable to truncate the y axis?
- On a line chart, often. Slope is the message, and a zero baseline can compress a real movement into a flat line. On a bar chart, almost never, because a bar encodes its value as length from the baseline and cutting the baseline changes that length arbitrarily. Index values and rates with no natural zero are the clearest case for truncation.
- Does adding a break symbol to the axis make truncation honest?
- Not reliably. Correll, Bertini and Franconeri tested designs carrying explicit cues that truncation had occurred and reported at CHI in 2020 that the subjective impact persisted anyway. A break glyph documents the choice for a careful reader and does very little for the impression the chart leaves on everyone else.
- How do you show a small percentage change without exaggerating it?
- Put the number in a sentence and show the series over enough history for the change to be judged against normal variation. A single pair of bars two per cent apart carries almost no information either way. Twelve editions on one line, with the current point marked, lets a reader see whether two per cent is unusual.