Skip to content

How to spot a breaking change in a data feed early

Unified dataUpdated 2026-08-188 min read

In short

A breaking change in a data feed is any change that alters the result of a query already running against it. Grade each change by blast radius: added columns nothing reads are informational, widened types are warnings, and a dropped or renamed column referenced downstream is a stop that holds the load.

Third alert of the month, 02:40, and the message says the registration feed schema changed. The person on call gets up, looks, and finds a new column called dietary_notes that no report has ever heard of. She goes back to bed. Two weeks later the same alert fires and gets muted from a phone at the kerb outside a hotel, and that one was a dropped column feeding the geographic breakdown.

A breaking change in a data feed is any change that alters what a query you already have returns. That definition does the work, because most changes to a file are not that, and treating all of them as equal is what trained the team to ignore the channel.

What counts as a breaking change in a data feed?

The test is downstream, in your own queries.

A column arrives. If nothing selects it, joins on it, filters on it or groups by it, the shape of your file moved and the answers did not. That is worth a log line. A column disappears, and if four reports reference it, four reports are now wrong, and one of them goes to exhibitors.

The awkward middle is types. A badge id that was an integer and now arrives as a string will still load, still display, and still fail to join against the integer version held in your warehouse, so the query runs, returns rows, and returns the wrong ones. A widened numeric type is usually harmless. A date arriving as a string in a different format is usually not. You cannot separate those two from the schema alone, which is why type changes belong in their own severity band with a person attached.

Sanderson, Freeman and Schmidt, in their 2025 book on data contracts, put the whole class of problem in terms of pipelines breaking and consumers losing trust, with monitoring and continuous integration among the components of the architecture they describe. The grading step below is what makes monitoring survivable, because monitoring that pages on everything gets switched off within a quarter.

Building the referenced column list once

The input you need is the set of columns anything downstream actually reads, and you can build it in an afternoon.

Dump the SQL behind every scheduled report, every warehouse view and every extract job. Collect the column identifiers. Intersect that set with the column list in your feed manifest. On one registration feed of 41 columns and a reporting estate of 168 queries, you might find 23 columns referenced and 18 that no query has ever touched, which is 56 per cent referenced and a surprise to everybody the first time it is counted.

Two traps in the parsing. A query using select star references every column, so any report built that way collapses the exercise and should be rewritten before you rely on the list. And identifiers get aliased through views, so a view selecting reg_country as country means the alert on reg_country has to know about the view. Building the list from the warehouse's own column level dependency graph, where you have one, is more reliable than a regular expression over query text.

Refresh the list monthly. It moves when reports are added, which is constantly in the eight weeks before a show, and a stale referenced list downgrades exactly the change you needed to catch.

The three severities and what each one does

Three bands is enough, and each one has a defined consequence written down before the alert fires.

  • Informational. A column was added, or a column was dropped that appears in no query. Log it, include it in a weekly digest, load the file. No interruption.
  • Warning. A type widened on a referenced column, or a new value appeared in a categorical column something filters on. Load the file, flag the load, and put it in front of a person the same working day.
  • Stop. A referenced column was dropped or renamed, a type changed on a join key, or the primary key stopped being unique. Hold the file. Nothing downstream runs on it until a named person clears it.

The mapping into tooling is direct. dbt Labs' documentation, in its 2026 revision, ships four built in generic tests, unique, not_null, accepted_values and relationships, and every test takes a severity config of warn or error, with error_if and warn_if available for custom failure thresholds. Your warning band is severity warn. Your stop band is severity error. The threshold configs are what let you say that three bad rows is a warning and eight thousand is an error, without writing two tests.

Who gets woken up, and for what?

Count a quarter and the case makes itself.

Say a portfolio feed produced 14 schema changes in three months: 9 additions of columns nothing reads, 3 type widenings on referenced columns, 2 dropped columns that four reports between them referenced. Paging on every schema change is 14 interruptions, 9 of which are dietary_notes at 02:40. Paging on the stop band alone is 2, and both of those genuinely needed somebody.

The 3 warnings sit in the middle and should not page at all. They go to a channel, they get looked at inside the working day, and if that does not happen the escalation books time in somebody's calendar.

There is one more variable, and it is the show calendar. The same dropped column costs almost nothing in a quiet month and costs a great deal on the Tuesday of show week, when badge printing is reading the file every fifteen minutes. Calendar aware severity is part of deciding what enforcement does at the boundary, which is where the quarantine rules get written in K29.

Write the alert for the person it wakes

An alert that grades correctly and reads badly still costs you. The 02:40 reader decides from a phone in under a minute, and the message should carry the decision's inputs in its first two lines: which column changed, what changed about it, which reports reference it, and what the pipeline has already done.

"Schema change detected on registration feed" forces the reader out of bed to go and look. "reg_country dropped, referenced by four reports including the exhibitor geographic breakdown, load held pending review" can be judged without getting up. Everything in the second version comes from work the grading already did. The referenced list knows the report names and the severity band knows the action taken. Formatting them into the template is an hour of work that pays back on the first night shift.

Include the timestamp of the last clean load as well, since the first question anybody asks is how much history is affected, and the alert can answer it before it is asked.

Replay last quarter through the bands

A grading policy nobody has tested gets its first test during an incident, which is the worst possible review meeting. Run the test in advance instead, on history you already hold.

Take the 14 changes from the quarter above and write down, for each, what the new bands would have done and what actually followed. The two dropped columns grade as stops, and those were the two that produced incidents, so policy and reality agree where it mattered most. The rows to study are the disagreements. If one of the three type widenings was a change that silently broke a join, that column was a join key all along, the stop band's join key clause should have caught it, and the replay has found the gap before the gap found you.

The replay also produces the sentence that sells the change to whoever owns the on call rota: fourteen interruptions a quarter become two, and the same two incidents still get caught.

Where the referenced list is wrong

The list is built from queries you can see, and the queries you cannot see are the ones that hurt.

An analyst has a saved query in a notebook. Finance has an Excel workbook with a live connection that predates you. A BI tool holds its own semantic layer, and the columns it references never appear in any SQL file in your repository. Every one of those is a real consumer, and every one of them is invisible to the parse. The practical mitigation is to treat query logs as the source rather than source code, since a warehouse query log holds what actually ran, including the notebook and the workbook.

The second gap is outbound. A column nothing internal reads may be in the file you send to a media partner, or in the export feeding an exhibitor portal, and dropping it breaks somebody else's report instead of yours. Outbound feeds need their own manifest and their own referenced list.

The third is honest uncertainty about intent. The list tells you a column is referenced. It cannot tell you the reference matters, and there are columns in every estate that appear in a query written for a board pack in 2023 that nobody has opened since. Reviewing the list once a year with the people who own the reports is the only fix I know, and it usually retires more queries than it retires columns.

Start by dumping the SQL for your scheduled reports and counting distinct column references against one feed's column list. The ratio tells you immediately how much of your alerting is noise: if 18 of 41 columns are unreferenced, then roughly two in five schema changes on that feed never needed to reach a person at all. Then wire the alert to check the list before it decides what to do, which needs the shape difference in the first place, from the hash comparison on arrival in K21 and the manifest that says what the columns should have been in K22. Grading is the layer that makes the rest of the unified data monitoring worth keeping switched on.

Questions people ask about breaking change in a data feed

What counts as a breaking change in a data feed?
A change that alters what an existing query returns. That definition is narrower than a change to the file, because most changes to a file touch columns nothing reads. Dropping or renaming a referenced column breaks a query outright, and changing the type of one can break a join while leaving the query technically valid.
How do you know which columns a change will break?
Parse the SQL behind your scheduled reports once and collect every column identifier that appears in it. That set is your referenced list. Intersect an incoming change with the list, and a change touching nothing on it can be logged rather than escalated, while a change touching a referenced column needs a person before the load.
Should every schema change page someone?
No, and paging on all of them is how a team learns to ignore the channel. Most changes on an event feed are additions that nothing reads. Reserve the interruption for changes to columns your reports reference, and let the rest accumulate in a log that somebody reviews weekly.

Related reading

All data quality articles