Finding the API history limits in your event systems before a five year backfill
An API history limit is the oldest record a vendor will return through its interface, which is often shorter than the retention advertised for the account. Find it by pulling the oldest edition of each show and recording what came back, before you promise a five year portfolio view.
Someone on the commercial side asks for five years of registration and engagement history across the portfolio, and the request sounds like a data engineering problem. Eight shows, five editions each, forty event instances, one warehouse. Budget it, schedule it, run it.
Then you check the API history limits on each system in the stack and the shape of the project changes, because two of the platforms will not give you five years and one of them will not tell you so in the documentation.
The question everyone asks, and the answer they get
The question usually goes to the account manager, phrased as "how much history do we have?" The answer comes back as the retention period on the contract, or as a screenshot of a dashboard showing 2021.
Both answers are about the interface. The API is a different surface with different rules, and the gap between them is where backfill projects go wrong. A platform can keep aggregated reporting for years while deleting the row-level records that fed it. Google is explicit about this for Analytics: the data retention setting for user and event data does not affect standard aggregated reports, so a report can keep rendering a period whose underlying event rows have expired (Google, 2026). The chart survives. The rows you wanted to load do not.
Ask instead what the earliest retrievable record is, through the API, for each object you need, and whether editions marked archived or closed are queryable at all. Those are three separate questions and vendors answer them differently.
What does an API history limit actually look like?
They come in a few recognisable shapes, and knowing which one you are dealing with changes the workaround.
A documented retention period. The cleanest case. Adobe publishes the numbers for Marketo Engage: activity and campaign membership data is stored for a rolling 25 months past the activity date, and high volume activity data is retained for a rolling 90 days by default, with the note that beyond these periods the data is no longer available through the interface (Adobe, 2026). You can plan against that. You can also see immediately that email opens and clicks from three years ago are gone whatever you do.
A limit nobody wrote down. Swapcard's public documentation for its content API covers rate limiting in detail, a cost based limit of 10,000 points per query against an allowance of 60,000 points per minute, plus query complexity ceilings of 20 levels of nesting and 2,000 tokens (Swapcard, 2026). It says nothing about how far back events remain queryable. Silence is not a promise of unlimited history. It means the limit, if there is one, is a property of the deployment rather than the contract, and you find it by asking for the oldest record and seeing what comes back.
An archive state. The most common shape in event platforms. Old editions still exist, still render in the interface, and drop out of the API's default collection because the event is closed. Sometimes a flag brings them back. Sometimes the records moved to a store the API does not address.
Working the calendar for one portfolio
The abstraction hides how quickly this bites, so put dates on it.
Take a show that runs every March. You are planning the backfill in August 2026 and you want the five most recent editions: March 2022, 2023, 2024, 2025 and 2026. The marketing automation platform holds 25 months of activity data on a rolling basis, so on the day you pull, the floor sits at July 2024. March 2025 and March 2026 are inside the window. March 2024 missed it by four months. Three of the five editions have no activity history to load, and no amount of retry logic changes that.
Now scale it to the portfolio. Eight shows, five editions each, forty instances. If the registration platform holds three years live, you get roughly 24 of the 40 instances from registration. If activity data holds 25 months, you get about 16. If the lead retrieval system was only introduced in 2024, you get 16 there too, and they are not the same 16, because the shows run at different times of year.
The honest output of that exercise is a grid: system down one axis, event instance across the other, and a yes or no in every cell. Forty instances against five systems is 200 cells. Filling it takes a couple of days of probing and it is the single most useful artefact the project produces, because it converts "we have five years of history" into a countable statement that survives contact with the board pack. Building anything on top of it, including a multi year registration baseline, depends on knowing which cells are real.
How do you test a limit the vendor has not documented?
Pull the oldest thing you can name and record exactly what came back. That is the whole method, and the discipline is in the recording.
Start with the oldest edition of one show. Request it by identifier rather than by date range, because a date filter on an empty collection returns the same empty result as a date filter on an archived one and you learn nothing. If the identifier returns a record, note the earliest created timestamp inside it. If it returns nothing, ask the vendor whether the edition is archived and what parameter unarchives it for a read.
Then repeat with a date range that you know brackets a real event, and compare. A vendor whose collection endpoint silently applies a default window will return fewer rows than the identifier lookup implies exist, and that difference is the limit you were looking for.
Three things are worth writing down for every system: the earliest retrievable record, the row count for one known edition, and the response when you ask for something older than the floor. That third one matters more than it sounds. An API that returns an empty list for out of range history behaves very differently in your loader than one that returns an error, and the empty list is the dangerous case, because a loader treats it as a successful pull of zero rows.
Probe with the cheapest query the API allows, because deep history and rate limits interact badly. On a cost based scheme like Swapcard's, a query that walks 20 levels of nesting to fetch one old registration burns points you will want later for the actual load, and the published ceiling of 10,000 points per query is easy to hit when you are guessing at the shape of an old record (Swapcard, 2026). Ask for two fields and an identifier while you are still finding the floor, then widen the selection once you know the edition responds at all.
One more habit worth building in at this stage: record the date you ran the probe next to the answer. A rolling retention window means the floor moves, and a note saying the earliest activity record was July 2024 is only meaningful if the reader knows you asked in August 2026.
What to do when the history is not there
A few options, in the order I would try them.
Ask for a file. Vendors who cannot serve five years through an API can often produce a one time extract from a backup, and this is a support ticket rather than an engineering project. It arrives as CSV, it has no schema guarantees, and it is still the cheapest path to depth.
Reconstruct from what you already hold. Post show reports, board packs, the registration extracts somebody saved to a shared drive every year. These are aggregates and they are lossy, and for a trend line at portfolio level they are frequently enough. Label them as reconstructed in the warehouse so nobody later joins them to row level data and wonders why the totals move.
Start the clock now. If the floor is 25 months, an extract taken every quarter from today means that in two years you have a history nobody can delete, held in your own storage. This is the option with the worst payoff this quarter and the best payoff every quarter after, and it is the one most often skipped because it does not solve today's request.
Getting the limit into the contract
The renewal conversation is where an undocumented history limit becomes a documented one, and the ask is narrow enough to be winnable: the earliest retrievable record by object, whether archived editions are queryable, and how much notice you get before a retention policy changes. What that clause should say, and the export format and frequency it needs to sit alongside, belongs with writing data export rights into the contract. Whether the answer you get is good enough to renew on is the subject of testing for vendor lock in.
Where this stops
Knowing the floor does not tell you the data below it is worth having, and there is a real argument that some of it is not.
Registration records from five years ago carry email addresses that have decayed, job titles that are wrong, and company names for firms that have been acquired twice. For counting, they are fine. For anything that treats a person as reachable, they are misleading, and a backfill that loads them without a recency flag creates a warehouse where a 2021 registrant and a 2026 registrant look equally current.
The other limit is that history depth is set by the shortest system in the chain, and averaging across the grid hides that. If registration reaches back five years and identity resolution depends on a field that only one of your systems has collected since 2024, your cross year matching effectively starts in 2024 whatever the registration table says. Work out which system is the binding constraint before you scope the integration work, because widening the wrong one costs the same and buys nothing.
Start this week by picking your oldest show edition and pulling it by identifier from every system in the stack. One afternoon, five requests, and you will know whether the five year number in the project plan is a fact or a hope.
Questions people ask about api history limits
- How far back can you pull data from an event platform API?
- It varies by vendor and by object, and the only reliable answer comes from testing. Adobe documents that Marketo Engage keeps activity and campaign membership data for a rolling 25 months, with high volume activity kept 90 days by default. Swapcard documents query cost limits with no stated history limit, which means you have to probe it yourself.
- Why does the interface show more history than the API returns?
- Interfaces often read from aggregated or archived stores that the API does not expose. Google states that its data retention setting does not affect standard aggregated reports in Analytics, so a chart can keep showing a period whose underlying event rows have already been deleted. Screenshots of an old dashboard are not evidence that the rows are retrievable.
- What should you test before committing to a five year backfill?
- Pull the single oldest edition of one show from every system in the stack and record the earliest record returned, the row count, and whether archived editions are queryable at all. Do this before the project plan is written. Each system will have a different floor, and the shortest one sets the real depth of your portfolio history.