Technical guide
Impact analysis: what breaks if you change a table or a column
The question always arrives at the worst moment: a field has to be renamed, a type changed or a table retired, and nobody knows for certain what depends on it. This guide explains how to answer it with data instead of from memory.
Last updated: 16 August 2026
What impact analysis actually is
It is walking the lineage downstream from the asset you want to change and listing everything that depends on it: processes that read it, tables built from it, datasets and paths that would end up affected.
Its twin is the upstream walk, which answers the opposite question: where does this data come from, and who am I trusting when I use it. Both directions come out of the same graph.
Why asking on chat does not scale
The usual method is to post in a channel: "does anyone use
curated.crm_orders?". It has three problems.
- It depends on who is available and on them remembering.
- It covers only what that person knows, not what exists. The forgotten consumers are exactly the ones that break.
- It leaves no trace: next time you ask again from scratch.
On a small platform it works. Once there are dozens of processes and several people have joined and left the team, that tribal knowledge is precisely the risk you need to remove.
What the lineage needs in order to answer
Not all lineage is good enough for impact analysis. For the answer to be usable, four properties are needed:
- Known coverage. Knowing which processes have been analysed and which have not. A graph that looks complete but covers half the jobs produces false negatives, and those are the dangerous ones.
- Configurable depth. One level answers "who reads this table"; three or four answer "which report comes out blank on Monday".
- Column granularity when the change is a field. Renaming a column does not affect every consumer of the table, only the ones that use that field.
- Confidence and evidence per relationship, to tell what is confirmed apart from what is an unverified inference.
A concrete example
You want to change the type of orders.amount from decimal to integer. The
downstream walk shows three things:
- A Glue Job reads it and aggregates it into
customer_360.total, with anaggregationtransformation. - A Redshift view exposes it to the reports.
- A Lambda writes it to an S3 path that another team consumes.
Without the lineage you would have seen the first two, which are the obvious ones. The third, the one that crosses into another team, is the one that produces the Friday afternoon phone call.
From the one-off check to continuous governance
Impact is not just a query before a change. The same graph answers governance questions that usually go unanswered: which pipelines have no consumer at all and are probably redundant, which assets are shared across several domains, who owns a dataset that half the platform depends on.
For that, the lineage has to be published and alive, not sitting in a slide deck from six months ago. That is the difference between documenting and governing.
How AI Data Lineage Mapper solves it
The lineage is extracted from the code, the SQL and the configuration of your Glue Jobs, Lambda functions and Step Functions, with evidence and a confidence level per relationship, and is published only after human review.
On top of it, the explorer lets you filter by node type, direction, depth from one to five levels, review state and minimum confidence. And the Copilot answers questions in plain language — which processes write a table, or what would be affected by a change — drawing on the persisted lineage.