Skip to main content
TrustedLineage AI Data Lineage for AWS
  • Guides
SpanishES EnglishEN
Request a demo
Home › Guides › Lineage in AWS Glue

Technical guide

How to get the data lineage of your AWS Glue Jobs

The Glue Data Catalog knows which tables exist, but not who writes them or where their columns come from. This guide explains where that information actually lives, why it is so hard to extract, and what it takes before you can rely on the result.

Last updated: 16 August 2026

The catalog is not the lineage

This is the most common confusion. The Glue Data Catalog is an inventory: database, table and column names, types and locations. It tells you that ready.customer_360 exists and has twelve columns.

What it does not tell you is who writes it, which source table feeds each of those twelve columns, or what transformation was applied along the way. That information does not live in the catalog: it lives in the job's code.

Where the relationship really is

In a PySpark Glue Job, the lineage is spread across several sources that have to be read together:

  • The code, in the read and write calls: spark.read.table(...), df.write.saveAsTable(...), spark.read.parquet(path).
  • The embedded SQL, which usually holds the column-level detail: the SELECTs with aliases, the JOINs and the aggregations.
  • The job parameters, because very often the database or the bucket arrives as an argument (--SOURCE_DB) and the code never mentions the real name.
  • The environment variables and dependencies, which fill in the rest when the job imports shared utilities.

That is why simple static analysis fails so often: if you only look at the code, the target is a variable; if you only look at the configuration, you do not know what is done with it.

Why the manual approach does not hold up

The usual alternative is to document by hand: a diagram, a spreadsheet, a Confluence page. It works the day it is written and starts ageing on the next deployment. On a platform with dozens of jobs, the documentation and reality diverge within weeks, and from then on nobody trusts it — which is worse than not having it.

The symptom is recognisable: when someone asks what happens if a column changes, the answer comes over chat and depends on who is online.

What it takes to automate it

Extracting the lineage of a Glue Job reliably requires four things:

  1. Discovering the jobs that exist in the account, without depending on someone keeping a list up to date.
  2. Retrieving their artefacts: the script, the SQL, the configuration and the parameters.
  3. Interpreting them together to infer sources, targets and transformations. This is where a language model earns its place, because the pattern differs in every team.
  4. Attaching its evidence to every relationship: the file and the line, the SQL statement or the specific parameter it rests on.

That fourth point is what separates a usable tool from a black box. Without evidence, automatically generated lineage is a claim nobody can check, and in data governance that is no use.

How far you can get at column level

When the SQL is explicit, column lineage can be determined precisely, and the transformation classified as well: direct, rename, cast, calculation, join or aggregation.

When it is not —a SELECT *, a DataFrame built dynamically— the honest answer is that there is not enough evidence. The right thing to do is flag it as such rather than invent a plausible relationship: lineage with declared gaps is useful, lineage with invented relationships is dangerous.

How AI Data Lineage Mapper solves it

The product connects to your AWS account in read-only mode, through a role assumed across accounts and protected with an ExternalId, or deployed as an appliance inside your own environment. It discovers the resources, you choose which processes to analyse, and on those it reads code, SQL, configuration and dependencies.

Every proposed relationship arrives with a confidence level and its evidence, and goes through human review before being published: it can be confirmed, rejected or left pending. The model can run on your own infrastructure, so your process code never leaves your network.

From there the lineage is persisted and explorable: graph, upstream and downstream traversal, impact analysis and queries in plain language.

Do you have undocumented Glue Jobs?

Tell us how many processes you have and what you would like to understand first. We will review your case and prepare a demo on the real product.

Request a demo

TrustedLineage

AI-powered Data Intelligence for modern cloud platforms.

Product Guides Legal notice and privacy Contact

© 2026 Gregorio Torrealba. All rights reserved.

AWS and the Amazon Web Services service names are trademarks of Amazon.com, Inc. or its affiliates.