Data Intelligence for AWS

Understand where every data point comes from⁠—and what breaks when it changes.

TrustedLineage discovers your AWS resources in read-only mode and builds explainable lineage, with evidence and confidence from process to column.

  • Read-only discovery
  • AWS-first
  • AI-assisted
  • Human validation
Data lineage diagram from AWS sources to destination tables, with evidence and confidence.
Product visualisation with demonstration data.
  • AWS-first
  • Read-only discovery
  • Column lineage
  • AI-assisted analysis
  • Human validation
  • Impact analysis

The problem

Your data platform changes faster than its documentation.

Manual lineage goes stale the moment someone deploys a new job. And when the knowledge lives in a handful of people, every change turns into a conversation instead of a query.

Answer

The graph walks the lineage upstream to the source: which process wrote the table, which S3 path or table it read from, and what evidence supports each hop.

What solves it Explore · lineage graph

Architectures that are hard to read

Dozens of jobs, functions and buckets with no single view relating them.

Invisible dependencies

Tables and datasets shared between processes that nobody has inventoried.

Out-of-date documentation

Diagrams that were right the day they were drawn and obsolete by the next deployment.

Changes with unknown impact

Changing a schema means asking on Slack instead of querying the lineage.

How it works

From AWS resources to lineage you can trust

Five stages. You decide what gets analysed and what gets published.

01 / 05

Discover

A read-only connection to your AWS environment to discover resources and spot candidate processes.

The product

From discovery to trusted lineage.

One continuous flow: scan, extraction, validation and exploration over the same knowledge base. These are the product screens, not a concept.

Three walkthroughs to start with. All nine areas remain available.

An inventory of your AWS ecosystem without modifying your resources.

  • AWS resource discovery
  • Inventory by account and region
  • Candidate process detection
AI Data Lineage Mapper
DataLake Intelligence Copilot

Dashboard

Resumen de cobertura, calidad y estado del linaje

Cuenta: acme-datalake Región: eu-south-2 Entorno: DEV Último escaneo: hoy, 16:34 COMPLETADO

Recursos descubiertos

259

▲ +125 vs. escaneo previo

Procesos detectados

61

▼ -21 vs. escaneo previo

Procesos con linaje

47

de 61 totales

Cobertura de linaje

77%

▲ +14 p.p. vs. previo

Confianza media

89%

confianza del proceso

Servicios detectados

6

▲ +3 vs. escaneo previo

Estado del linaje

61
Procesos totales
  • Con linaje validado 47 (77 %)
  • Con linaje no validado 9 (15 %)
  • Extracción en curso 2 (3 %)
  • Error de extracción 1 (2 %)
  • Sin linaje extraído 2 (3 %)
Ver todos los procesos →

Calidad del linaje

  • Relaciones totales 170
  • Confirmadas 142 (84 %)
  • Pendientes 24
  • Rechazadas 4
  • Baja confianza (<85 %) 18

Cobertura por tipo de recurso (nº de relaciones)

  • Tablas159
  • Columnas43
  • Procesos9
  • Almacenes14
Ver calidad de linaje →

Tendencia de cobertura

Últimos 7 escaneos

Requiere atención

4 alertas
  • Procesos con error de extracción 1
  • Procesos sin linaje extraído 2
  • Relaciones pendientes de revisar 24
  • Relaciones con baja confianza 18

Últimos escaneos

hoy, 16:34 COMPLETADO

111122223333 / eu-south-2 / DEV

Procesos

61

Con linaje

47

Confianza

89 %

Ver detalle →

Cambios recientes

  • Relación validada hace 2 h
  • Relación validada hace 5 h
  • Relación rechazada ayer
Ver historial →

✦ Copilot (Linaje)

Ver historial

Haz una pregunta sobre el linaje…

✦ ¿Qué procesos tienen baja confianza?

✦ ¿Qué procesos no tienen linaje extraído?

✦ ¿Qué cambió en el último escaneo?

✦ ¿Cuáles son los activos más críticos?

Recursos destacados

Ver todos →
  • ready.customer_360Tabla
  • glue-orders-curatedGlue Job
  • s3://acme-raw/orders/Ruta S3

Mapa resumido de la plataforma

Ver mapa completo
  • Fuentes externas4
  • Amazon S314
  • AWS Glue Jobs28
  • AWS Athena0
  • Amazon Redshift96

AWS Glue Jobs → Redshift: 156 relaciones · Amazon S3 → Redshift: 12 relaciones. Ruta directa sin pasar por Athena.

Every alert is a link: it leads to the list of processes or relationships that produce it, not to a report someone has to interpret.

Product views over a demonstration environment. The account, the processes and the table and dataset names are fictitious.

Column-level lineage

Down to the column, when the evidence is there

When the code and the SQL allow it, the analysis goes down to column level and identifies the type of transformation. When the evidence is not enough, the relationship is flagged as such instead of being taken as sound.

  • direct
  • rename
  • cast
  • calculation
  • join
  • aggregation
  • + other detectable patterns

Four relationships detected between the two tables. Pick one to see its type.

curated.crm_orders

ready.customer_360

direct order_id → order_id

The column is copied unchanged. It is the easiest relationship to support: the SELECT names it the same in source and target.

Use cases

What it is used for inside a data team

What it lets you answer

What breaks if I touch this table or this column?

Know what might break before changing a table, a column or a job.

Do any of these scenarios look like your platform?

Request a demo

Security & AI

Designed with enterprise security in mind

  • Read-only discovery

    The connection to your AWS account is read-only: it discovers resources and processes without modifying anything, and only what you select is analysed.

  • Your table contents are not analysed

    Rows, samples and column contents never enter the analysis or leave your account. What is analysed is how data moves, not the data itself.

  • The model runs where you decide

    Ollama locally, with nothing leaving your network, or Claude or OpenAI under an egress policy declared class of data by class of data.

We would rather be explicit: today we claim no certifications or compliance seals, and access uses the product's own users, not your corporate SSO yet. The specific security, deployment and data processing requirements are reviewed case by case during the evaluation.

View technical details Connection and access, what the AI does, providers and the full data egress matrix.

Connection and access to your AWS account

Analysis only on selection

Only the processes you explicitly select are analysed.

Control of the scope

You decide which accounts, regions, environments and processes go into each scan.

You create a read-only role in your account with the trust policy the product gives you, protected with an ExternalId. No credentials of yours are stored: the role is assumed for each operation.

In both cases

Verified permissions

Every connection is validated and shows what it can and cannot do: discovery, Glue Jobs, catalogue, Lambda, Step Functions or reading code from S3. If a permission is missing you see it before scanning instead of failing halfway through.

Adoption model

Start with one domain. Extend it when it pays off.

The product is in a validation phase with its first organisations. That means direct access to the person building it, and priorities that genuinely make the roadmap.

01

Pilot

We analyse one domain or a limited set of pipelines to prove the value against your own lineage.

02

Scale

Progressive extension to more accounts, regions, environments and processes.

03

Customize

Tuning of analysis rules, integrations and capabilities to particular needs.

A platform that adapts to your architecture We do not impose a single architecture. The product core is the same for everyone; the analysis rules, integrations and deployment adapt to how you work.

What you get from the first scan

  • Read-only AWS discovery
  • AI-assisted lineage extraction
  • Confidence, evidence and review
  • Graph, process detail and Copilot
  • Users, roles and per-account scope

What we adapt to your environment

  • AI model provider
  • Assume Role or appliance
  • Confidence target
  • Analysis depth
  • AWS services
  • Accounts, regions and environments
  • Naming conventions
  • ETL / ELT patterns
  • Governance requirements
  • Roles and scope per connection
  • Data egress policy
  • Integrations

Lineage is persisted in PostgreSQL as the source of truth and, optionally, the graph is synchronised in Neo4j. It can be adopted as a product, as an accelerator for a governance project, or as a platform configured for one specific environment.

Tell us about your architecture and we will assess the fit.

Evaluate a pilot

Roadmap

Data lineage is only the beginning.

The same technical understanding that makes the lineage possible enables other services on top of the platform. We draw a clear line between what exists today and what is on the roadmap.

  • AWS resource discovery (read-only)
  • Process lineage
  • Tables, datasets and S3 paths
  • Detected transformations
  • Column lineage where there is evidence
  • Confidence and evidence per relationship
View all 15 available capabilities
  • Review and publishing
  • AI model running locally or on an external API
  • Catalogue and inventory with domain and owner
  • Per-resource documentation in Markdown
  • Graph exploration
  • AI Data Copilot
  • Users, three roles and scope per AWS account
  • Declared data egress policy
  • Multi-account connection with verified ExternalId

Outside of that: access uses the product's own users. Integration with a corporate directory —SSO, SAML or OIDC— is planned, and the authorisation layer is already built so as not to depend on how you authenticate, but it is not there today.

Capabilities marked as roadmap are not available yet and their scope may change. If one of them is a priority for your team, tell us.

Frequently asked questions

Data lineage on AWS, explained

What is data lineage?

Data lineage is the record of how information moves and is transformed across systems: which source a piece of data comes from, which process reads it, what transformations it goes through and which table or dataset it ends up in. It lets you answer impact and data governance questions without depending on what one particular person happens to remember.

How do you document data lineage on an AWS platform?

Done by hand, it means diagrams and spreadsheets that go stale the moment someone deploys a new job. AI Data Lineage Mapper takes it from the environment itself: it discovers the account's resources in read-only mode and analyses the code, the SQL and the configuration of the processes you select. The processes analysed to extract lineage are Glue Jobs, Lambda functions and Step Functions, and the resulting relationships connect Glue Data Catalog tables, Amazon S3 paths and Redshift objects.

What is column-level lineage?

It is lineage that goes down from the table level to the field level: which source column feeds each target column, and through which transformation (direct, rename, cast, calculation, join or aggregation). It is identified when the code and the SQL contain enough evidence; when they do not, the relationship is flagged as such instead of being taken as sound.

How do I find out which processes read from or write to a table?

Once the lineage is published you query it in the graph or ask the Copilot in plain language. For any table, dataset or S3 path you can see its producing and consuming processes, and walk the lineage upstream and downstream as many levels as you need.

How do I run an impact analysis before changing a table or a column?

Impact analysis walks the lineage downstream from the asset you want to change and lists the processes, tables and datasets that depend on it. That way you know what might break before you touch a schema, instead of asking on chat and waiting for someone to remember the answer.

Can lineage extracted with artificial intelligence be trusted?

Every relationship is proposed with a confidence level and the specific evidence it rests on: file and line of code, SQL statement or job parameter. It then goes through human review before being published, and can be confirmed, rejected or left pending. The AI proposes and a person decides, so there is no black box.

Does the AI analysis send my process code to an external service?

Only if you choose to. The model provider is configured for each extraction: you can use Ollama running on your own infrastructure, in which case neither the code nor the SQL leaves your network, or the Anthropic or OpenAI APIs if you prefer their capability. The decision is explicit and is made before the analysis is launched.

Does it need write permissions on my AWS account?

No. Discovery runs in read-only mode and does not modify any resource in your account. The connection can be made in two ways: a read-only role assumed across accounts, protected with an ExternalId and without storing any credentials of yours, or by deploying the product as an appliance inside your own account. On top of that, only the processes you explicitly select are analysed, and you decide which accounts, regions and environments go into each scan.

Does the product read the contents of my tables?

No. Lineage is inferred from the code, the SQL and the schema metadata —database, table and column names— never from the rows. Data contents, samples and column values are classified as a category that cannot leave towards any model, external or local, and they do not enter the analysis either. AWS tags are excluded too, because they tend to contain email addresses, cost centres or ticket numbers that lineage does not need. The product analyses how data moves, not the data.

How is it controlled who can view or modify the lineage?

With the product's own users and three cumulative roles. Read grants access to the dashboard, the inventory, the catalogue, the lineage and the Copilot without being able to change anything. Operate adds launching scans, extracting lineage and reviewing relationships. Administer adds managing AWS connections, AI providers and users. Beyond the role, each user is assigned the AWS connections they are allowed to work on, so a team can see the development account and not production. Authorisation is checked on the server on every request: hiding a button is a convenience, not the protection. Access does not integrate with a corporate directory over SSO today.

Contact

Let's talk about your data platform.

Every architecture is different. Tell us what you want to understand, document, optimise or govern, and we will assess how to adapt the platform to your environment.

  • A technical review of your case, no strings attached
  • A demo on the real product
  • A proposal for a scoped pilot

Add details Optional: role, need, message and architecture.