01
Pilot
We analyse one domain or a limited set of pipelines to prove the value against your own lineage.
Data Intelligence for AWS
TrustedLineage discovers your AWS resources in read-only mode and builds explainable lineage, with evidence and confidence from process to column.
The problem
Manual lineage goes stale the moment someone deploys a new job. And when the knowledge lives in a handful of people, every change turns into a conversation instead of a query.
Answer
The graph walks the lineage upstream to the source: which process wrote the table, which S3 path or table it read from, and what evidence supports each hop.
What solves it Explore · lineage graph
Answer
The analysis reads the process code, SQL, configuration and parameters, and from that it determines which job or function writes each table.
What solves it Analyze · code and SQL analysis
Answer
Impact is traced downstream from the column: which processes read it, which tables it ends up in, and through what transformation. At column level where there is evidence to back it.
What solves it Explore · impact analysis
Answer
The same walk, the other way round: the processes that read the dataset and the destinations it reaches.
What solves it Explore · downstream lineage
Answer
The inventory collects discovered resources and their detected runs, so a process with no recent runs shows up for what it is.
What solves it Discover · inventory and runs
Answer
Every relationship carries its confidence level and the evidence it was inferred from. You confirm, reject or leave pending: the published lineage is the one you validated.
What solves it Validate · confidence, evidence and human validation
Dozens of jobs, functions and buckets with no single view relating them.
Tables and datasets shared between processes that nobody has inventoried.
Diagrams that were right the day they were drawn and obsolete by the next deployment.
Changing a schema means asking on Slack instead of querying the lineage.
How it works
Five stages. You decide what gets analysed and what gets published.
01 / 05
A read-only connection to your AWS environment to discover resources and spot candidate processes.
02 / 05
You pick the processes, accounts and regions you want analysed. Nothing is analysed without being selected first.
03 / 05
The AI analyses code, SQL, configuration, parameters and dependencies of the selected process.
04 / 05
You review the confidence and evidence of each relationship: confirm, reject or leave pending.
05 / 05
You publish the lineage and explore it with the graph, impact analysis, process detail and the Copilot.
The product
One continuous flow: scan, extraction, validation and exploration over the same knowledge base. These are the product screens, not a concept.
Three walkthroughs to start with. All nine areas remain available.
An inventory of your AWS ecosystem without modifying your resources.
Every relationship arrives with its confidence level and its evidence.
Navigate the ecosystem and ask questions in plain language.
Linaje de datos
Vista general de relaciones entre procesos, tablas, rutas S3 y datasets escaneados
Procesos con linaje
12
Datasets
42
Tablas
35
Rutas S3
7
Relaciones canónicas
47
Columnas mapeadas
78
Revisadas
11%
Baja confianza
38
Relaciones destacadas · 47
| Origen | Relación | Destino | Cons. | Prod. | Confianza | Estado |
|---|---|---|---|---|---|---|
curated.crm_orders | Lee | glue-orders-curated |
1 | 0 | 96% | Confirmada |
glue-orders-curated | Escribe | ready.customer_360 |
0 | 1 | 92% | Confirmada |
accounts.name | Transforma | customer_360.client |
1 | 1 | 91% | Inferida |
crm_orders.amount | Transforma | customer_360.total |
1 | 1 | 84% | Pendiente |
s3://acme-raw/legacy/ | Lee | curated.accounts |
1 | 0 | 61% | Rechazada |
Exportable to CSV and navigable with the keyboard: it is the accessible alternative to the graph. The five review states are Pending, Confirmed, Inferred, Rejected and All. Nothing is published until a person reviews it.
Dashboard
Resumen de cobertura, calidad y estado del linaje
Recursos descubiertos
259
▲ +125 vs. escaneo previo
Procesos detectados
61
▼ -21 vs. escaneo previo
Procesos con linaje
47
de 61 totales
Cobertura de linaje
77%
▲ +14 p.p. vs. previo
Confianza media
89%
confianza del proceso
Servicios detectados
6
▲ +3 vs. escaneo previo
Estado del linaje
Calidad del linaje
Cobertura por tipo de recurso (nº de relaciones)
Ver calidad de linaje →Tendencia de cobertura
Últimos 7 escaneosRequiere atención
4 alertasÚltimos escaneos
Procesos
61
Con linaje
47
Confianza
89 %
Cambios recientes
✦ Copilot (Linaje)
Ver historialHaz una pregunta sobre el linaje…
✦ ¿Qué procesos tienen baja confianza?
✦ ¿Qué procesos no tienen linaje extraído?
✦ ¿Qué cambió en el último escaneo?
✦ ¿Cuáles son los activos más críticos?
Recursos destacados
Ver todos →ready.customer_360Tablaglue-orders-curatedGlue Jobs3://acme-raw/orders/Ruta S3Mapa resumido de la plataforma
Ver mapa completoAWS Glue Jobs → Redshift: 156 relaciones · Amazon S3 → Redshift: 12 relaciones. Ruta directa sin pasar por Athena.
Every alert is a link: it leads to the list of processes or relationships that produce it, not to a report someone has to interpret.
Copilot
Asistente inteligente para explorar linajes, procesos y dependencias
Conversaciones
Nuevaready.customer_360gold?¿Qué procesos escriben ready.customer_360?
✦ Claude · claude-sonnet-4-6
Dos procesos escriben esa tabla:
glue-orders-curated · Glue Jobfn-events-enrich · LambdaBasado en 32 relaciones publicadas del escaneo Discovery dev.
Haz una pregunta sobre el linaje…
Contexto de la conversación
Resumen
Discovery dev scan with 259 resources and 6 services. Context sent to the model: 30 resources.
Metadatos clave
Acciones rápidas
The Copilot answers about lineage that is already published and cites where every statement comes from. With the local provider no data leaves the account; with an external one, only what the AI providers policy allows.
Descubrimiento de recursos AWS
Explora y descubre servicios y recursos en tu entorno AWS para generar linaje de datos
Contexto del escaneo
Qué obtendrás
Servicios incluidos
Seleccionar todosEsta cuenta está validada con avisos: no se analizará Glue Jobs ni Step Functions por permisos insuficientes. El escaneo puede ejecutarse igualmente.
Historial reciente · últimos 5
v3 Discovery dev Completadov2 Discovery dev Completadov1 Aceptación instalación limpia CompletadoThe scan modifies nothing: it discovers resources, detects candidates and compares against the previous run. What gets analysed is your choice in step 2, and nothing is published until step 4.
Inventario completo
Censo de la cuenta: todos los recursos, tengan linaje o no
Total de recursos
259
Tablas
33
12,7 % del total
Vistas
45
17,4 % del total
Jobs / Lambdas
39
15,1 % del total
Dashboards / Reportes
0
0,0 % del total
Otros recursos
142
54,8 % del total
Filtros rápidos
Tipo de recurso
Servicio / origen
Etiquetas (tags)
259 recursos encontrados
Orden: Última actualización| Nombre del recurso | Tipo | Servicio | Propietario | Dominio | Criticidad | Linaje |
|---|---|---|---|---|---|---|
ready.customer_360 | Tabla | AWS Glue | data-platform | dwh | Alta | Sí |
curated.crm_orders | Tabla | AWS Glue | data-platform | dwh | Media | Sí |
glue-orders-curated | Job / Lambda | AWS Glue | data-platform | dwh | Media | Sí |
acme-raw | Bucket S3 | S3 | data-platform | — | — | No |
fn-events-enrich | Lambda | AWS Lambda | — | — | — | No |
Showing 1–12 of 259 resources · page 1 of 22. The inventory is the COMPLETE census of the account, with lineage or without it; the catalogue, in the next section, is only what has already been analysed.
Catálogo de recursos
Lo ya analizado, con su confianza y su estado de análisis
Recursos catalogados
259
Procesos
47
Tablas
33
Columnas
43
Servicios AWS
6
Escaneos disponibles
4
Filtros
Tipo de recurso
Entradas del catálogo
Orden: Confianza| Entrada | Tipo | Servicio | Análisis | Confianza |
|---|---|---|---|---|
glue-orders-curated | glue:job | glue | Analizado | 92% |
ready.customer_360 | glue:table | glue | Analizado | 96% |
fn-events-enrich | lambda:function | lambda | Analizado | 88% |
dwh.sales_fact | redshift:table | redshift | Pendiente | — |
The catalogue only holds what has already been analysed. Anything still «Pending» appears, but without confidence: its lineage has not been extracted yet.
Conexiones AWS
Cuentas que el producto puede explorar. Siempre en solo lectura
Acme Ingest DEV
VálidaModo: Appliance
Capacidades concedidas
Data Lake DEV
VálidaModo: Assume role
Capacidades concedidas
Both connections are read-only. In Assume role mode the product assumes a role in your account with an ExternalId that you generate, and the screen states whether the client requires it in its trust policy and when it was last verified. A connection that stops validating is not used: it is flagged and stopped.
Proveedores IA
Modelos que el producto puede usar para el Copilot y la extracción de linaje
Claude API externa
ValidadoThe credential and the model have been confirmed against the provider.
OpenAI API externa
ValidadoThe credential and the model have been confirmed against the provider.
Ollama Local
PredeterminadoIt runs inside the perimeter: no data leaves the account.
Qué datos pueden salir de la cuenta
v1 · cc7549dbe2c4Egress to external providers enabled. The authorised providers are claude and openai. Every class of data leaves with the transformation its rule states, and anything marked «Never leaves» never crosses the perimeter.
| Clase de dato | Regla | Qué incluye |
|---|---|---|
| Identificadores de AWS | Sale pseudonimizado | ARN y Account ID, sustituidos por seudónimos estables. |
| Credenciales | No sale | Claves, contraseñas y tokens. Prohibido de forma inamovible: no configurable. |
| Valores de datos | No sale | Filas, muestras y contenido de columnas. |
| Etiquetas de recursos | No sale | Tags de AWS: suelen llevar correos, centros de coste o tickets. |
| Metadatos de recursos | Sale tal cual | Nombres, tipos y configuración de los recursos inventariados. |
| Metadatos de esquema | Sale tal cual | Nombres de base de datos, tabla y columna. Nunca su contenido. |
| Contenido sintético | Sale tal cual | Texto de prueba que genera el producto para validar un proveedor. |
| Código fuente | Secretos redactados | Código de los Glue Jobs, Lambdas y scripts que se analizan. |
| Sentencias SQL | Secretos redactados | Consultas y DDL encontradas en el código o en el catálogo. |
| Errores y diagnóstico | Secretos redactados | Mensajes de error y logs, recortados a su causa. |
| Pregunta del usuario | Secretos redactados | El texto que se escribe en el chat, y el hilo reciente. |
This policy is deployment configuration: it is defined in config/ai-policy.yml, reviewed in change control and cannot be modified from the application. The fingerprint identifies the exact document being operated with, and is stored with every extraction.
Usuarios
Quién puede entrar, con qué rol y sobre qué conexiones AWS
Administración DEV
TúOperador de evidencia
Debe cambiar contraseñaTres roles, acumulativos
| Puede | Consulta | Operación | Administración |
|---|---|---|---|
| Consultar dashboard, inventario, catálogo, linaje y Copilot | ✔ | ✔ | ✔ |
| Lanzar escaneos y extraer linaje | ✔ | ✔ | |
| Revisar y publicar relaciones | ✔ | ✔ | |
| Validar una conexión AWS | ✔ | ✔ | |
| Gestionar conexiones, proveedores de IA y usuarios | ✔ |
Scope is declared per AWS connection: a user can see one account and not another. Hiding a button is not the protection: every request is authorised on the server.
Product views over a demonstration environment. The account, the processes and the table and dataset names are fictitious.
Column-level lineage
When the code and the SQL allow it, the analysis goes down to column level and identifies the type of transformation. When the evidence is not enough, the relationship is flagged as such instead of being taken as sound.
Four relationships detected between the two tables. Pick one to see its type.
curated.crm_orders
ready.customer_360
direct order_id → order_id
The column is copied unchanged. It is the easiest relationship to support: the SELECT names it the same in source and target.
rename customer_name → client
Same data, different name. The SELECT alias is what makes it traceable; without it the two columns would look unrelated.
cast order_ts → order_date
The type changes, not the content: from timestamp to date. The conversion shows up in the expression, and that is what fixes the transformation type.
aggregation amount → total_amount
The target value summarises several source rows. The relationship holds, but it is no longer one to one, and that changes how impact must be read.
Use cases
What it lets you answer
What breaks if I touch this table or this column?
Know what might break before changing a table, a column or a job.
What it lets you answer
Where does this data come from and who answers for it?
Understand the origin, destination and transformation of the information you govern.
What it lets you answer
What is deployed in this account and how does it relate?
Document an existing AWS platform without starting from hand-drawn diagrams.
What it lets you answer
Which process left this data wrong, and when?
Follow dependencies and data producers to get to the source of the problem.
What it lets you answer
What depends on the thing I want to migrate or retire?
Understand the current systems before migrating, consolidating or refactoring them.
What it lets you answer
What evidence backs up that the data flowed this way?
Provide traceability and evidence on how information travels between systems.
Do any of these scenarios look like your platform?
Request a demoSecurity & AI
The connection to your AWS account is read-only: it discovers resources and processes without modifying anything, and only what you select is analysed.
Rows, samples and column contents never enter the analysis or leave your account. What is analysed is how data moves, not the data itself.
Ollama locally, with nothing leaving your network, or Claude or OpenAI under an egress policy declared class of data by class of data.
We would rather be explicit: today we claim no certifications or compliance seals, and access uses the product's own users, not your corporate SSO yet. The specific security, deployment and data processing requirements are reviewed case by case during the evaluation.
Only the processes you explicitly select are analysed.
You decide which accounts, regions, environments and processes go into each scan.
You create a read-only role in your account with the trust policy the product gives you, protected with an ExternalId. No credentials of yours are stored: the role is assumed for each operation.
The product is deployed in your own AWS environment and works from the inside. For organisations that allow no access at all from outside their perimeter.
In both cases
Every connection is validated and shows what it can and cannot do: discovery, Glue Jobs, catalogue, Lambda, Step Functions or reading code from S3. If a permission is missing you see it before scanning instead of failing halfway through.
Identity and session
Passwords stored with Argon2id. The session lives on the server and the browser only receives an HttpOnly, Secure and SameSite cookie; the token is stored hashed, so a copy of the table does not hand over reusable sessions. It expires on inactivity and has an absolute ceiling on top of that, and repeated failed attempts lock the account temporarily.
Authorisation
Read covers the whole product without being able to change anything. Operate adds launching scans, extracting lineage and reviewing relationships. Administer adds managing connections, AI providers and users. They are cumulative, checked on the server on every request, and a role the code does not recognise ends up with no permissions, not with all of them.
Scope
The permission is not the only thing checked: each user is assigned the connections they are allowed to work on. A team can see its development account and not production. The installation is dedicated per client and serves several AWS accounts.
The metadata does not always contain the relationship: often it is written in the code, in the SQL or in the job configuration. The AI reads those artefacts and infers the relationship, putting in writing what it based that on.
Every proposed relationship carries confidence and evidence, and goes through review before being published. No black box.
Your code does not have to leave
Analysing lineage requires reading the code and the SQL of your processes. That is why you choose where the model runs: Ollama locally, with nothing leaving your network, or the Anthropic or OpenAI APIs if you prefer their capability.
It is chosen per extraction, together with the confidence target and whether the analysis should include SQL, configuration, environment variables and dependencies.
And if you choose an external API
There is no single «allow external calls» switch. Each class of data has its own rule, and the product applies it before talking to the provider:
The product analyses how data moves, not the data: table contents never enter the analysis. The policy is versioned configuration, the interface shows it without being able to edit it, and if it is missing or has a typo all external egress is blocked rather than assuming the most permissive reading.
Adoption model
The product is in a validation phase with its first organisations. That means direct access to the person building it, and priorities that genuinely make the roadmap.
01
We analyse one domain or a limited set of pipelines to prove the value against your own lineage.
02
Progressive extension to more accounts, regions, environments and processes.
03
Tuning of analysis rules, integrations and capabilities to particular needs.
Lineage is persisted in PostgreSQL as the source of truth and, optionally, the graph is synchronised in Neo4j. It can be adopted as a product, as an accelerator for a governance project, or as a platform configured for one specific environment.
Tell us about your architecture and we will assess the fit.
Evaluate a pilotRoadmap
The same technical understanding that makes the lineage possible enables other services on top of the platform. We draw a clear line between what exists today and what is on the roadmap.
Outside of that: access uses the product's own users. Integration with a corporate directory —SSO, SAML or OIDC— is planned, and the authorisation layer is already built so as not to depend on how you authenticate, but it is not there today.
Analysing jobs, functions and SQL to flag complexity, redundancy, duplicated code and expensive transformations.
Today it generates per-resource documentation in Markdown, with its confidence level. What is missing is raising that to architecture documentation with diagrams and keeping it in sync with every scan.
Detecting pipelines with no consumers, duplicated processes, obsolete assets and critical dependencies.
Comparing scans to identify new processes, removed resources, schema changes and their potential impact.
Linking a quality problem to the dataset, the producing process and the source, to speed up root-cause analysis.
Crossing architecture with execution to locate oversized processes, redundant jobs and under-used resources.
Assessing existing platforms with AI and lineage before migrating, refactoring, consolidating or retiring processes.
Proposing SQL tests, contract validations, regression tests and data quality checks from the lineage.
Evolving the lineage into a knowledge graph connecting data, documentation, owners, quality, cost, policies and incidents.
Capabilities marked as roadmap are not available yet and their scope may change. If one of them is a priority for your team, tell us.
Frequently asked questions
Data lineage is the record of how information moves and is transformed across systems: which source a piece of data comes from, which process reads it, what transformations it goes through and which table or dataset it ends up in. It lets you answer impact and data governance questions without depending on what one particular person happens to remember.
Done by hand, it means diagrams and spreadsheets that go stale the moment someone deploys a new job. AI Data Lineage Mapper takes it from the environment itself: it discovers the account's resources in read-only mode and analyses the code, the SQL and the configuration of the processes you select. The processes analysed to extract lineage are Glue Jobs, Lambda functions and Step Functions, and the resulting relationships connect Glue Data Catalog tables, Amazon S3 paths and Redshift objects.
It is lineage that goes down from the table level to the field level: which source column feeds each target column, and through which transformation (direct, rename, cast, calculation, join or aggregation). It is identified when the code and the SQL contain enough evidence; when they do not, the relationship is flagged as such instead of being taken as sound.
Once the lineage is published you query it in the graph or ask the Copilot in plain language. For any table, dataset or S3 path you can see its producing and consuming processes, and walk the lineage upstream and downstream as many levels as you need.
Impact analysis walks the lineage downstream from the asset you want to change and lists the processes, tables and datasets that depend on it. That way you know what might break before you touch a schema, instead of asking on chat and waiting for someone to remember the answer.
Every relationship is proposed with a confidence level and the specific evidence it rests on: file and line of code, SQL statement or job parameter. It then goes through human review before being published, and can be confirmed, rejected or left pending. The AI proposes and a person decides, so there is no black box.
Only if you choose to. The model provider is configured for each extraction: you can use Ollama running on your own infrastructure, in which case neither the code nor the SQL leaves your network, or the Anthropic or OpenAI APIs if you prefer their capability. The decision is explicit and is made before the analysis is launched.
No. Discovery runs in read-only mode and does not modify any resource in your account. The connection can be made in two ways: a read-only role assumed across accounts, protected with an ExternalId and without storing any credentials of yours, or by deploying the product as an appliance inside your own account. On top of that, only the processes you explicitly select are analysed, and you decide which accounts, regions and environments go into each scan.
No. Lineage is inferred from the code, the SQL and the schema metadata —database, table and column names— never from the rows. Data contents, samples and column values are classified as a category that cannot leave towards any model, external or local, and they do not enter the analysis either. AWS tags are excluded too, because they tend to contain email addresses, cost centres or ticket numbers that lineage does not need. The product analyses how data moves, not the data.
With the product's own users and three cumulative roles. Read grants access to the dashboard, the inventory, the catalogue, the lineage and the Copilot without being able to change anything. Operate adds launching scans, extracting lineage and reviewing relationships. Administer adds managing AWS connections, AI providers and users. Beyond the role, each user is assigned the AWS connections they are allowed to work on, so a team can see the development account and not production. Authorisation is checked on the server on every request: hiding a button is a convenience, not the protection. Access does not integrate with a corporate directory over SSO today.
Contact
Every architecture is different. Tell us what you want to understand, document, optimise or govern, and we will assess how to adapt the platform to your environment.