Technical guide
Analysing your pipeline code with AI without it leaving your infrastructure
Extracting real lineage means reading the code and the SQL of your processes. That is the objection that stops many projects, and it has a solution: run the model inside your own network.
Last updated: 16 August 2026
Why the code has to be read
Metadata does not contain the relationship. The catalog knows a table exists, but the
sentence that says "this column comes from that one, through a cast" is
written in a PySpark script or in a SQL statement inside the job.
Any tool promising column lineage without looking at the code is either inferring it from names or asking you to write it yourself. Reading the artefacts is not a design preference: it is the only place the information exists.
What sending it to an external API involves
If the analysis runs through a third-party API, your process code leaves your network. It is worth being precise about what that means, without dramatising or minimising it:
- ETL scripts usually include table, schema and path names that reveal the structure of the business.
- They sometimes include business logic that is itself sensitive information: how a score is calculated, which rules determine a segment.
- Even if the provider does not train on your data, one more data processor appears and, if it sits outside the EEA, an international transfer to document.
For many organisations that is perfectly acceptable with a contract in place. For others —banking, healthcare, the public sector— it is a flat no, and the conversation ends there.
The alternative: the model inside your network
A language model running locally, for example through Ollama, changes the picture completely. The code and the SQL are analysed on your own infrastructure and never travel to a third party. There is no new provider, no international transfer, and no data processing annex to negotiate before starting a trial.
The real trade-off is capability: the models that fit on your own machine are less powerful than the best commercial ones. On a narrow task such as interpreting the reads and writes of an ETL script, with the right context, the result is usable — but it is honest to say there is a difference.
How to decide in each case
It does not have to be a permanent choice, nor the same one for the whole platform:
- Local model when the code is sensitive, when the client allows no data egress, or for a first trial without going through procurement and legal.
- External API when the code is not particularly sensitive and you want maximum precision on complex processes.
What matters is that the decision is explicit and made per extraction, not a condition imposed by the tool.
What does not change: evidence and review
Running the model locally does not make it infallible. AI-generated lineage still needs the same guarantees: a confidence level per relationship, evidence traceable to the file and the line, and human review before publishing.
The AI proposes; a person confirms, rejects or leaves pending. Without that, it does not matter where the model runs: it is still a black box.
How AI Data Lineage Mapper solves it
The model provider is chosen for each extraction: Ollama running locally, or the Anthropic or OpenAI APIs. Alongside that choice you configure the confidence target and whether the analysis should include SQL, configuration, environment variables and dependencies.
The AWS connection is read-only, through a role assumed with an ExternalId or as an appliance deployed in your own account, and the resulting lineage is persisted in PostgreSQL with optional graph synchronisation in Neo4j. The whole path can stay inside your perimeter.