Complete Visibility From Mainframe to Cloud
The organization is a global financial institution with decades of operational history, a highly regulated business environment, and a data infrastructure that reflects both. Core banking operations run on mainframe systems and on premises databases that have accumulated layers of complexity over time. Alongside those established systems, the organization has been actively modernizing its infrastructure: migrating workloads to Snowflake and Databricks, building out Python and Spark based transformation pipelines, and scaling AI and machine learning programs across risk, compliance, and operational functions.
The CDO organization is accountable for governance across both environments, navigating regulatory frameworks including BCBS 239 while enabling the AI and analytics initiatives the business depends on to remain competitive. That accountability requires lineage and governance infrastructure that can operate continuously and at the resolution regulators and AI accountability requirements demand.
In global financial services, data lineage is not a governance aspiration. It is a regulatory requirement. Frameworks including BCBS 239 establish clear expectations for risk data aggregation, traceability, and the ability to explain how data contributed to decisions and reports.
For this financial institution, meeting those expectations had come to depend on a process that could not scale: manual documentation, consulting driven audit preparation, and static lineage artifacts that were outdated almost as soon as they were produced. Audit preparation cycles consumed months of effort, governance capacity was directed at documentation rather than risk management, and exposure grew as AI and machine learning initiatives scaled on data whose lineage was incomplete.
"Audit readiness required continuously updated visibility into how data moved across the organization."
The technical root of the problem was coverage. Existing catalog solutions provided meaningful value in the environments they were built for, primarily SQL based platforms with relatively straightforward transformation logic. They could not follow data as it moved through mainframe systems, through Java and Scala application code, through ETL pipelines with embedded transformation logic, and into cloud environments where Python notebooks and Spark jobs applied further transformations before data reached the models and reports that regulators and business leaders depended on.
Every gap in that coverage was a gap in traceability. When a regulator asked how a risk metric was derived, or when an AI model produced a result that required explanation, the answer had to be assembled manually from documentation that may or may not have reflected the current state of the environment.
The organization's AI and machine learning programs were expanding rapidly across risk scoring, fraud detection, and operational analytics. Regulators in financial services have clear and growing expectations for explainability, data provenance, and the ability to trace a model's inputs back to their origins. Scaling those programs on a lineage foundation that was manually maintained and structurally incomplete was accumulating governance debt that would eventually need to be addressed.
"Manual lineage processes could not keep pace with the scale and complexity of banking systems."
The organization had invested in catalog tooling and governance initiatives. Those investments were not wasted. The problem was that they operated at the wrong layer of the stack and on a timeline that could not match the pace of change. Catalog solutions read metadata and query logs. They cannot read the transformation logic embedded in mainframe JCL, Java application code, Scala pipelines, or Python notebooks, which is precisely where the most consequential transformations in a complex banking environment happen.
Foundational operates differently. Source code analysis across SQL, Java, Scala, Python, COBOL, Spark, and mainframe environments means lineage reflects what the code actually does rather than what metadata suggests. Column level dependencies are captured as a function of the transformation logic itself, not inferred from query patterns.
By deploying Foundational, the organization replaced its manual lineage model with automated, continuously updated visibility across its full hybrid stack, from mainframe source code and established Oracle pipelines through Java and Scala application layers to Snowflake, Databricks, and ML environments. Because Foundational analyzes source code directly rather than relying on query logs or catalog metadata, coverage extended to the application layer where much of the organization's most critical data transformation logic lives.
The most consequential change was not what the lineage covered but how it was maintained. Audit readiness shifted from a periodic project to a continuously updated operational capability. Lineage documentation that had required months of manual effort to assemble for each audit cycle was now generated and kept current by the platform. When regulatory requests arrived, the information needed to respond was available rather than requiring fresh discovery.
"AI governance starts with understanding the origin, movement, and transformation of every critical data element."
With complete lineage across the environments where AI and ML workloads were operating, the organization gained the governance foundation that regulatory expectations for AI explainability require. Training data and operational data inputs to models could be traced to their origins, and the transformation logic applied along the way was documented and auditable.
For the first time, the organization had complete lineage visibility from mainframe source code and established Oracle pipelines through Java and Scala application layers, ETL and orchestration systems, and into Snowflake, Databricks, and ML environments.
For AI and advanced analytics programs, the organization gained complete lineage coverage across AI and ML pipelines, including training data provenance and operational input traceability, governance controls that support regulatory expectations for AI explainability and data accountability, and a documented, auditable record of data provenance that scales with the growth of AI programs rather than lagging behind them.
For the organization overall, the outcome was an improved posture against BCBS 239 and related regulatory frameworks requiring risk data traceability, reduced operational risk from lineage gaps across the hybrid environment, and a governance foundation that scales with cloud adoption and AI growth rather than requiring recurring manual investment to maintain.