ReportTransformation

Why AI stalls on data: the foundations banks need before agents can scale

Models are no longer the constraint. Metadata, a shared semantic layer, knowledge graphs and lineage are — and banks that already invest in BCBS 239 compliance are closer to AI-ready data than they may think.

7 min read By · Report
60%
of AI projects unsupported by AI-ready data will be abandoned through 2026, Gartner predicts1

Key takeaways

  • 63% of organisations lack, or are unsure they have, the right data management practices for AI, and Gartner expects 60% of AI projects without AI-ready data to be abandoned through 20261.
  • Banks have a head start they rarely use: BCBS 239 already demands lineage and data quality — yet only 2 of 31 global systemically important banks were fully compliant in 20233.
  • AI-ready data is contextual data: metadata, a shared ontology, and knowledge graphs give agents the meaning that raw tables lack.
  • Context engineering — deciding what information an agent sees at each step — is becoming a core data-management discipline.

When AI pilots stall in banks, the post-mortem rarely blames the model. It blames the data: definitions that differ between finance and risk, customer records that cannot be joined, documents with no metadata, lineage that stops at the data warehouse. Agents make the problem sharper. A copilot that misreads a table produces a bad answer that a person may catch; an agent that misreads it takes a bad action.

Gartner puts numbers on the risk. In a survey of 248 data management leaders, 63% of organisations said they either do not have, or are unsure whether they have, the right data management practices for AI, and Gartner predicts that through 2026 organisations will abandon 60% of AI projects unsupported by AI-ready data1. Among its recommended actions is to move metadata from passive to active2.

Banks already have the mandate — but not the finish

Banks have been under a supervisory obligation to fix their data for over a decade. The Basel Committee's principles for effective risk data aggregation and risk reporting, BCBS 239, require accurate, complete and timely risk data with clear governance and architecture. Yet the Committee's November 2023 progress report found that only two of the 31 G-SIBs assessed were fully compliant with all principles, and that no single principle was fully implemented across all banks3. The European Central Bank's 2024 guide raised the bar further, expecting complete and up-to-date data lineage at data-attribute level — from capture through extraction, transformation and loading — for the risk indicators in scope4.

Exhibit 1

BCBS 239: a decade on, full compliance is rare

Global systemically important banks assessed by the Basel Committee, 2023 (banks)

Note: 31 G-SIBs assessed; 'not yet fully compliant' is the remainder (31 − 2).

Source: Basel Committee on Banking Supervision (BIS), “Progress in adopting the Principles for effective risk data aggregation and risk reporting” (2023)

The overlap with AI readiness is large. The lineage, data dictionaries, quality controls and ownership that BCBS 239 requires are exactly what an agent needs to know where a number came from, what it means and whether it can be trusted. Banks that treat BCBS 239 as a reporting exercise miss the chance to make it their AI data foundation.

From tables to meaning: metadata, ontology and knowledge graphs

Large language models are good at language and weak at institutional meaning. They do not know that ‘exposure’ in the credit-risk mart is post-mitigation while in finance it is gross, or that two customer identifiers refer to the same legal entity. Three layers supply that meaning:

  • Active metadata. Technical, business and operational metadata — definitions, owners, quality scores, usage — captured continuously rather than documented once and forgotten.
  • An ontology or semantic layer. A shared model of the bank's core concepts (customer, account, facility, exposure, product) and how they relate, so that finance, risk and operations query the same meaning.
  • Knowledge graphs. Entities and relationships instantiated from the ontology — ownership chains, guarantor links, product hierarchies — that agents can traverse.

Knowledge graphs also improve retrieval. Microsoft Research found that GraphRAG, which builds a knowledge graph from source documents, substantially outperformed baseline vector-search RAG on questions that require connecting disparate information or summarising themes across a whole dataset5. In banking, those are precisely the questions that matter: who ultimately owns this counterparty, and what is our total exposure to the group?

Context engineering is the new data discipline

Anthropic describes context engineering as the set of strategies for curating and maintaining the optimal set of tokens an LLM sees during inference6. For a bank, that translates into data-management questions: which policies, records and definitions should an agent retrieve for this task; which fields must be masked; how fresh must the data be; and how is each retrieved item traced back to source? Poor context engineering is the root cause of many agent errors that get blamed on the model.

Exhibit 2

What AI-ready data adds to BCBS 239 foundations

Mapping supervisory data requirements to agent needs

FoundationSupervisory anchorWhat agents need from it
Attribute-level lineageECB RDARR guide4Provenance to cite in every answer and action
Data quality controlsBCBS 239 principles3Confidence signals to decide when to escalate to a human
Active metadataGartner AI-ready data actions2Definitions, owners and freshness at retrieval time
Semantic layer / ontologyConsistent risk and finance definitionsShared meaning across finance, risk and operations
Knowledge graph + RAGEntity and ownership resolutionMulti-hop reasoning over relationships5

Note: DaasLabs synthesis; citations indicate the anchoring source for each row.

Source: European Central Bank – Banking Supervision, “Guide on effective risk data aggregation and risk reporting” (2024)

An agent is only as good as the context it is given. In a bank, context is lineage, definitions and entitlements — which is to say, data management. (DaasLabs view)

Two further foundations are often overlooked. The first is entitlements: an agent retrieving context on behalf of a user must see only what that user — and the agent itself — is permitted to see, which means access policies must be expressed in the data layer, not only in applications. The second is unstructured content. Much of the knowledge agents need sits in credit memos, policies, contracts and emails. Bringing those documents under the same metadata, classification and lineage disciplines as structured data is what allows an agent to combine a covenant clause with a live exposure figure and cite both.

Data quality, finally, needs to become machine-readable. Quality scores and freshness indicators that an agent can read at retrieval time allow it to decide when to proceed, when to caveat an answer and when to hand the task to a human. That is a small change in data engineering with a large effect on the reliability of agentic workflows.

A pragmatic sequence

Banks do not need to finish an enterprise ontology before starting. The sequence that works is use-case led: pick a domain where agents will act — reconciliation, KYC, regulatory reporting — model its core concepts, connect lineage and quality scores for the data it uses, and expose that through a governed retrieval layer. Each domain adds to the shared semantic layer, and the second domain is faster than the first.

For executives

What this means for your bank

  1. Assess AI data readiness per priority use case — lineage, quality, metadata, entitlements — rather than enterprise-wide in the abstract.
  2. Reposition BCBS 239 and RDARR programmes as the foundation for AI, with shared funding and shared attribute-level lineage.
  3. Start a bank ontology with the ten to twenty concepts that matter most for your first agentic use cases, and grow it domain by domain.
  4. Build a governed retrieval layer that combines knowledge-graph queries with document search, and logs the provenance of every item an agent sees.
  5. Make context engineering an explicit responsibility of data owners, not only of AI engineers.
Put it to work

How DaasLabs can help

Benchmark data and AI readiness with our maturity diagnostic.

Take the maturity assessment

Explore the DaasLabs Data Fabric Framework: metadata, semantic layer, knowledge graph and lineage.

Explore the framework

Data strategy and governance accelerator with automated quality scoring.

Explore the accelerator

Data engineering, governance and AI services.

See data engineering services

Sources

  1. 1
  2. 2
  3. 3
  4. 4
    Guide on effective risk data aggregation and risk reporting (opens in a new tab) European Central Bank – Banking Supervision, 3 May 2024
  5. 5
  6. 6

Figures are drawn from the cited public sources. Opinions labelled “DaasLabs point of view” are our own.

Stay informed

Get new banking insights in your inbox

New perspectives on AI, data and transformation in banking — a few times a month. Browse all insights.

AI
AI Analystagentic

I'm the DaasLabs AI Analyst, working with tools rather than from memory. I can:

  • Query the live platform APIs (disputes, fraud, AML, recon, revenue…)
  • Report what the digital workforce is doing: runs, approvals, overrides
  • Start an agent run on a real case and hand you the link to watch it

Every answer shows the tools it used.