Data Governance & Management September 1, 2026 · 21 min read

Start With Uses, Not Words: Why Most Critical Data Element Programs Fail Before They Begin

A regulatory cross-reference of nine major frameworks reveals that usage-first CDE identification, starting from reports, models, and regulatory submissions and tracing backward, is structurally superior to glossary-first approaches. Part 1 of the Critical Data Element Practitioner's Guide.

By Vikas Pratap Singh
#critical-data-elements #data-governance #bcbs-239 #data-quality #regulatory-compliance #data-lineage #financial-services #healthcare #insurance

CDE Practitioner’s Guide: Overview | Part 0 | Part 1 | Part 2 | Part 3 | Part 4 | Part 5 | Part 6

The $7.4 Billion Lesson

In October 2020, the OCC assessed a $400 million civil money penalty against Citibank for deficiencies in enterprise-wide risk management, compliance, Data Governance, and internal controls. The bank had policies. It had a business glossary. It had stewards on an org chart. What it did not have was the ability to trace data from its regulatory reports back through transformation layers to authoritative source systems.

Four years later, the OCC and Federal Reserve levied an additional $135.6 million ($75M OCC; $60.6M Federal Reserve) for “insufficient progress” in remediating those same Data Governance gaps. By that point, Citigroup had spent $7.4 billion between 2021 and 2023 on the technology, consultants, and compensation for its risk-and-controls overhaul and other modernization efforts (American Banker, 2024).

By the time this article went live on September 1, 2026, the Citi story had moved. On December 18, 2025, the Federal Reserve closed three related risk-control notices, and on December 19 the OCC withdrew its 2024 amendment to the 2020 consent order. The 2020 order itself, tied to the $400 million penalty, remains in effect. Regulators eased the reporting burden. They did not declare the data problem solved.

The Citigroup story is not an outlier. It is the most expensive example of a pattern that plays out across regulated industries: organizations build elaborate glossaries and governance frameworks on paper, then discover during an audit or regulatory examination that they cannot answer the one question that actually matters. “Can you show me the data feeding this report, prove it is accurate, and trace it back to its source?”

That question starts from the use. Not from the glossary.

Across the Critical Data Element programs I have worked on, the first instinct was usually the same: define every business term, build a comprehensive glossary, and then decide what qualifies as “critical.” The result was a catalog with thousands of entries, no clear prioritization, and no connection to the regulatory reports or business-critical dashboards the organization actually needed to defend. Programs built that way stalled within a year. When the approach shifted to start from Tier 1 uses, the scope defined itself. The teams identified the regulatory submissions, risk models, and executive dashboards that carried the highest stakes, then traced backward to the source data elements feeding those outputs. Within about three months, they had a defensible CDE inventory of roughly 200 elements with clear ownership, quality rules, and lineage documentation. What follows draws on that composite of programs, not a single case.

This article examines why that shift works, and why the evidence, from regulatory frameworks to enforcement actions to industry benchmarks, consistently supports usage-first CDE identification over glossary-first approaches.

Two Schools of Thought

Part 0 of this series defines what a CDE structurally is: a business-level concept, not a physical column, mapped through a three-layer meta model to its physical implementations. This article assumes that foundation.

Usage-First vs. Glossary-First CDE Identification

The CDE identification debate reduces to a sequencing question: where do you start?

The Glossary-First Approach

The glossary-first school begins with business vocabulary. Organizations define their terms, build a comprehensive business glossary, document definitions and ownership, and then designate certain terms as “critical.” The logic is intuitive: you need to know what your data means before you can decide what matters.

In practice, this looks like:

  1. Convene a working group to define business terms across domains
  2. Build or populate a business glossary (often in a Data Catalog tool)
  3. Apply criticality criteria to each term (regulatory relevance, financial impact, operational dependency)
  4. Tag qualifying terms as Critical Data Elements
  5. Build governance controls around the tagged elements

The appeal is conceptual clarity. You build shared vocabulary first, then layer governance on top. The problem is scope. In practice, a mid-size bank can carry tens of thousands of candidate business terms across its domains, and an insurer with multiple product lines, legacy systems, and acquired entities can carry more. Deciding which of those terms are “critical” without a clear reference point for criticality becomes a political exercise, not an analytical one.

The Usage-First Approach

The usage-first school starts from the other end. Rather than defining all terms and then selecting which matter, it begins with the outputs that carry the highest organizational stakes:

  • Regulatory reports
  • Tier 1 risk models
  • Financial close processes
  • Executive dashboards
  • CCAR submissions
  • Solvency II Quantitative Reporting Templates
  • USCDI-mandated clinical data exchanges

From those outputs, the approach traces backward:

  1. Identify Tier 1 uses: regulatory submissions, risk models, SOX-relevant financial reports, compliance dashboards
  2. Map each use to its constituent data elements (the specific fields that populate the report or feed the model)
  3. Trace lineage from the output back through transformation layers to source systems
  4. Every element on that path is a CDE candidate
  5. Score candidates by usage count and impact (how many Tier 1 uses depend on this element, and what breaks if it is wrong?)
  6. Build the semantic layer (business definitions, ownership, cross-domain mappings) around the identified CDEs

The scope defines itself. You do not need to boil the ocean of business terminology because your starting point, the set of high-stakes outputs, is finite and known. In practice, most enterprises land at roughly 200 to 250 CDEs through this process, a manageable inventory that directly connects to audit and compliance requirements.

What the Regulators Actually Examine

The most compelling argument for usage-first CDE identification is not theoretical. It is regulatory. When an examiner walks into a bank, an insurer, or a healthcare organization, they do not ask to see the business glossary. They ask to see the data feeding specific reports, and they ask you to prove it is accurate and traceable.

I examined nine major regulatory frameworks across financial services, insurance, and healthcare. The pattern is consistent.

RegulationSectorCDE RequirementExplicit or ImplicitApproach Favored
BCBS 239 + ECB RDARR (2024)BankingIdentify CDEs for risk indicatorsExplicit (ECB 2024)Usage-first
SR 11-7BankingAssess quality of all model inputsImplicitUsage-first
SOX 302/404 + PCAOB AS 2201All public companiesKey data in financial controls (IPE testing)ImplicitUsage-first
CCAR/DFAST (FR Y-14)BankingPopulate prescribed data fieldsEffectively explicitUsage-first (regulator-defined)
HIPAAHealthcare18 PHI identifiers, designated record sets, USCDIExplicitMixed
NAIC MAR / ORSAInsuranceMaterial data for statutory reportingImplicitUsage-first
OCC Heightened Standards (12 CFR 30 App D)Banking ($50B+)Data governance for risk frameworkImplicit, enforcement-testedUsage-first
GDPR Article 30All (EU operations)Categories of personal data per processing activityPartially explicitUsage-first
IFRS 17 / LDTIInsurancePolicy-level actuarial data elementsImplicitUsage-first

Every regulation except one portion of HIPAA pushes toward usage-first identification. Here is why each matters.

BCBS 239: The Standard That Made CDEs Explicit

The Basel Committee’s Principles for Effective Risk Data Aggregation and Risk Reporting, published in 2013, established the foundational expectation that banks must be able to aggregate accurate risk data and produce reliable risk reports. The original text uses “material risk data” rather than “Critical Data Elements,” but the direction is clear. Start from risk reports (Principles 7 through 11) and ensure the data feeding them meets accuracy, completeness, and timeliness requirements (Principles 3 through 6).

The ECB’s May 2024 Guide on Effective Risk Data Aggregation and Risk Reporting removed any remaining ambiguity. It explicitly defines Critical Data Elements as “those data elements that are used to calculate the key risk indicators and have a direct or significant impact on the value of the indicator or technical routine.” That is textbook usage-first language: start from the indicator, trace to the elements feeding it.

The ECB Guide also requires complete, up-to-date data lineages “on data attribute level” (starting from data capture through extraction, transformation and loading), not merely at the system level, and requires banks to “clearly outline the scope of application of their data governance framework by explicitly identifying the included reports, models, risk data, and critical data elements.” Reports first, then CDEs. Not the other way around.

The compliance data tells its own story.

BCBS 239 compliance reality:

  • Only 2 of 31 G-SIBs fully comply (BIS, 2023)
  • Not a single principle fully implemented across all banks (PwC, 2023)
  • Six of 11 principles deteriorated or stalled between 2019 and 2022 (BIS, 2023)
  • ECB lists RDARR as top supervisory priority for 2025 through 2027

A decade after the compliance deadline, the world’s largest banks still cannot fully trace their risk data. The ones that documented glossaries and frameworks without operationalizing lineage from reports to source are the ones that stalled.

SR 11-7: Models Define What Data Matters

The Federal Reserve’s SR 11-7 guidance on Model Risk Management requires banks to demonstrate that “the data and information used are suitable for the model” through “a rigorous assessment of data quality and relevance.” The guidance defines a model as having three components: an information input component (data), a processing component (logic), and a reporting component (output).

SR 11-7 never uses the term “Critical Data Element.” But its structure is inherently usage-first: you identify data requirements starting from each model, then ensure quality controls exist for the specific elements feeding that model. The model inventory drives the data inventory.

SOX and PCAOB AS 2201: Financial Reports Drive Data Requirements

SOX Section 302 requires CEO and CFO certification that internal controls ensure “material information” is made known to officers. Section 404 requires management to “assess the effectiveness of their company’s internal controls over financial reporting.”

The operational mechanism is PCAOB Auditing Standard 2201, which requires testing of “Information Produced by the Entity” (IPE). Every report used as audit evidence must have its specific data fields validated for accuracy and completeness.

The IPE testing methodology traces backward through four layers:

  1. Financial statement line item
  2. Significant account
  3. Key business process
  4. Specific data elements within the process

That is a usage-first CDE identification flow embedded directly in the audit methodology.

CCAR/DFAST: The Regulator Defines Your CDEs

For the largest US banks, the Federal Reserve prescribes exact data fields through the FR Y-14 report family. The FR Y-14A (annual), Y-14Q (quarterly), and Y-14M (monthly) templates define hundreds of specific data fields covering balance sheets, income projections, loan-level portfolio data, and capital components.

What this looks like in practice. Every field in these templates is, by definition, a Critical Data Element for stress testing. Banks that cannot populate a field accurately face a direct consequence: the Federal Reserve “applies conservative assumptions to portfolios that cannot be modeled because of missing data.” Missing or poor-quality data translates directly to higher capital requirements. The regulator is not asking banks to build a glossary. It is handing them a CDE list and saying “make sure these are right.”

HIPAA: The One Exception That Proves the Rule

HIPAA is the only regulation in this analysis that partially supports a glossary-first approach. The 18 PHI identifiers (names, Social Security numbers, medical record numbers, and so on) are an enumerated, closed list defined independent of any specific use case. That is glossary-first by design.

But HIPAA’s “designated record set” concept is usage-first. It is defined as records “used, in whole or in part, by or for the covered entity to make decisions about individuals.” The scope is determined by how the data is used, not by what it is called.

And healthcare’s most significant CDE framework, USCDI (United States Core Data for Interoperability), is essentially a regulator-curated CDE list organized by clinical use context. USCDI v5, published in July 2024, added 16 new data elements and 2 new data classes (Observations and Orders). The ONC did not build a glossary and then tag items as critical. It identified the data elements required for interoperable health information exchange and standardized them. Usage-first, codified as a federal standard.

Insurance: Solvency II, ORSA, and IFRS 17

Insurance regulation follows the same pattern.

Solvency II. EIOPA’s Solvency II framework requires data for technical provisions to be “accurate, complete, and appropriate,” with documented quality assessment processes. The Quantitative Reporting Templates define what must be reported. Insurers trace backward from those templates to the policy, claims, and investment data elements feeding them.

ORSA. Own Risk and Solvency Assessment requires proving robust governance through quality data across risk types feeding capital adequacy calculations.

IFRS 17 / LDTI. IFRS 17 requires policy-level granularity with traceability from source to reporting. LDTI (ASU 2018-12) similarly demands accurate, traceable data as the backbone of compliance.

The Insurance Thought Leadership consortium puts it directly: “A policy effective date that drives billing cycles or a loss development factor in a statutory filing is far more ‘critical’ than a seldom-used rating variable, even if both sit in the same table” (Insurance Thought Leadership). Criticality is a business condition determined by downstream use, not a technical attribute.

Why Glossary-First Programs Fail

The regulatory evidence explains why usage-first works. The industry evidence explains why glossary-first doesn’t.

The Scope Problem

A glossary-first program confronts an unbounded problem. How many business terms does a mid-size bank have? Tens of thousands. A large insurer with multiple product lines and legacy acquisitions? Potentially hundreds of thousands. Defining all of them before deciding which are critical creates what practitioners call the “boil the ocean” problem: unlimited scope, no natural stopping point, and no mechanism for prioritization.

The EDM Council’s 2023 Global Data Management Benchmark found that 80% of organizations have Data Governance programs in progress or established. Having a program and having one that produces outcomes are different things, and that gap is where glossary-first initiatives get stuck. They produce documentation. They do not produce outcomes.

Gartner’s Prediction Maps to a Specific Failure Mode

Gartner’s February 2024 prediction that 80% of Data and Analytics governance initiatives will fail by 2027 identifies the root cause as “a lack of a real or manufactured crisis.”

That failure has a familiar shape: too much scope, too little focus, too long a wait for first tangible outcomes. That is a precise description of what happens when you start with a glossary. You spend twelve months defining terms. You hold quarterly working sessions to align definitions across domains. You build a beautiful catalog. And at the end of year one, the CRO asks “so which data elements are we actually monitoring?” and the answer is none, because you never got past definitions.

For a deeper analysis of why governance maturity models focused on formalization fail to produce outcomes, see The Data Governance Maturity Model Most Organizations Get Wrong.

For practitioners: usage-first avoids this by anchoring to business outcomes from day one. The first CDE you identify is one that feeds a regulatory report. The first quality rule you implement monitors an element that, if wrong, triggers a material misstatement. The program demonstrates value before it asks for expansion.

The Enforcement Evidence

Three enforcement actions from 2024 alone illustrate what happens when organizations cannot connect their data to its uses:

Citigroup Entity: Citibank N.A. / Penalty: $400M (2020) + $135.6M (2024) / Root cause: Could not trace data from regulatory reports back through transformation layers to authoritative source systems. / CDE lesson: Governance artifacts without operational lineage do not survive regulatory examination.

The OCC found that Citibank “failed to implement and maintain an enterprise-wide risk management and compliance risk management program, internal controls, or a data governance program commensurate with the Bank’s size, complexity, and risk profile.” The consent order required the bank to “perform an analysis of its current data quality, aggregation, and management and regulatory reporting policies, procedures, and processes to identify all gaps.” Citigroup had governance artifacts. It could not operationalize them because it could not trace data from reports to source. The bank has since spent $7.4 billion between 2021 and 2023 on the technology, consultants, and compensation for its risk-and-controls overhaul and other modernization efforts (American Banker, 2024).

JPMorgan Chase Entity: JPMorgan Chase Bank N.A. / Penalty: $250M OCC + $98.2M Fed = $348.2M (2024) / Root cause: Data elements required for trade surveillance were not systematically identified or fed into surveillance platforms. / CDE lesson: Failing to identify CDEs for a specific use (trade surveillance) left billions of trading instances unmonitored for nine years.

The OCC found that JPMC “operated with gaps in trading venue coverage and without adequate data controls” required to maintain an effective trade surveillance program. The bank failed to surveil “billions of instances of trading activity on at least 30 global trading venues” from 2014 through 2023, a gap the Federal Reserve’s coordinated action also cited in assessing its own $98.2 million penalty. The trading data was not correctly feeding into surveillance platforms. This is a CDE identification failure at scale: the bank did not know which data elements were critical to a specific use and therefore did not ensure those elements were flowing correctly.

Freedom Mortgage Entity: Freedom Mortgage Corporation / Penalty: $3.95M (2024) / Root cause: Systemic compliance management failures caused widespread errors across numerous HMDA reporting fields. / CDE lesson: A usage-first CDE program would have identified every HMDA reporting field as critical and monitored it continuously.

The CFPB found that Freedom Mortgage’s HMDA data submission contained “widespread errors across numerous data fields” because of “systemic problems with its compliance management systems.” This was a repeat offense: in a 2019 order, the CFPB found Freedom Mortgage had submitted inaccurate race, ethnicity, and sex data because loan officers were instructed to select non-Hispanic white whenever an applicant did not provide the information, regardless of accuracy. Freedom Mortgage apparently had neither the identification nor the monitoring.

Where the Glossary Layer Fits

None of this means the glossary is unnecessary. It means the glossary is sequenced wrong when it comes first.

Once you have identified your CDEs through usage tracing, the semantic layer becomes essential for three reasons:

Semantic consistency across uses. The same data concept, “net revenue” for example, may appear in a CCAR submission, a SOX-relevant financial report, and an executive dashboard. If the definition differs across those uses, you have a consistency problem that usage-first identification alone will not catch. The glossary provides the canonical definition that all uses must align to.

Cross-domain discovery. A CDE identified in the Risk domain (“customer credit score feeds the CCAR model”) may also be critical in the Compliance domain (“customer credit score feeds fair lending analysis”). The glossary provides the abstraction layer for discovering these cross-domain dependencies. As TDAN’s analysis notes, the same data element may exist 10 to 70 times across an enterprise. Reporting Data Quality at the business term level, rather than at each individual physical instantiation, is the only scalable approach.

Data contracts as the bridge. Data contracts formalize CDE expectations between producers and consumers: agreed-upon schemas, quality thresholds, SLAs, and change management procedures. The contract requires both a usage context (who consumes this element and what do they need?) and a semantic anchor (what does this element mean across the enterprise?). Usage-first gives you the former. The glossary gives you the latter. Together, they create an enforceable governance mechanism.

The right sequence is: identify CDEs through usage tracing, then build the business glossary around those CDEs. You define the 150 to 500 terms that your CDE inventory requires (the number varies by organizational complexity; see Part 2’s sizing heuristic), not 500,000 terms that nobody will maintain. The glossary stays focused because it serves the CDE program, rather than the CDE program trying to emerge from an unfocused glossary.

For practitioners: identify CDEs through usage tracing (the count depends on your regulatory footprint and organizational complexity). Then define the glossary around those CDEs. Not 500,000 terms. Only the terms your CDE program requires.

Industry-Specific Starting Points

The usage-first approach looks different depending on your sector, because the high-priority uses differ.

Financial Services

Start with:

  • CCAR/DFAST submissions: The FR Y-14 templates define your stress testing CDEs. Every field is a candidate.
  • Risk reports covered by BCBS 239: Key risk indicators and the elements feeding them, as the ECB now explicitly requires.
  • SOX-relevant financial close processes: Data elements flowing through internal controls to financial statements. PCAOB IPE testing will validate these.
  • AML/BSA surveillance: Transaction monitoring data elements. The JPMorgan case demonstrates the cost of gaps here.

The APRA 100 CRDE Pilot offers the most directly applicable model. In 2019, Australia’s prudential regulator asked major banks to identify their 100 most critical data elements based on business impact, map Data Lineage for each, document controls, and resolve quality issues. APRA later extended this approach to life insurers and superannuation entities. The pilot produced six recommendations centered on identifying Critical Data Elements, remediating data issues, enhancing technology platforms, simplifying legacy architecture, and making data more accessible. This is the closest any regulator has come to prescribing a CDE program methodology, and it is explicitly usage-first.

Insurance

Start with:

  • Solvency II Quantitative Reporting Templates (QRTs): These define what must be reported on balance sheet, technical provisions, capital requirements, and own funds. Trace backward from each QRT field.
  • IFRS 17 / LDTI actuarial calculations: Contractual Service Margin, Risk Adjustment, discount rates, and policy-level cash flow data elements. These require traceability from source to reporting.
  • ORSA submissions: Data elements supporting capital adequacy calculations across underwriting, credit, market, operational, and liquidity risks.
  • Statutory filings under NAIC Model Audit Rule: The MAR mirrors SOX 404 for insurers. Key processes and controls over financial reporting drive backward to specific data elements.

Three practitioner-level complications deserve attention here.

The actuarial judgment boundary. Loss Development Factors and claims reserves embed actuarial assumptions governed by Actuarial Standards of Practice (ASOPs). Governing the data element itself (the number in the system) is straightforward. Governing the methodology that produced it is not your CDE program’s job; that belongs to the actuarial function. Draw that boundary explicitly in your CDE register: the CDE program governs the data; the actuarial function governs the judgment. Without that boundary, your governance council will spend meetings debating reserve adequacy instead of Data Quality.

Reinsurance data sharing. CDEs that feed ceded premium calculations or treaty bordereau reports have external stakeholders (reinsurers) who depend on Data Quality but sit outside your governance structure. Quality failures in reinsurance CDEs create both financial risk (miscalculated ceded premiums) and relationship risk (reinsurers who lose confidence in your data may reprice or decline to renew). Your CDE quality monitoring must account for these external dependencies, even though the reinsurer has no seat on your governance council.

Emerging global standards. For globally active insurance groups, the IAIS Insurance Capital Standard (ICS) introduces another regulatory driver for CDE identification. The ICS requires standardized valuation and capital calculations across jurisdictions, which means additional data elements that must be consistent, traceable, and quality-controlled at the group level.

Healthcare

Start with:

  • USCDI-mandated data elements: The ONC has already built the CDE list for interoperability. USCDI v5 covers patient demographics, clinical notes, medications, lab results, social determinants of health, and more. This is a pre-built, regulator-curated CDE framework.
  • CMS reporting requirements: T-MSIS applies a large battery of Data Quality checks to state-submitted Medicaid data, prioritized by severity. The specific fields subject to those checks are your CDEs.
  • HIPAA designated record sets: Records used to make decisions about individuals. If the data drives a clinical or coverage decision, it is a CDE.
  • Quality reporting (HEDIS, CMS Star Ratings): Data elements feeding quality measures and payment adjustments.

Healthcare has an advantage that financial services and insurance lack: a regulator-curated standard (USCDI) that essentially hands you a starter CDE inventory. The NIH’s CDE Repository goes further for clinical research, providing standardized data element definitions built to ISO/IEC 11179 metadata registry standards.

Three practitioner-level complications are worth calling out.

EHR interoperability and system-specific mappings. Epic and Cerner (now Oracle Health) use different internal data models. A “Diagnosis Code” CDE that maps cleanly to ICD-10 in one system may require transformation logic in another. Your CDE register must document system-specific mappings, not just the canonical code set. If you record only “ICD-10-CM” as the standard and ignore how each EHR stores and surfaces that code, your lineage documentation will fail the first time someone traces a reported diagnosis back to its source system.

Clinical coding currency. ICD-10-CM updates annually. CPT codes change on a regular cycle. SNOMED CT publishes biannual releases. CDE quality rules must include code set currency checks: is this code active in the current version of the relevant standard? A diagnosis code that was valid in 2023 but retired in 2024 is not just a data entry issue; it is a quality failure that can cascade into claims denials, quality measure miscalculation, and compliance gaps. Build version-awareness into your CDE quality rules from the start.

The clinician-as-steward problem. Physicians are the ultimate domain experts for clinical CDEs, but they will not participate in governance councils or respond to stewardship workflows. They have patients to see. The practical model is to designate a clinical informaticist or nurse informaticist as the Domain Steward for clinical CDEs. Physician sign-off should be reserved for the certification level only: annual CDE certification reviews, not day-to-day quality monitoring. Programs that list physicians as data stewards on paper but never get their engagement end up with unowned CDEs in practice.

What Has Changed Since March 2026

  • Europe: still slow. The ECB’s supervisory priorities for 2026 to 2028, published in November 2025, say progress on structural risk-data deficiencies “remains slow,” with no improvement in the average sub-score and persistent weaknesses in data governance frameworks, IT architecture, and data accuracy.
  • A US order closed. On March 30, 2026, the OCC terminated JPMorgan Chase’s trade-surveillance consent order, stating that its continued existence was no longer required for safety and soundness.
  • Healthcare’s starter list moved on. ASTP/ONC published USCDI v6 on July 24, 2025, adding six data elements, and as of September 2026 a v7 draft is posted on the same page. Read the USCDI references in this series as v6.
  • The BCBS 239 baseline has not been refreshed. The BIS progress report this article relies on dates from November 2023 and is the most recent one I could find, so its compliance figures remain the newest public data point.

None of this changes the argument. Usage-first identification is how you answer the examiner’s question whether or not an examination is on the calendar.

Do Next

PriorityActionWhy It Matters
Start hereInventory your Tier 1 uses: regulatory reports, risk models, SOX-relevant processes, and compliance dashboards that carry the highest stakesWithout this list, CDE identification defaults to the unbounded glossary approach, the same too-much-scope pattern behind Gartner’s 80% failure forecast.
Start hereAudit your current CDE approach: are you starting from glossaries or from downstream uses?Every major regulation examined favors usage-first identification; glossary-first programs can resequence now.
ThenMap one regulatory report end-to-end as a proof of concept, tracing from the report’s fields back through transformation layers to source systemsOne traced report surfaces lineage gaps and produces CDE candidates in weeks, not months.
ThenIdentify your CDE program sponsor: CRO, CDO, or head of compliancePrograms without executive sponsorship stall at the first resource request.
NextReview enforcement actions relevant to your sector (OCC consent orders, CFPB actions, ECB supervisory findings) for Data Governance failuresEnforcement patterns reveal what examiners actually test: traceability from reports to source.
AdvancedBuild the business case using Gartner’s 80% failure prediction and your own organization’s audit history to justify usage-first over glossary-firstInternal audit findings paired with regulatory cross-references make the strongest resequencing argument.

What Comes Next

This article establishes the “why” of usage-first CDE identification. The remaining parts of this series cover the “what,” “how,” and “at what scale”:

Part 2: Building Your CDE Inventory covers the step-by-step methodology for usage-first identification, including lineage-based tracing, scoring formulas (criticality = usage count multiplied by impact rating), the DAMA-NL weighted scoring model, building the CDE register, and layering the semantic glossary around your identified elements.

Part 3: Operationalizing CDEs addresses the governance controls that turn a CDE list into an operating capability: tiered quality SLAs, column-level lineage requirements, monitoring and remediation workflows, the three-layer accountability model (owner, steward, domain lead), and technology enablement.

Part 4: Scaling and Sustaining covers maturity stages from 10 CDEs to 5,000+, the 12-to-18-month program roadmap, eight documented anti-patterns with real-world examples, AI-assisted CDE discovery, and connecting CDE quality to business outcomes.

Part 5: Measuring What Matters tackles CDE program KPIs versus KRIs, reporting hierarchies from steward dashboards to board summaries, mapping CDEs to risk assessment units and RCSAs, regulatory examination readiness, and the bridge between the CDO’s metrics and the CRO’s risk language.

The glossary-first crowd is not wrong about needing shared vocabulary. They are wrong about where to start. Start with the uses that carry the highest stakes. Trace backward. Build the vocabulary around what you find. That is how you build a CDE program that survives its first audit.

Sources & References

  1. ECB Guide on Effective Risk Data Aggregation and Risk Reporting (RDARR)(2024)
  2. BIS: Progress in Adopting BCBS 239 Principles (Report d559)(2023)
  3. OCC: $400 Million Civil Money Penalty Against Citibank(2020)
  4. OCC: $75 Million Additional Penalty Against Citibank(2024)
  5. Federal Reserve: Enforcement Action Against Citigroup ($60.6 Million)(2024)
  6. American Banker: Citi Fined $136 Million for Alleged Violations of 2020 Consent Orders(2024)
  7. OCC: $250 Million Penalty Against JPMorgan Chase (Trade Surveillance)(2024)
  8. Federal Reserve: Enforcement Action Against JPMorgan Chase ($98.2 Million, Trade Surveillance)(2024)
  9. CFPB: Action Against Freedom Mortgage for HMDA Data Errors(2024)
  10. APRA: Quality Data as an Asset for Boards, Management, and Business(2022)
  11. Federal Reserve SR 11-7: Model Risk Management Guidance(2011)
  12. Gartner: 80% of D&A Governance Initiatives Will Fail by 2027(2024)
  13. McKinsey: BCBS 239 2.0 Resurgence(2024)
  14. PwC: Not a Single BCBS 239 Principle Fully Implemented by All Banks(2023)
  15. ONC: United States Core Data for Interoperability (USCDI)(2024)
  16. EIOPA: Solvency II Quantitative Reporting Templates
  17. EDM Council: 2023 Global Data Management Benchmark Report(2023)
  18. EY: Why BCBS 239 Compliance Is Essential in 2025(2025)
  19. Capco: ECB Final Guidelines Complement BCBS 239(2024)
  20. Solidatus: ECB Expectations on End-to-End Data Lineage(2024)
  21. Data Crossroads: Critical Data Elements, a Practitioner's Perspective(2024)
  22. LightsOnData: Critical Data Elements, Why Important and How to Measure
  23. Insurance Thought Leadership: CDEs Transform Insurance Decisions
  24. PCAOB Auditing Standard 2201: Internal Control Over Financial Reporting
  25. 12 CFR Part 30, Appendix D: OCC Heightened Standards

Stay in the loop

Get new articles on data governance, AI, and engineering delivered to your inbox.

No spam. Unsubscribe anytime.