CDP VS DATA CATALOG: YOU'RE PROBABLY ASKING FOR ONE AND NEEDING THE OTHER
Customer data platforms resolve identities into unified profiles for activation; data catalogs document and govern data for trust and discovery. Which to deploy first, how they compound, and the failure modes of confusing them.
A customer data platform (CDP) answers 'who is this customer, across all our systems, right now?' - identity resolution into unified, activation-ready profiles. A data catalog answers 'what data do we have, what does it mean, can I trust it, and may I use it?' - documentation, lineage, ownership, and access governance over your estate. Deploy the catalog first when your problem is trust and findability ('we can't find/agree on data'); deploy the CDP first when your problem is fragmentation of one specific entity ('we can't see one customer'). Most enterprises voicing CDP demands are actually bleeding from catalog problems.
The confusion is structural, not stupid
Both products 'unify data,' both demo beautifully, both get sold to the same buyer. But they unify different things: a CDP unifies RECORDS about one entity type (customers) into profiles; a catalog unifies KNOWLEDGE about all data into an inventory. A CDP makes data actionable; a catalog makes data trustworthy. Marketing wants the first; everyone silently depends on the second.
What a CDP actually does
- →Identity resolution: deterministic + probabilistic matching across email, phone, device, account IDs - the hard, unglamorous core
- →Profile assembly: one timeline of traits and events per resolved person
- →Consent tracking: which uses this person permitted, enforced at activation
- →Activation: segments pushed to ads, email, personalization, support
Failure mode: pointing identity resolution at sources whose quality nobody governs. Garbage in, confidently-merged garbage out - two customers fused into one is a lawsuit, not a rounding error. CDP quality is capped by upstream data quality, which is… a catalog-and-contracts problem.
What a catalog actually does
- →Inventory + search: every table/report/pipeline, findable
- →Semantics: definitions, owners - which 'revenue' is THE revenue
- →Lineage: where a number came from, what breaks if a column changes
- →Trust signals: freshness, quality checks, certification status
- →Access governance: who may see what, requestable in-place
Failure mode: shelfware - cataloging as a documentation project instead of wiring it into workflows (PR checks that block undocumented schema changes, BI that displays certification badges). A catalog nobody's tools consult is a wiki with delusions.
Sequencing without regret
If reports disagree, analysts rebuild the same numbers, and 'can I use this?' takes a meeting - catalog first; a CDP on that swamp just activates the swamp. If your data layer is sane but marketing/support each see a different fragment of the customer - CDP first, scoped to the few sources that matter, catalog those sources as you go. They compound: the catalog certifies the sources the CDP consumes; the CDP's profiles become a governed, cataloged asset themselves. (Both live in our Data Platform for exactly that reason.)
Questions that cut through vendor fog
- Show me identity resolution on OUR two ugliest sources - precision/recall, not adjectives (CDP).
- What happens at activation when consent is absent? (CDP - the right answer involves 'blocked'.)
- Show lineage for one KPI back to raw tables, live (catalog).
- What ENFORCES documentation staying current? (catalog - cadence answers mean shelfware.)