How To Integrate A CDP With Databricks
Blog
9/03/26
How To Integrate A CDP With Databricks
Integrating a CDP with Databricks in 2026 means choosing between two architectures, not simply selecting integration tools.
Path A is the external CDP model. A packaged or composable CDP connects to Databricks as the customer data layer. Databricks remains the system of record for customer profiles, behavioral events, ML scoring, and governance, while the CDP or adjacent activation tools handle identity resolution, segmentation, orchestration, and channel activation.
Path B is the native embedded model. Databricks CustomerLake, launched in June 2026 and currently in Private Preview, runs CDP capabilities directly inside the Databricks lakehouse. Instead of moving customer data into a separate CDP store, CustomerLake uses the Databricks environment itself for Customer 360 identity resolution, audience building, AI assisted activation, and governance through Unity Catalog.
That path decision must come before tool selection.
An organization that starts building an external CDP integration and then decides six months later to evaluate CustomerLake is not making a small connector change. It is reconsidering the operating model for customer data, identity resolution, audience activation, governance, and AI agent access.
At Stable Kernel, we advise enterprise teams to treat Databricks as an exceptionally capable CDP data foundation, but not automatically as a complete CDP program by itself. Databricks provides powerful native capabilities for ingestion, transformation, governance, ML, natural language querying, and zero copy data sharing. Lakeflow Connect supports managed ingestion from enterprise sources. Structured Streaming supports near real time event processing. Mosaic AI supports in lakehouse scoring for churn, lifetime value, propensity, and next best action use cases. Unity Catalog provides governance, lineage, access control, and auditability across Databricks data assets. Delta Sharing enables zero copy analytics access for partners, BI teams, and clean room workflows.
But on Path A, Databricks does not replace every CDP function. It does not automatically provide production tested journey orchestration, marketer friendly no code audience workflows, every prebuilt marketing connector, consent capture from customer facing channels, or sub second Profile API serving for in session personalization.
CustomerLake is designed to close more of those gaps for organizations that choose Path B. But because CustomerLake is still in Private Preview, the right answer depends on maturity, risk tolerance, use case timing, marketing team needs, identity resolution requirements, and AI roadmap.
The question is not, “Can Databricks support a CDP?”
It can.
The stronger question is, “Should the organization connect an external CDP to Databricks, or evaluate CustomerLake as the embedded CDP alternative?”
The Two Path Decision Comes Before Integration Tools
The two path decision determines every downstream architecture choice.
If the organization chooses Path A, the design work focuses on connecting Databricks to a packaged CDP, composable CDP, reverse ETL platform, experimentation system, orchestration tool, and activation destinations. Databricks becomes the governed customer data layer, while external tools execute parts of the CDP operating model.
If the organization chooses Path B, the design work focuses on evaluating CustomerLake’s embedded capabilities, validating identity resolution against the organization’s own data, confirming activation pathway maturity, designing agentic workflow governance, and determining whether the Private Preview timeline aligns with business needs.
Both paths can be valid. The mistake is deciding based on vendor momentum instead of architecture fit.
Decision Factor 1: Databricks Maturity
The first question is whether Databricks is already the primary data platform for customer data.
- Path A is usually the better starting point when Databricks is only one of several data platforms. Many enterprises still operate across Snowflake, BigQuery, legacy warehouses, CRM databases, ecommerce platforms, and other systems. In that environment, the CDP may need to unify customer data from several places, not only Databricks.
- Path A also makes sense when Databricks is present but not yet mature. If Unity Catalog governance is incomplete, Delta Lake tables are not cleanly organized, the medallion architecture is inconsistent, or the data engineering team has not yet built trusted customer gold tables, an external CDP or composable layer may help the organization deliver near term value while Databricks maturity catches up.
- Path B becomes more compelling when Databricks is already the canonical customer data platform. The customer data the CDP needs already lives in Databricks. Unity Catalog governs access. Delta Lake tables are organized and maintained. Data engineering actively owns the lakehouse. In that context, moving customer data into a separate CDP store may create unnecessary duplication.
Decision Factor 2: GA Readiness
The second question is whether the organization can accept a Private Preview product for a CDP program.
CustomerLake’s Private Preview status is not a small footnote. It matters for procurement, compliance, support, SLAs, training, internal change management, and executive risk tolerance.
- Path A is the safer choice when the organization needs a generally available, production tested CDP architecture today. If the CDP must support active campaigns this quarter, pass compliance review, satisfy vendor SLA requirements, and serve business critical activation workflows, Private Preview risk may be unacceptable.
- Path B warrants evaluation when the organization has a 6 to 18 month CDP roadmap, can participate in a beta or preview program, and has enough internal engineering capacity to validate immature capabilities before full production dependence.
Most organizations in 2026 will continue on Path A while monitoring CustomerLake’s GA timeline.
Decision Factor 3: Marketing Team Self Service
The third question is whether marketing needs to operate without SQL or data engineering support.
- Path A is often stronger when marketing operations requires a mature no code or low code interface for segment building, journey configuration, campaign audience management, and activation. Packaged CDPs and mature customer engagement platforms have spent years refining marketer facing workflows, permissions, training patterns, and operational playbooks.
- Path B may work when the marketing team can operate through Genie, Campaign Agents, structured audience workflows, or a data engineering supported operating model. That may be appropriate for organizations where marketing and data teams already collaborate closely, where self service does not mean full independence, or where agent assisted audience building is part of the future operating model.
The practical question is not whether a marketer can ask a natural language question. It is whether the marketing team can reliably build, validate, approve, activate, monitor, and iterate customer audiences inside the workflow without creating governance gaps.
Decision Factor 4: Agentic AI Priority
The fourth question is whether agentic AI is a near term CDP requirement.
- Path A is usually sufficient when the current objective is reliable segmentation, suppression, activation, reporting, and personalization for known workflows. If AI assisted activation is a 24 to 36 month roadmap item, it should influence architecture flexibility, but it does not need to drive every decision today.
- Path B becomes more compelling when agentic AI is a 12 to 18 month requirement. AI agents need governed, fresh, unified customer data. If agents must build audiences, recommend offers, optimize campaigns, and adjust journeys, running those capabilities closer to the lakehouse reduces reconciliation risk between an external CDP store and the data platform.
CustomerLake’s potential advantage is that Profile Agents and Campaign Agents can operate directly against Databricks governed data rather than relying on a round trip through a separate CDP environment.
Decision Factor 5: Identity Resolution Requirements
The fifth question is how complex identity resolution needs to be.
- Path A is often better when identity resolution requires probabilistic matching, household graphs, cross brand relationships, partner identity, low confidence identifiers, or match rate targets above what an internal SQL based implementation can reliably support. In those cases, a packaged CDP or dedicated identity resolution tool may provide stronger out of the box capability.
- Path B deserves evaluation when identity resolution is mostly deterministic. If the organization primarily matches on hard identifiers such as email, phone, loyalty ID, account ID, or authenticated customer ID, and the data engineering team can validate match quality inside Databricks, CustomerLake may be a better fit.
The decision rule is straightforward. Choose Path A when the organization needs GA readiness, mature marketer self service, complex probabilistic identity resolution, or CDP coverage across multiple data platforms. Evaluate Path B when Databricks is already the primary governed customer data platform, Private Preview risk is acceptable, marketer workflows can operate through Genie or Campaign Agents, and identity resolution can be validated inside the lakehouse.
The Databricks Native CDP Capability Map
Before selecting third party CDP tools, enterprise teams should understand what Databricks already provides.
A mature Databricks environment may already contain 60 to 70 percent of the technical foundation required for a CDP program. The remaining gaps may be CDP schema design, identity resolution implementation, reverse ETL activation, hot store serving, or marketing workflow design rather than a full new platform.
Lakeflow Connect For Managed Ingestion
Lakeflow Connect supports managed data ingestion from enterprise sources into Databricks.
For CDP programs, this can bring CRM records, ecommerce transactions, loyalty activity, support data, financial data, and operational tables into Delta Lake without building every pipeline from scratch. In a Path A architecture, Lakeflow Connect feeds Databricks so the external CDP or reverse ETL layer can operate against governed customer tables. In Path B, Lakeflow Connect becomes part of the CustomerLake ingestion foundation.
The limitation is important. Lakeflow Connect is an ingestion capability, not an activation capability. It moves data into Databricks. It does not replace reverse ETL, CDP destination connectors, or customer engagement tools that send audiences to downstream systems.
Lakeflow Declarative Pipelines For Medallion Transformation
Lakeflow Declarative Pipelines support the transformation layer that moves customer data from raw records to CDP ready outputs.
In a Databricks medallion architecture, bronze tables store raw source data, silver tables store cleaned and standardized records, and gold tables store identity resolved customer profiles, segment membership, and scored attributes.
For CDP architecture, this is critical. External CDP tools should not read raw bronze tables. They should read the gold layer, where customer data has been cleaned, deduplicated, governed, and resolved into canonical profile structures.
Lakeflow Pipelines can maintain this transformation logic continuously. dbt on Databricks can complement the pattern for SQL based transformations, testing, documentation, and quality checks.
Structured Streaming For Near Real Time Events
Databricks Structured Streaming supports near real time event processing from sources such as Kafka, Kinesis, Azure Event Hubs, and other streaming systems.
This matters when CDP use cases require more than daily batch refresh. Cart abandonment triggers, churn signals, recent converter suppression, session behavior, and loyalty updates may need data freshness measured in minutes rather than hours.
Structured Streaming is best suited for Tier 2 use cases, where the response window is roughly 5 to 60 minutes.
It should not be confused with sub second profile serving. For Tier 1 use cases, such as in session personalization or live AI agent profile lookup inside a customer interaction, Databricks usually needs to feed a hot store such as Redis or Lakebase. The Profile API then serves from that low latency store instead of querying Delta Lake directly during the customer request.
Mosaic AI For In Lakehouse Scoring
Mosaic AI gives CDP teams a way to run machine learning workflows inside the Databricks lakehouse.
That can include churn risk scoring, lifetime value modeling, propensity scoring, product recommendation inputs, next best action models, and automated model training. MLflow supports experiment tracking and model versioning, while scores can be written back into gold layer Delta tables for downstream activation.
The practical pattern is simple: compute the score in Databricks, store it in the customer profile or ML score table, and activate it through reverse ETL or serve it through a hot store when the use case requires real time access.
Mosaic AI is strongest for batch and near real time scoring. It should not be assumed to replace a low latency inference or serving architecture for every in session decision.
Unity Catalog For Governance
Unity Catalog is the governance layer that makes Databricks viable as a CDP foundation.
It supports access control, lineage, audit logs, row level and column level policies, and governance across tables, models, volumes, functions, and AI assets. For customer data, this matters because CDP governance must control who can see PII, which systems can activate which attributes, how consent restricted data is handled, and how customer records are deleted or audited.
In Path A, Unity Catalog governs how external tools access Databricks. Reverse ETL platforms and CDPs should connect through service accounts with limited access to the specific gold tables and columns they need.
In Path B, CustomerLake uses Unity Catalog as the native governance layer for customer profiles, AI agents, and activation workflows.
Unity Catalog does not replace a consent management platform. Consent still needs to be captured from customer facing channels and written into Databricks. Unity Catalog enforces access and governance once the data is there.
Delta Sharing And Lakehouse Federation For Analytics Access
Delta Sharing enables zero copy access to Databricks data for external consumers, analytics partners, clean room workflows, and non Databricks systems.
For CDP programs, this is useful for BI access, partner measurement, clean room attribution, third party enrichment, and privacy preserving collaboration. A clean room can query customer conversion data and media impression data without requiring raw PII to move freely between organizations.
Lakehouse Federation supports the inverse problem: querying external systems such as Snowflake, BigQuery, or operational databases from Databricks without moving the data first. That can help organizations with customer data split across platforms.
The limitation is that these are analytics and query access patterns. They do not replace operational activation. Sending audiences to ESPs, CRMs, ad platforms, or service tools still requires reverse ETL or activation connectors.
Genie And Agent Bricks For Agentic Workflows
Genie gives business users a natural language interface for querying governed Databricks data. Agent Bricks and the Mosaic AI Agent Framework support agentic workflows, including agents that can reason over governed data and perform structured tasks.
For CDP programs, these capabilities matter most in Path B. CustomerLake uses natural language and agentic interfaces to support audience building, customer profile analysis, campaign decisioning, and AI assisted activation.
The limitation is maturity and governance. Natural language querying works best when the underlying tables are well modeled, documented, and governed. Agentic workflows also require clear approval rules. The architecture must define which AI initiated audience changes require human approval, which actions are logged, which data the agent can access, and how decisions are audited.
Path A: Four Integration Patterns For External CDPs On Databricks
For organizations that choose Path A, Databricks becomes the customer data foundation and external tools connect around it.
The strongest Path A programs do not use one pattern for every use case. They route each use case by latency, governance, activation requirement, and cost.
Pattern 1: Batch Ingestion Into Databricks
Pattern 1 moves source system data into Databricks on a schedule.
CRM exports, ecommerce transactions, loyalty records, support tickets, billing data, POS records, and historical event data flow into bronze Delta Lake tables through Lakeflow Connect, Fivetran, Auto Loader, or file based ingestion. Lakeflow Declarative Pipelines and dbt then transform those records into silver and gold tables.
This pattern fits Tier 3 use cases, including:
- Daily audience refresh
- Weekly campaign segmentation
- Historical customer analysis
- ML model training datasets
- Compliance reporting
- Batch identity resolution
- Executive dashboards
The failure mode is latency. If a customer abandons a cart at 2:14 PM and the next batch runs overnight, the CDP cannot use that signal for a 15 minute recovery trigger.
Use batch when batch is enough. Do not force daily ingestion into use cases where the value of the signal decays within minutes.
Pattern 2: Near Real Time Streaming With Structured Streaming
Pattern 2 streams behavioral events into Databricks and processes them through Structured Streaming.
Web events, mobile events, server side events, cart activity, login events, purchase events, loyalty redemptions, and other high value behavioral signals can flow through Kafka, Kinesis, or Azure Event Hubs into Databricks. Structured Streaming processes those events and writes results into Delta Lake tables with seconds to minutes of latency.
This pattern fits Tier 2 use cases, such as:
- Cart abandonment detection within a short trigger window
- Churn signal detection
- Recent converter suppression
- Loyalty tier updates
- Session level behavior summaries
- Near real time segment refresh
The failure mode is overextending streaming into Tier 1 serving. Structured Streaming can keep Databricks fresh. It does not automatically make Databricks a sub 100 millisecond Profile API.
For in session personalization, Pattern 2 should feed a hot store. The live experience should read from Redis, Lakebase, or another low latency serving layer.
Pattern 3: Reverse ETL Activation From Databricks
Pattern 3 moves computed customer data out of Databricks and into operational tools.
This is the activation pattern. Databricks computes the segment, score, suppression list, customer attribute, or audience membership. Reverse ETL then writes that output to Braze, Iterable, Salesforce, Google Ads, Meta, The Trade Desk, Customer.io, support tools, or a packaged CDP.
Common tools include Hightouch, Fivetran Activations, GrowthLoop, and RudderStack Warehouse Actions.
This pattern fits any use case where Databricks is the computation layer but another system executes the customer interaction.
Examples include sending a churn risk score to Salesforce, delivering a propensity audience to Google Ads, syncing a loyalty segment to Braze, or updating a packaged CDP with Databricks computed attributes.
The failure mode is assuming this remains zero copy. Reverse ETL creates downstream copies. That may be necessary, but it must be governed. Databricks output tables need data contracts so schema changes in Lakeflow or dbt do not silently break activation mappings.
Pattern 4: Delta Sharing And Lakehouse Federation For Analytics
Pattern 4 supports analytics, collaboration, and zero copy access.
Delta Sharing allows partners, BI tools, clean rooms, and external systems to query Databricks Delta tables without creating physical copies. Lakehouse Federation allows Databricks to query external platforms without migrating all data first.
This pattern fits:
- Clean room measurement
- Partner attribution
- BI dashboards
- Cross platform analytics
- Third party data enrichment
- Multi platform customer data exploration
- Privacy preserving audience analysis
The failure mode is using analytics access as if it were activation. Delta Sharing can let another system read an audience. It does not send that audience to an ESP or ad platform. Operational activation still requires Pattern 3.
The routing rule is clear. Tier 3 batch use cases use Pattern 1. Tier 2 near real time use cases use Pattern 2. Activation use cases use Pattern 3. Analytics and collaboration use cases use Pattern 4. Tier 1 in session use cases require a hot store fed by Pattern 2 and served through a Profile API.
Path B: Databricks CustomerLake Honest Assessment
CustomerLake is the most complete expression of the embedded in lakehouse CDP model. It deserves serious evaluation for enterprises already running Databricks at scale.
It also requires honest scrutiny.
What CustomerLake Is
CustomerLake is an agentic CDP embedded inside Databricks.
Instead of operating as a separate customer data platform with its own storage layer, CustomerLake runs CDP capabilities where the customer data already lives. Its core promise is that Customer 360, identity resolution, audience building, AI assisted campaign decisions, and activation can happen inside the Databricks lakehouse under Unity Catalog governance.
The key components include Profile Agents, Campaign Agents, Infinity Campaigns, Genie, Mosaic AI, Lakeflow, and Unity Catalog.
Profile Agents support customer identity resolution and Customer 360 profile construction. Campaign Agents support audience building, next best action recommendations, and activation workflows. Infinity Campaigns represent always on personalization loops rather than discrete campaign cycles. Genie provides natural language interaction with customer data. Unity Catalog governs access, lineage, and auditability.
The Architecture Advantage
CustomerLake’s strongest advantage is that it removes the reconciliation problem.
In Path A, Databricks and the external CDP may both hold versions of customer truth. Databricks may have the freshest transaction data. The CDP may have a different identity graph. The activation tool may have another audience membership state. Quarterly audits often reveal differences that require engineering time to reconcile.
CustomerLake reduces that duplication by making the lakehouse the CDP execution environment.
This matters even more for AI agents. Agentic workflows need fresh, governed, unified customer data. If the agent must query a CDP that then queries Databricks, or act on a profile that was synced from Databricks hours earlier, the agent is operating against a lagging representation of the customer.
CustomerLake’s promise is direct access to governed customer data in the same environment where AI and ML workflows already run.
The Honest Limitations
CustomerLake’s advantages are meaningful, but several limitations need to be validated before enterprise commitment.
- First, CustomerLake is in Private Preview. Organizations that need a production ready, generally available CDP today may not be able to accept the risk. SLAs, support models, compliance reviews, and training materials may not yet match mature packaged CDP vendors.
- Second, marketer facing workflow maturity needs validation. Genie and Campaign Agents are compelling, but enterprise marketing teams need more than query access. They need approval workflows, segment governance, journey management, campaign QA, role based permissions, training, and operational confidence.
- Third, activation pathway maturity must be tested. Mature packaged CDPs often have large catalogs of production tested marketing, CRM, analytics, and engagement connectors. CustomerLake’s activation pathways should be validated against the organization’s top destinations, expected audience volumes, sync cadence requirements, suppression workflows, and failure recovery needs.
- Fourth, identity resolution quality must be benchmarked on the organization’s own data. Industry claims are not enough. The evaluation should test CustomerLake against actual first party data, including known customers, anonymous to known transitions, duplicate records, loyalty IDs, app identifiers, email hashes, and edge cases.
The CustomerLake decision should be practical. Evaluate Path B when Databricks is the primary governed data platform, Private Preview risk is acceptable, the roadmap includes agentic AI, and the organization can validate identity and activation quality before production dependence. Continue with Path A when the business needs a GA ready CDP today, marketer self service is non negotiable, activation complexity is high, or Databricks is not the primary customer data platform.
The Component Map For A Path A CDP Databricks Stack
A Path A CDP Databricks stack usually includes six layers.
Not every layer is required on day one. The right sequence depends on use case maturity.
Layer 1: Event Collection
The event collection layer captures behavioral events from web, mobile, server side, ecommerce, product, loyalty, and edge systems.
RudderStack, Snowplow, Segment, server side SDKs, Kafka, Kinesis, and Azure Event Hubs may all participate. High frequency, lower value events can be routed through batch or lower cadence ingestion. High value behavioral signals that drive activation should use streaming when freshness matters.
Layer 2: Bronze Raw Data Layer
The bronze layer stores source data exactly as received.
This layer should function as the append only audit log. It gives the organization replayability, traceability, and a way to rebuild downstream profiles if transformation logic changes.
Bronze data should not be read directly by the CDP for activation. It is too raw, inconsistent, and source specific.
Layer 3: Silver And Gold Transformation Layer
The silver layer cleans and standardizes customer data. The gold layer creates CDP ready outputs.
Gold tables should include canonical customer IDs, unified profiles, segment membership, consent attributes, derived scores, and activation ready audience structures. This is where Lakeflow Pipelines, dbt on Databricks, Databricks SQL, and identity resolution tooling become central.
The external CDP or reverse ETL tool should read from gold tables only.
Layer 4: Mosaic AI Scoring
Mosaic AI adds predictive intelligence to the CDP stack.
Churn risk, lifetime value, propensity to purchase, product affinity, and next best action inputs can be calculated in Databricks and written back to gold tables. Those scores can then activate through reverse ETL or serve through a hot store when latency requires it.
The key is freshness governance. A churn score may refresh daily. A cart urgency signal may need minutes. A live personalization decision may need pre computed outputs in a hot store.
Layer 5: Hot Store For Tier 1 Use Cases
The hot store is optional, but essential for Tier 1 use cases.
Databricks is the cold store and computation layer. Redis, Lakebase, DynamoDB, or another serving layer can hold the small set of attributes needed for sub second profile access.
This layer should not be added for sophistication. It should be added because a specific customer experience requires it.
Examples include in session personalization, fraud decisions before transaction completion, live AI agent context, cart state decisions, and real time consent checks.
Layer 6: Activation
Activation moves Databricks computed outputs into downstream tools.
Hightouch, Fivetran Activations, GrowthLoop, RudderStack Warehouse Actions, or packaged CDP connectors can deliver customer data from Databricks gold tables to marketing, CRM, paid media, support, and personalization destinations.
Every activation path should include data contracts, record count reconciliation, destination error monitoring, and ownership. The risk is not only whether data moves. The risk is whether the right data arrives in the right destination at the right time in the right format.
A practical Path A sequence is to build Layers 1 through 3 first, then add activation through Layer 6. Layer 4 comes next for ML scoring. Layer 5 is added only when Tier 1 latency requirements appear.
How Stable Kernel Designs CDP Databricks Integration Architectures
Stable Kernel designs CDP Databricks integration architectures across both Path A and Path B without defaulting to a single vendor position.
The work begins with the two path decision.
The Two Path Assessment Comes First
Stable Kernel evaluates Databricks maturity, GA readiness, marketing team self service requirements, agentic AI priority, and identity resolution complexity before recommending any integration pattern or vendor.
For organizations that qualify for CustomerLake evaluation, Stable Kernel helps define the Private Preview assessment plan. That includes identity match rate benchmarking against the organization’s own data, activation pathway validation for priority destinations, and agentic workflow governance design.
For organizations routed to Path A, Stable Kernel specifies the full external CDP on Databricks architecture: event collection, medallion architecture, identity resolution, Mosaic AI scoring, reverse ETL, data contracts, observability, and hot store requirements where needed.
Native Capability Assessment Prevents Overbuying
Before recommending third party tools, Stable Kernel audits which Databricks capabilities are already licensed, configured, and mature.
That includes Lakeflow Connect, Lakeflow Pipelines, Structured Streaming, Unity Catalog, Mosaic AI, Delta Sharing, Databricks SQL, and existing data quality controls.
In many mature Databricks environments, the missing element is not another platform. It is the CDP architecture that connects existing capabilities into a governed operating model.
Implementation Prioritizes The Golden Record
The first implementation milestone is the golden record.
Before activation, AI scoring, personalization, and journey orchestration can work, the organization needs a trusted customer profile in the gold layer. That means source ingestion, schema governance, identity resolution, consent attributes, survivorship logic, and profile validation must come before broad activation.
Once the golden record is reliable, the team can add segments, scores, reverse ETL, clean room sharing, and Tier 1 hot store serving.
Stable Kernel designs CDP Databricks integration architectures from the two path decision through medallion architecture specification, identity resolution implementation, Mosaic AI scoring, activation, and CustomerLake evaluation for organizations whose Databricks maturity and roadmap make Path B the right fit.
Reflection Questions For Executives
- Which path is the organization actually choosing: external CDP connected to Databricks, or CustomerLake embedded inside Databricks?
- Is Databricks already the canonical customer data platform, or is customer data still fragmented across several systems?
- Can the CDP program accept Private Preview risk, or does it need a generally available platform today?
- Does marketing require mature no code self service, or can teams operate through Genie, Campaign Agents, SQL supported workflows, or a hybrid model?
- Are the most important identity resolution requirements deterministic, probabilistic, cross brand, or partner based?
- Which use cases require daily, hourly, near real time, or sub second response?
- Which Databricks capabilities are already licensed and mature enough to support the CDP roadmap?
- Does the activation layer have data contracts, record count reconciliation, and destination health monitoring?
FAQ
How Do You Integrate A CDP With Databricks?
Integrating a CDP with Databricks begins with the two path decision. Path A connects an external packaged or composable CDP to Databricks as the customer data layer. Path B evaluates Databricks CustomerLake as the native embedded CDP alternative. For Path A, the four core integration patterns are batch ingestion with Lakeflow Connect or Fivetran, near real time streaming with Structured Streaming and Kafka, reverse ETL activation from Databricks gold tables, and Delta Sharing for analytics and collaboration.
What Is Databricks CustomerLake?
Databricks CustomerLake is an agentic CDP embedded natively inside the Databricks lakehouse. It runs Customer 360 identity resolution, audience building, AI assisted campaign decisions, and activation workflows inside Databricks rather than requiring a separate CDP data store. CustomerLake uses Databricks capabilities such as Unity Catalog, Lakeflow, Mosaic AI, and Genie. As of July 2026, it is in Private Preview, so organizations should validate maturity before depending on it for production CDP operations.
What Is The Difference Between Path A And Path B For CDP Databricks Integration?
Path A uses Databricks as the customer data foundation while an external CDP, reverse ETL tool, or activation platform handles CDP workflows around it. Path B uses CustomerLake to run CDP capabilities natively inside Databricks. Path A is usually better when the organization needs GA readiness, mature marketing self service, complex identity resolution, or multi platform data integration. Path B is worth evaluating when Databricks is already the primary governed data platform and the organization can accept Private Preview risk.
What Is The Medallion Architecture And Why Does It Matter For CDP Programs?
The medallion architecture organizes Databricks data into bronze, silver, and gold layers. Bronze stores raw source data. Silver stores cleaned and standardized records. Gold stores identity resolved customer profiles, segment membership, ML scores, consent attributes, and activation ready tables. CDP tools should read from the gold layer, not bronze, because gold tables represent the governed customer profile that is ready for segmentation, scoring, and activation.
How Does Unity Catalog Support CDP Governance On Databricks?
Unity Catalog supports CDP governance by managing access control, lineage, audit logs, row level security, column level permissions, and governance across Databricks data assets. For CDP programs, Unity Catalog helps control who can access customer profile tables, which service accounts can activate which attributes, how data lineage is audited, and how sensitive data is governed. It applies to both external CDP integrations and CustomerLake embedded architectures.
How Does Structured Streaming Enable Near Real Time CDP Use Cases?
Structured Streaming processes events from Kafka, Kinesis, Azure Event Hubs, and other streaming sources into Databricks Delta Lake tables with seconds to minutes of latency. It supports Tier 2 CDP use cases such as cart abandonment detection, churn signal detection, recent converter suppression, loyalty updates, and near real time segment refresh. For Tier 1 in session use cases that require sub second response, Structured Streaming should feed a hot store rather than serve the live profile request directly.
How Does Reverse ETL Connect Databricks To Marketing Tools?
Reverse ETL reads customer segments, attributes, suppression lists, and ML scores from Databricks gold tables and writes them into downstream tools such as ESPs, CRMs, paid media platforms, customer engagement systems, and packaged CDPs. Tools such as Hightouch, Fivetran Activations, GrowthLoop, and RudderStack Warehouse Actions can support this pattern. Reverse ETL is the activation layer of a Path A CDP Databricks architecture.
What Is Delta Sharing And How Does It Support CDP Analytics?
Delta Sharing allows external consumers to query Databricks Delta tables without creating a physical copy. For CDP analytics, it supports BI access, clean room measurement, partner attribution, and privacy preserving data collaboration. It is useful for analytics and data sharing, but it does not replace reverse ETL because Delta Sharing does not write activation payloads into ESPs, CRMs, or ad platforms.
When Does A CDP Databricks Architecture Need A Hot Store?
A CDP Databricks architecture needs a hot store when the customer experience requires sub second or millisecond profile access. Examples include in session personalization, fraud checks before transaction completion, live AI agent context, cart state decisions, and real time consent enforcement. Databricks remains the source of truth, while Redis, Lakebase, DynamoDB, or another low latency store serves the live Profile API request.
Can Stable Kernel Help With A CDP Databricks Integration Project?
Yes. Stable Kernel designs CDP Databricks integration architectures by starting with the two path decision, assessing native Databricks capabilities, designing the medallion architecture and golden record, defining identity resolution, specifying Mosaic AI scoring, selecting reverse ETL and activation patterns, implementing data contracts, and supporting CustomerLake evaluation when Databricks maturity and roadmap make Path B appropriate.