CDP API Architecture And Integration Patterns
Blog
8/25/26
CDP API Architecture And Integration Patterns
A customer data platform API is not one endpoint.
It is an interoperability layer that connects source systems, customer profiles, segmentation logic, activation tools, personalization engines, and increasingly, AI agents. Treating that layer as a generic API problem is one of the fastest ways to create brittle CDP architecture.
An ingestion endpoint that works during normal traffic may collapse during a campaign spike. A profile endpoint served from the warehouse may time out when a personalization engine needs a response before page render. A bulk audience export may send more PII than the destination needs. A webhook may retry without idempotency and create duplicate downstream actions. An AI agent may query customer profile data at machine speed before the profile API, authorization model, or audit layer is ready.
The architecture question is not, “Does the CDP have an API?”
The better question is, “Which API type does this use case require, which integration pattern fits the data flow, and what latency, throughput, PII, consent, authentication, and failure handling requirements must be designed before production?”
At Stable Kernel, we advise enterprise organizations that CDP API architecture should start with the use case portfolio, not the vendor’s default integration model. Each use case has a direction, a latency tier, a data sensitivity level, and an operational failure mode. The API and integration pattern should follow those requirements.
A mature CDP integration architecture usually includes four API types:
- Ingestion APIs for inbound events and records
- Profile APIs for real time profile reads
- Audience APIs for segment membership and bulk audience access
- Activation APIs for outbound delivery, triggers, and webhooks
It also includes six integration patterns:
- SDK based event streaming
- REST API pull
- Webhook push
- Reverse ETL batch or micro batch
- Native connectors
- MCP for AI agent access
The rest of this guide explains how to design each API type, when to use each integration pattern, and how to handle the cross cutting decisions that determine whether the CDP integration layer can operate safely at enterprise scale.
The Four CDP API Types
A CDP’s integration surface should be organized by function. Each API type exists because it solves a different integration problem. When teams blur these functions together, they often choose the wrong latency target, storage layer, authentication model, or error handling pattern.
Ingestion API
The ingestion API receives customer behavioral events and source system records into the CDP.
It supports web events, mobile app events, server side events, POS transactions, CRM records, loyalty platform records, support events, consent events, and batch imports from source systems.
The ingestion API should be designed for throughput first. At enterprise scale, it should be provisioned for at least three times expected peak volume at 18 month scale, not current average volume. For high volume event streams above roughly 5,000 events per second per instance, Go or Rust is often a better HTTP layer than Node.js or Python because ingestion becomes a throughput and efficiency problem.
The ingestion API should acknowledge receipt quickly, usually within 100 milliseconds, then process events asynchronously. Producers should not wait for downstream identity resolution, segmentation, enrichment, or warehouse writes before receiving an acknowledgment.
Key design decisions include:
- Every event should carry a unique event_id for idempotency
- Tier 1 events should have higher burst limits than lower priority behavioral events
- Streaming sources should route to Kafka, Kinesis, Pub/Sub, or Redpanda
- Batch sources should enter through a batch ingestion path that feeds the same downstream model
- Schema validation should happen at the boundary
- Tier 1 identity field violations should be hard rejected
- Tier 2 completeness violations can be quarantined
- PII filters should reject data that should never enter the CDP, such as credit card numbers or health information
The ingestion API is the front door of the CDP. If it is under provisioned, inconsistent, or weakly governed, every downstream layer inherits the problem.
Profile API
The profile API exposes unified customer profiles for real time lookup.
Personalization engines, customer support tools, POS systems, mobile apps, and AI agents use the profile API to retrieve current customer state at the moment of decision.
For Tier 1 real time use cases, the profile API should serve from a Redis or DynamoDB hot profile store, not directly from the warehouse. Warehouses are excellent for analytics, historical queries, model training, and reporting. They are not designed to sit in the critical path of in session personalization or AI agent execution.
A profile API for real time use cases should target sub 10 millisecond p95 read latency from the hot store. It should support lookup by canonical identifiers such as customer_id, email, loyalty ID, phone number, or another approved identifier, then return the resolved unified profile.
The profile API should also support field selection. The caller should receive only the attributes needed for the use case, not the full customer profile. A support tool may need service history, loyalty status, and contact information. A personalization engine may need current segment membership, recent behavior, suppression status, and consent state. An AI agent may need only the specific fields authorized for its role.
Key design decisions include:
- Serve Tier 1 reads from Redis or DynamoDB
- Keep the full historical profile in the warehouse
- Update the hot store within the Tier 1 event SLA, often 30 seconds
- Support partial responses to reduce payload size and PII exposure
- Provision for AI agent query load if agentic AI is on the 18 to 24 month roadmap
- Use scoped authorization so callers receive only approved profile fields
The profile API is where CDP architecture becomes operational. If it is slow, stale, or overexposed, real time personalization and AI activation will fail in production even if the CDP looks strong in reporting.
Audience API
The audience API exposes segment membership lists and computed audience attributes for bulk export.
It supports use cases such as paid media audience sync, suppression list export, CRM campaign membership updates, email platform list management, and lookalike seed audience export.
Unlike the profile API, the audience API is not optimized for individual low latency lookups. It is optimized for bulk movement of segment data, often on an hourly, daily, or micro batch cadence.
The audience API should support cursor based pagination for large audiences and differential sync so downstream systems receive only records that changed since the last export. Re exporting a full audience every time increases payload size, destination API cost, and processing overhead.
Key design decisions include:
- Align sync cadence with segmentation refresh cadence
- Include segment version and computation timestamp in responses
- Support differential sync for audience entries and exits
- Filter fields by destination to minimize unnecessary data exposure
- Provide bulk export patterns for millions of records
- Avoid using per profile REST calls for large audience movement
Audience API failures often appear as audience size discrepancies. The CDP says an audience contains 500,000 records, but the destination receives fewer because of pagination issues, missing identifiers, destination match constraints, or rate limits.
A well designed audience API makes those differences measurable and recoverable.
Activation API
The activation API delivers customer events, profile changes, and segment changes to downstream operational systems.
It supports cart abandonment triggers, paid media suppression updates, email platform events, CRM lifecycle changes, customer service notifications, AI orchestration triggers, and real time personalization actions.
The activation API usually uses webhooks, event delivery, or scheduled outbound syncs depending on latency needs.
For Tier 1 activation, webhooks are often the correct pattern. If a customer withdraws marketing consent, completes a purchase, abandons a cart, enters a high value segment, or triggers an AI action, the CDP should notify the appropriate destination instead of waiting for the destination to poll.
Key design decisions include:
- Use exactly once delivery for transaction critical events where possible
- Use at least once delivery with idempotency keys for non critical events
- Configure exponential backoff retry for failed deliveries
- Route failed deliveries to a dead letter queue
- Alert when dead letter queue depth exceeds threshold
- Include only the fields required for the destination use case
- Propagate consent context with every activation payload
- Use webhook signature verification to prevent forged events
The activation API is where governance meets execution. It is not enough to send the event. The CDP must send the right fields, in the right format, to the right destination, at the right time, with the right consent state.
The Six CDP Integration Patterns
Once the API type is clear, the integration pattern determines how data moves.
The pattern should match the direction of data flow, the latency tier, the source or destination capability, and the operating burden the enterprise is willing to own.
Pattern 1: SDK Based Event Streaming
SDK based event streaming is the right pattern when digital touchpoints generate continuous customer behavioral events that must reach the CDP within seconds.
This pattern fits:
- Web behavior
- Mobile app events
- Server side event tracking
- POS or kiosk event streams
- Loyalty actions
- Consent events
- Checkout events
Common tools include Snowplow, RudderStack, Segment Track API, and custom server side event APIs. Kafka, Kinesis, Pub/Sub, or Redpanda often sits downstream as the event bus.
Use SDK based event streaming when the source produces high frequency behavioral events and the downstream use case requires Tier 1 or Tier 2 freshness.
Avoid this pattern when the source system naturally produces batch data, such as CRM nightly exports or SFTP vendor files. Streaming a batch source usually adds cost without adding business value.
Pattern 2: REST API Pull
REST API pull is the right pattern when a downstream system needs to retrieve specific customer data on demand.
This pattern fits:
- Profile lookup at page load
- Support sidebar profile enrichment
- POS personalization lookup
- Mobile app profile retrieval
- AI agent profile access when MCP is not available
- Low volume profile or segment queries
The downstream system initiates the request. For example, when a support ticket opens, the support tool queries the CDP profile API for the customer’s unified profile. When a personalization engine loads a page, it queries the profile API for current segment membership and relevant attributes.
Use REST API pull when the downstream system needs current state for one customer or a small set of records.
Avoid REST API pull for bulk audience export. Pulling millions of customer profiles one request at a time is inefficient and expensive. Use the audience API, reverse ETL, or a bulk export pattern instead.
Pattern 3: Webhook Push
Webhook push is the right pattern when the CDP needs to notify a downstream system that something happened.
This pattern fits:
- Cart abandonment triggers
- Consent withdrawal propagation
- Segment entry or exit notifications
- Fraud or risk alerts
- Purchase completion events
- Customer lifecycle changes
- AI action triggers
A webhook reverses the relationship of a standard API call. Instead of the destination asking the CDP for updates, the CDP notifies the destination when a defined event occurs.
Use webhook push when the downstream system should react immediately or within minutes.
Avoid webhook push when the destination cannot reliably receive webhooks, when the use case is bulk audience sync, or when the destination needs to fetch the full current profile at the moment of action. In that case, combine the webhook with a profile API call.
Pattern 4: Reverse ETL Batch Or Micro Batch
Reverse ETL moves processed customer data from the warehouse or CDP profile layer into operational tools on a schedule.
This pattern fits:
- Paid media audience sync
- Suppression list export
- CRM enrichment
- Email platform segment sync
- Customer success attribute updates
- Lookalike seed audience export
Common tools include Hightouch, Fivetran Activations, RudderStack Warehouse Actions, and custom dbt plus scheduled export pipelines.
Use reverse ETL when audience membership and computed attributes need to move from the warehouse to a destination on a scheduled cadence. This is usually correct for Tier 2 and Tier 3 use cases.
Avoid reverse ETL for sub minute activation. A cart abandonment action, consent suppression event, or live AI trigger should not wait for a batch sync when the use case requires immediate action.
Pattern 5: Native Connector
A native connector is the right pattern when the CDP platform offers a certified, maintained connector that matches the integration requirement.
This pattern fits:
- Standard source systems
- Standard destinations
- Common marketing tools
- Common analytics platforms
- Lower customization requirements
- Teams that want to reduce custom maintenance burden
Native connectors can reduce engineering effort significantly. If the connector supports the right schema, latency, retry behavior, consent handling, and field mapping, it may be better than building a custom integration.
Use native connectors when the requirement matches what the connector provides.
Avoid native connectors when the connector introduces unacceptable latency, lacks required schema control, does not support the needed consent propagation, or creates vendor lock in that conflicts with the organization’s architecture strategy.
Pattern 6: MCP For AI Agent Access
MCP, or Model Context Protocol, is the emerging integration pattern for AI agent access to CDP data.
The MCP server sits between the AI agent runtime and the CDP’s profile API or audience API. The agent calls a tool such as get_customer_profile or get_segment_membership. The MCP server routes that call to the CDP, applies authorization, enforces field level access controls, rate limits the request, logs the access, and returns the approved data to the agent.
This pattern fits:
- AI agent customer service workflows
- Agentic personalization
- Autonomous retention workflows
- AI assisted next best action
- AI agents querying customer profiles during live interactions
- AI agents retrieving segment membership or behavioral signals
Use MCP when the AI agent runtime supports the protocol and the enterprise needs a governed access point for customer profile data.
Avoid MCP when the agent runtime does not support it, when the use case is offline model training, or when the MCP server cannot enforce access controls appropriate to the sensitivity of the profile data.
MCP is not just another API wrapper. It is the governance layer between AI agents and customer data.
The MCP Pattern In Depth
MCP matters because AI agents change the access pattern for CDP data.
A human operated personalization system may query profile data at page load. An AI agent may query profiles repeatedly during a live task, ask for related context, retrieve segment memberships, compare scores, and trigger downstream actions. That creates new requirements for access control, latency, rate limiting, and auditability.
How MCP Works In CDP Architecture
In a CDP architecture, the MCP server sits between the AI agent runtime and the CDP APIs.
The MCP server may expose tools such as:
- get_customer_profile
- get_segment_membership
- get_recent_behavioral_signals
- get_consent_status
- trigger_activation
The AI agent does not directly call the CDP’s internal profile API. It calls the MCP tool. The MCP server then routes the request to the appropriate CDP API, checks the agent’s permissions, applies rate limits, filters the response, and logs the access.
This is important because customer data is sensitive. An AI agent should not receive the full profile by default. It should receive the minimum attributes required for the task.
MCP Requires Enterprise Grade Authorization
For production CDP use, MCP access should use OAuth 2.1 with PKCE S256. Token passthrough should not be allowed. The MCP server should issue and manage its own authorization context rather than simply forwarding a caller token to the CDP profile API.
Field level access control is essential.
For example:
- A customer service agent may receive service history, loyalty status, and consent status
- A marketing agent may receive segment membership and campaign eligibility
- A personalization agent may receive recent behavior and product affinity
- A finance adjacent workflow may be blocked from receiving marketing behavior unless explicitly authorized
Each query should be logged with the agent ID, requested tool, parameters, fields returned, timestamp, and purpose. This creates the audit trail needed for governance and compliance review.
When To Build A Custom CDP MCP Server
Organizations should consider a custom MCP server when they are using a composable CDP, custom CDP, or warehouse native architecture that does not already provide governed MCP access.
Building a CDP MCP server usually means:
- Wrapping the existing profile API and audience API
- Defining tool schemas for agent queries
- Implementing OAuth 2.1 authorization
- Configuring field level access scopes
- Applying rate limits for agent speed traffic
- Logging every profile and segment query
- Routing real time requests to the hot profile store, not the warehouse
The rate limit design matters. AI agents can query at machine speed. A profile API sized for human session traffic may not survive agentic AI workloads without additional capacity planning.
Cross Cutting Design Decisions For CDP APIs
The following design decisions apply across every API type and integration pattern. They determine whether the CDP integration layer is secure, reliable, scalable, and usable in production.
PII Handling
PII handling must be designed at each integration point:
- At the ingestion API, the CDP should reject fields that should never enter the platform, such as credit card numbers or health information. Email, phone, IP address, device IDs, and similar identifiers should be flagged as PII and governed accordingly.
- At the profile API, callers should use field selection so they receive only the attributes they need. This matters even more for MCP, where AI agents should receive only approved fields based on their role.
- At the audience API, bulk exports should include only the identifiers and attributes required by the destination. Paid media destinations often need hashed email or hashed phone. They do not need the full customer profile.
- At the activation API, every webhook or outbound payload should be minimized. A cart abandonment email may need email address and cart contents. It does not need service history, full behavioral history, or unrelated profile attributes.
The design rule is simple: send the least sensitive version of the least amount of data that can accomplish the use case.
Consent Propagation
Consent should travel with the event and govern the activation.
Every event entering the CDP should carry consent context, such as analytics consent, marketing consent, functional consent, and any other categories used by the organization’s consent management platform.
The activation API should check consent before delivering any event or audience to a destination. If a customer has not consented to marketing, their event should not flow to an email platform or marketing destination.
Consent changes should propagate in real time. An opt out is a Tier 1 event. If the CDP waits for a nightly reverse ETL sync, the customer may continue receiving marketing messages after withdrawing consent.
Authentication By API Type
Each CDP API type has a different threat model:
- For ingestion APIs, server to server integrations can use API keys when the key is stored securely in a trusted environment. Client side SDK flows should use OAuth 2.0 PKCE because an API key embedded in browser JavaScript or a mobile binary is exposed.
- For profile APIs, use OAuth 2.0 with scoped access tokens. Different callers should have different scopes.
- For audience APIs, scheduled export jobs can use service accounts or API keys. Interactive admin based integrations should use OAuth authorization flows.
- For activation APIs, outbound webhooks should use HMAC SHA 256 signature verification so the destination can confirm the payload came from the CDP.
- For MCP servers, use OAuth 2.1 with PKCE S256 and avoid token passthrough.
Rate Limiting And Error Handling
Rate limiting should be designed separately for each API type.
Ingestion APIs need per source limits with burst capacity for peak events. Promotional campaigns, app releases, and QSR rush periods can create short spikes that normal average based limits will not handle.
Profile APIs need per application limits. AI agent applications need higher capacity planning than human session based applications because agents can query repeatedly during a task.
Audience APIs need destination aware limits for bulk exports. Differential sync reduces pressure by sending only changed records instead of full audiences.
Activation APIs need retry policies, idempotency keys, and dead letter queues. A failed webhook should not disappear silently. It should retry with backoff, fail into a monitored queue, and alert the owner when delivery errors exceed threshold.
The goal is not to prevent every failure. The goal is to make failures safe, observable, and recoverable.
How Stable Kernel Designs CDP API Architecture And Integration Patterns
Stable Kernel designs CDP API architecture across packaged CDPs, composable CDPs, and custom built CDPs without defaulting to a specific vendor implementation.
The architecture starts with the use cases.
API Type And Integration Pattern Audit
Stable Kernel begins by inventorying each planned CDP use case and assigning it to an API type and integration pattern.
Most enterprise CDP programs include several categories:
- High frequency digital event ingestion through SDK based streaming and ingestion APIs
- Real time profile serving through profile APIs backed by Redis or DynamoDB
- Bulk audience synchronization through audience APIs and reverse ETL
- Time sensitive activation through webhooks and activation APIs
AI agent access through MCP where agentic AI is on the roadmap
This audit prevents the common mistake of applying one pattern everywhere. REST polling is not right for high frequency ingestion. Reverse ETL is not right for sub minute activation. A warehouse query is not right for in session profile serving.
CDP Specific API Architecture Standards
Stable Kernel designs ingestion APIs with idempotency keys on every event so retries do not create duplicate processing.
Stable Kernel designs Tier 1 profile APIs to serve from a Redis or DynamoDB hot store, never directly from the warehouse, with capacity planning for AI agent query overhead when applicable.
Stable Kernel designs activation APIs with dead letter queues, exponential backoff retries, idempotency keys, delivery monitoring, and consent propagation.
For MCP based AI agent access, Stable Kernel designs the server as a governed access layer with field level authorization, OAuth 2.1 with PKCE S256, agent specific rate limits, and full audit logging.
Integration Specifications By Source, Consumer, And Destination
The final output is a practical integration specification.
For each source system, Stable Kernel defines the ingestion pattern, authentication model, schema validation requirements, consent context, throughput assumptions, retry behavior, and PII handling.
For each profile consumer, Stable Kernel defines the profile API fields, hot store requirements, access scopes, latency target, and rate limit.
For each destination, Stable Kernel defines the activation or audience pattern, payload fields, PII minimization rules, sync cadence, retry behavior, and delivery monitoring.
Stable Kernel designs CDP API architecture and integration patterns that match the integration mechanism to the use case’s latency requirement, data flow direction, and PII handling constraint, from the ingestion layer through the profile store to every activation destination.
Reflection Questions For Executives
- Are we treating the CDP API as one generic integration layer, or have we separated ingestion, profile, audience, and activation APIs?
- Which use cases require real time profile reads, and are those reads served from a hot store rather than the warehouse?
- Which integrations should use webhooks instead of polling?
- Which audience syncs can run through reverse ETL without harming business outcomes?
- Where does PII enter the CDP, and which destinations actually need it?
- Can consent changes propagate fast enough to stop in flight activation?
- Are our ingestion APIs sized for peak traffic at 18 month scale?
- Do we have an MCP strategy for AI agent access to CDP profiles?
- Are rate limits designed for human traffic, agent traffic, or both?
- Which integration failures are observable, retryable, and assigned to an owner?
FAQ
What Are The Four Types Of CDP APIs?
The four types of CDP APIs are ingestion APIs, profile APIs, audience APIs, and activation APIs. Ingestion APIs receive inbound customer events and source records. Profile APIs expose unified customer profiles for real time reads. Audience APIs expose segment membership and computed attributes for bulk export. Activation APIs deliver events, profile updates, and segment changes to downstream operational systems.
What Are The Main CDP Integration Patterns?
The main CDP integration patterns are SDK based event streaming, REST API pull, webhook push, reverse ETL batch or micro batch, native connectors, and MCP for AI agent access. SDK streaming is used for high frequency inbound events. REST pull is used for on demand profile reads. Webhook push is used for event triggered outbound delivery. Reverse ETL is used for scheduled audience sync. Native connectors reduce custom maintenance when they fit. MCP gives AI agents governed access to CDP data.
When Should A CDP Use A Webhook Instead Of API Polling?
A CDP should use a webhook when the downstream system needs to react to a specific event or state change. Examples include cart abandonment, consent withdrawal, purchase completion, fraud alerts, and segment entry. API polling is better when the downstream system needs to retrieve current state on demand, such as a personalization engine querying a profile at page load.
What Is The MCP Integration Pattern For CDP AI Agent Access?
The MCP integration pattern uses a Model Context Protocol server as the governed access point between AI agents and CDP profile or audience APIs. The agent calls a tool such as get_customer_profile, and the MCP server authenticates the request, applies field level access control, enforces rate limits, routes the query to the CDP, and logs the access for audit.
What Are The Key Design Decisions For A CDP Ingestion API?
A CDP ingestion API should be designed for three times expected peak volume at 18 month scale, sub 100 millisecond acknowledgment, idempotency keys on every event, schema validation at the boundary, per source rate limits, burst capacity for peak events, appropriate authentication, and PII filtering before data enters the CDP pipeline.
Why Should A CDP Profile API Serve From A Hot Store Instead Of The Warehouse?
A CDP profile API should serve Tier 1 real time use cases from a Redis or DynamoDB hot store because warehouse query latency is too slow for in session personalization and AI agent execution. The warehouse should remain the full historical system of record, while the hot store holds the current profile state needed for low latency decisions.
How Does Reverse ETL Fit Into CDP Integration Architecture?
Reverse ETL moves processed customer data from the warehouse or CDP profile layer into operational tools such as ESPs, CRMs, ad platforms, and customer success platforms. It is best for Tier 2 and Tier 3 audience sync use cases that can run on a five minute, hourly, daily, or scheduled cadence. It is not the right pattern for sub minute activation.
How Should A CDP Handle PII Across Integration Patterns?
A CDP should handle PII by filtering sensitive fields at ingestion, minimizing profile API responses through field selection, sending hashed identifiers to paid media destinations when plain text is not required, limiting activation payloads to fields required for the use case, and logging access or export events. Credit card numbers and health information should not enter the CDP pipeline.
What Authentication Model Should Each CDP API Type Use?
Ingestion APIs can use API keys for trusted server to server connections and OAuth 2.0 PKCE for client side flows. Profile APIs should use scoped OAuth access tokens. Audience APIs can use service accounts or OAuth depending on whether the integration is automated or interactive. Activation APIs should use HMAC signed webhooks. MCP servers should use OAuth 2.1 with PKCE S256.
Can Stable Kernel Help Design CDP API Architecture And Integration Patterns?
Yes. Stable Kernel designs CDP API architecture by mapping use cases to API types and integration patterns, specifying throughput and latency targets, designing hot profile store access, defining PII and consent handling, selecting the right pattern for each source and destination, and preparing the architecture for MCP based AI agent access where needed.