How To Build A Real-Time Customer Profile Service

Blog

9/08/26

How To Build A Real-Time Customer Profile Service

A real time customer profile service is a purpose built data layer that maintains a continuously updated, unified view of each customer and makes that view available to downstream systems at sub 100 millisecond to sub 500 millisecond latency, depending on the use case. Unlike a customer data platform, which manages the broader lifecycle of data ingestion, identity resolution, segmentation, and activation, a real time customer profile service is the serving layer. It answers one operational question: who is this customer right now, and what do we know about them?

That distinction matters because many enterprise CDPs were not originally designed to serve customer profile reads inside the live customer experience.

A traditional CDP may be excellent for daily segmentation, lifecycle campaigns, audience activation, attribution, and customer analytics. But when the business needs a personalization engine, fraud model, AI agent, or recommendation system to read customer state during the current interaction, batch profile updates are not enough.

A customer adds an item to cart. A recommendation system needs to know that immediately. A customer signs in from a new device. A fraud model needs current device history and behavior. A customer interacts with an AI agent. The agent needs the latest consent state, lifecycle stage, purchase history, suppression rules, and current intent signals before it recommends the next best action.

If the profile only updates hourly or overnight, the system is not operating on the customer’s current state. It is operating on a delayed representation of the customer.

At Stable Kernel, we advise enterprise teams to treat real time customer profile services as architecture decisions, not feature requests. The question is not only whether the business wants real time personalization. The better question is whether the use case truly requires a sub 100 millisecond or sub 500 millisecond Profile API, and whether the organization has the infrastructure and governance to support that requirement in production.

Why Batch CDPs Cannot Serve Every Real Time Use Case

Batch CDPs are not broken. They are built for a different set of use cases.

A batch CDP updates customer profiles on a scheduled cadence, such as nightly, hourly, or every 15 to 30 minutes. That cadence can work well for weekly campaigns, daily lifecycle journeys, audience exports, and reporting. It does not work as well when the downstream system needs to act during the customer’s current session.

The Batch Profile Freshness Problem

Profile freshness is the gap between what the customer just did and what the customer profile currently reflects.

For campaign segmentation, that gap may not matter. If a weekly email campaign uses yesterday’s loyalty tier, the customer experience is usually not harmed. But for in session personalization, fraud prevention, and AI decisioning, stale profile data can create wrong outcomes.

Common examples include:

  • A product recommendation engine promoting an item the customer just added to cart because the cart event has not reached the profile yet.
  • A fraud model missing a new device or location signal because the profile still reflects last night’s customer state.
  • An AI agent recommending an offer that conflicts with a customer’s current journey, recent conversion, or updated consent state.
  • A personalization engine showing a generic experience because anonymous session behavior has not been linked to the known profile yet.

The issue is not only speed. It is decision quality. When the profile is stale, every downstream system that depends on it becomes less reliable.

The Latency Tier Decision

Not every customer data use case needs the same latency.

Real time bidding and fraud scoring may require a sub 100 millisecond response. In session web personalization and recommendation reranking may tolerate 200 to 500 milliseconds. Push notifications and session close triggers may work on a seconds to minutes cadence. Email campaigns and weekly lifecycle segmentation may only need minutes to hours.

That difference determines the architecture.

A team that builds sub 100 millisecond infrastructure for weekly email segmentation has over engineered the system. A team that uses daily batch profiles for in session personalization has under engineered the customer experience. The right architecture starts by mapping each consuming use case to its latency requirement.

The Five Layer Architecture Of A Real Time Customer Profile Service

A real time customer profile service has five core layers:

  • Streaming event ingestion
  • Stream processing and feature computation
  • Real time identity resolution
  • Hot and cold profile store
  • Profile API and serving layer

Each layer has a specific job. The system only works when all five are designed together.

Layer 1: Streaming Event Ingestion

The streaming event ingestion layer captures customer events as they happen and delivers them into the profile service pipeline.

These events may come from web applications, mobile apps, server side systems, ecommerce platforms, loyalty systems, POS systems, support tools, fraud systems, and consent platforms.

Common tools include Kafka, Confluent Cloud, Redpanda, AWS Kinesis, server side SDKs, mobile SDKs, and change data capture connectors for operational databases.

Key Design Decision: Delivery Semantics And Partitioning

The first ingestion decision is whether events require exactly once or at least once processing.

Financial transactions, identity events, consent updates, and purchase confirmations usually need stronger guarantees because duplicate processing can corrupt customer state. Behavioral events such as product views, clicks, or page views may tolerate at least once delivery when deduplication exists downstream.

The second decision is partitioning. For most profile services, events should be partitioned by customer ID, anonymous ID, or another identity key so events for the same customer are processed in order.

Out of order processing can create incorrect state. If a cart_abandoned event is processed before the add_to_cart event, the profile may show abandonment without the cart context that explains it.

What Can Go Wrong

The ingestion layer fails when it captures events quickly but does not preserve order, validate schemas, deduplicate records, or route events by business value.

A real time pipeline that accepts malformed events simply moves bad data faster. Data contracts at the ingestion boundary are essential because they prevent schema violations from becoming broken profile attributes, incorrect model features, or unusable activation signals.

Layer 2: Stream Processing And Feature Computation

The stream processing layer turns raw events into usable customer profile updates and model features.

A raw event says a customer viewed a product. A computed feature says the customer has viewed three running shoe products in the last 12 minutes, returned to the same category twice, and now has high short term purchase intent.

That transformation is what makes real time personalization useful.

Common tools include Apache Flink, Spark Structured Streaming, RisingWave, and Kafka Streams. Flink is often used for stateful stream processing, session windows, exactly once processing, and low latency joins. Spark Structured Streaming can be useful where the organization already has a Spark and Databricks ecosystem. RisingWave can support streaming SQL and materialized views. Kafka Streams may fit lightweight processing close to Kafka.

Key Design Decision: Stateful Versus Stateless Processing

Some profile updates are simple. A consent status update may only need to replace the current consent value. Other updates require state.

Stateful processing is needed when the system must compare the current event to prior behavior. Examples include session duration, current session category, number of recent views, cart state, recent search pattern, and velocity of account actions.

A session window must also be defined. For web behavior, a 30 minute inactivity gap is often used as a session boundary. The exact window matters because it determines which events count as current intent versus historical behavior.

What Can Go Wrong

The biggest risk is training and serving skew.

If offline model features are computed one way in the warehouse and online features are computed differently in the streaming layer, the model may behave well in testing and poorly in production. The customer’s real time feature values no longer match the feature definitions used during model training.

The profile service should define features once, document their window logic, and validate that offline and online calculations remain consistent.

Layer 3: Real Time Identity Resolution

Real time identity resolution determines whether incoming events belong to an existing customer profile, a new anonymous profile, or a known profile connected through a shared identifier.

This is one of the hardest parts of the architecture.

A customer may browse anonymously on mobile web, log in through the app, click an email on desktop, and purchase in store using a loyalty account. Without identity resolution, those events become separate profiles. With identity resolution, they become one customer journey.

Deterministic Matching

Deterministic matching links profiles using exact shared identifiers.

Common identifiers include:

  • Email address
  • Phone number
  • Loyalty ID
  • Account ID
  • Customer ID
  • Hashed device ID
  • Authenticated app ID

Deterministic matching is precise and usually the right foundation for regulated industries or high stakes customer experiences. The tradeoff is coverage. It only links identities when a reliable shared identifier exists.

Probabilistic Matching

Probabilistic matching links profiles using statistical similarity.

It may use device signals, browser attributes, IP patterns, behavioral similarity, location signals, session timing, or other indirect evidence. This can improve match coverage for anonymous users, but it introduces false positive risk.

A false positive merge can be worse than no merge at all. It may combine two people’s profiles, expose irrelevant personalization, distort fraud decisions, or create compliance concerns.

For most enterprises, the best pattern is deterministic foundation with probabilistic enrichment. Deterministic matches can be treated as durable. Probabilistic matches should be scored, time bounded, monitored, and reversible.

Merge Rule Design

The merge rule defines what happens when two profile nodes are linked.

A golden record merge combines attributes into one canonical customer profile. This is operationally simple, but the system must define survivorship rules. Which email wins? Which loyalty status wins? Which address, consent state, or customer tier becomes authoritative?

A linked graph model keeps profiles as separate nodes with relationships between them. This is safer when merges may need to be reversed, but live graph traversal can add latency. For sub 100 millisecond Profile API targets, merged views should usually be pre computed and cached in the hot store rather than assembled through graph traversal during the request.

Layer 4: Hot And Cold Profile Store

The hot and cold profile store split determines where customer profile data lives based on latency requirements.

The hot store serves the live Profile API. It is optimized for fast reads. The cold store holds the full historical profile and supports analytics, model training, compliance, and batch segmentation.

Common hot store options include Redis, DynamoDB, Cassandra, and similar low latency serving systems. Common cold store options include Snowflake, Databricks, BigQuery, and other warehouse or lakehouse platforms.

What Belongs In The Hot Store

The hot store should contain only the attributes needed for real time serving.

That usually includes:

  • Current consent status
  • Active suppression flags
  • Current lifecycle stage
  • Active segment membership flags
  • Loyalty tier and active rewards
  • Pre-computed ML scores
  • Recent session behavior
  • Current cart state
  • Next best action output
  • Recent identity resolution result

The practical test is simple: does a consuming system need this attribute during the current customer interaction? If yes, it belongs in the hot store. If not, it likely belongs in the cold store.

What Belongs In The Cold Store

The cold store should contain the complete customer history.

That includes full event logs, historical transactions, long term behavior, model training datasets, attribution history, compliance archives, and profile attributes that do not need millisecond access.

The cold store is also where heavier analytical queries belong. A dashboard analyzing 90 days of behavior should not query Redis. A machine learning training job should not depend on the hot serving layer.

What Can Go Wrong

The most common failure is hot store over provisioning.

Teams add attributes to the hot store because they may be useful. Over time, the hot store becomes a second warehouse at a much higher cost profile. Memory utilization rises, evictions begin, and p95 or p99 latency starts to degrade.

The solution is lifecycle governance. Session events should have TTL rules. Aging attributes should move to cold storage. Hot store contents should be audited against actual Profile API query patterns.

Layer 5: Profile API And Serving Layer

The Profile API is the interface that downstream systems use to retrieve current customer state.

It may serve personalization engines, fraud models, AI agents, recommendation systems, customer service tools, experimentation platforms, mobile apps, web applications, and activation systems.

Common implementation patterns include a thin API service in Go or Node.js, REST for high volume point lookups, GraphQL when consumers need flexible attribute selection, and gRPC for low latency internal service to service calls.

Designing The Latency Budget

The Profile API latency budget should be assigned from the use case SLA.

  • For a sub 100 millisecond end to end use case, the Profile API itself may only have 10 to 20 milliseconds. The rest of the time is consumed by DNS, TLS, API gateway overhead, service processing, serialization, network hops, and downstream decisioning.
  • For a 200 to 500 millisecond personalization use case, the Profile API may have 20 to 50 milliseconds. For push triggers or session close events, 100 to 500 milliseconds may be acceptable. For email segmentation, seconds may be fine.

The key rule is to measure p95 and p99, not just average latency. A system with a 50 millisecond average and a 5 second p99 is not reliable for real time use cases. At high request volume, even 1 percent of slow requests can represent thousands of broken customer experiences per day.

What Can Go Wrong

The Profile API often works in testing and fails under load.

A service that meets latency targets at 1,000 requests per second may miss them at 10,000 requests per second if connection pools, hot store capacity, serialization overhead, or API gateway limits are not tested.

Before production, the Profile API should be load tested at 2 to 3x expected peak traffic. Circuit breakers should return a degraded profile, cached response, or safe default when the hot store slows down, instead of allowing the entire request path to collapse.

Build Versus Buy: When A Custom Profile Service Is Justified

A custom real time customer profile service is not always the right answer.

It is powerful, but it introduces real operating responsibility. The organization must maintain streaming ingestion, stateful processing, identity resolution, a hot store, Profile API infrastructure, schema governance, latency monitoring, and incident response.

When A Custom Build Makes Sense

A custom build is justified when three conditions are true.

  • First, the business has one or more use cases with a sub 500 millisecond Profile API requirement that packaged or composable tools cannot meet reliably.
  • Second, the engineering team has the maturity to operate streaming infrastructure, identity graph logic, and low latency serving systems in production.
  • Third, the organization’s data volume, customization needs, or pricing model makes a packaged CDP less attractive than owned infrastructure.

Custom builds make the most sense when the profile service is a strategic product capability, not only a marketing data utility.

When A Composable CDP Is The Better Answer

A composable CDP or warehouse native approach is often better when the most demanding use case works on a 5 to 60 minute cadence.

Triggered email, lifecycle segmentation, daily suppression, churn scoring, LTV modeling, and campaign activation often do not need sub 100 millisecond serving. In those cases, the organization may be better served by Snowflake or Databricks as the customer data foundation, dbt for identity and transformation, reverse ETL for activation, and Redis or DynamoDB only for the small number of use cases that truly need low latency reads.

The decision should not be framed as modern versus outdated. It should be framed as latency requirement versus operating cost.

The Three Question Decision Framework

Before building a custom real time customer profile service, leaders should ask:

  • What is the latency requirement of the most demanding consuming use case?
  • Does the engineering team have the operational capacity to run the stack reliably?
  • At what data volume or licensing level does the custom build become more economical than packaged or composable alternatives?

If the answer to the first question is minutes rather than milliseconds, the custom build may be unnecessary.

How Stable Kernel Approaches Real Time Customer Profile Service Design

Stable Kernel designs real time customer profile services from the latency budget outward.

The engagement begins by mapping every consuming use case to its required latency tier. That includes personalization engines, AI agents, fraud systems, mobile apps, web experiences, service tools, and activation systems.

Phase 1: Architecture And Latency Budget Design

Stable Kernel defines the five layer stack for the client’s cloud environment, data volume, governance requirements, and use case roadmap.

That includes the streaming ingestion pattern, stream processor, identity graph design, merge rule, hot and cold profile store split, and Profile API budget.

Phase 2: Build And Integration

Stable Kernel builds or integrates the selected components, including Kafka or Redpanda, Flink or equivalent stream processing, identity resolution logic, hot store lifecycle policies, and the Profile API serving layer.

The goal is not only to connect the stack. It is to validate that the stack meets the required p95 and p99 latency under expected production load.

Phase 3: Governance And Operational Handoff

Stable Kernel also designs the governance needed to keep the service reliable after launch.

That includes data contracts at ingestion, hot store attribute governance, identity match monitoring, Profile API circuit breakers, latency dashboards, p99 alerts, and operational ownership.

Stable Kernel does not sell a CDP platform. The recommendation is vendor agnostic. For some clients, that means a full custom profile service. For others, it means a composable CDP architecture with a targeted hot serving layer. The right answer is the architecture that matches latency, scale, cost, and operational capacity.

FAQ

What Is A Real Time Customer Profile Service?

A real time customer profile service is a purpose built data layer that maintains a continuously updated, unified view of each customer and makes that profile available to downstream systems at low latency. It is used by personalization engines, AI agents, fraud systems, recommendation tools, and activation platforms that need to know who the customer is right now and what context should influence the next decision.

What Are The Five Layers Of A Real Time Customer Profile Service?

The five layers are streaming event ingestion, stream processing and feature computation, real time identity resolution, hot and cold profile storage, and the Profile API serving layer. Streaming ingestion captures events as they happen. Stream processing turns those events into usable features. Identity resolution links events to the right customer. The hot store serves current profile state quickly, while the cold store preserves full history. The Profile API exposes the current profile to downstream systems.

How Is A Real Time Customer Profile Service Different From A CDP?

A CDP manages the broader lifecycle of customer data, including ingestion, identity resolution, segmentation, activation, analytics, and governance. A real time customer profile service is a narrower architectural layer focused on serving current customer state at low latency. A packaged CDP may include this capability, but enterprises sometimes build a custom service when they need lower latency, more control, or more specialized profile serving than their CDP provides.

What Latency Should A Profile API Target?

The correct latency target depends on the use case. Fraud scoring and real time bidding may need sub 100 millisecond p95 response. In session web personalization may tolerate 200 to 500 milliseconds. Push triggers may work on a seconds to minutes cadence. Email campaigns and lifecycle segmentation may tolerate minutes or hours. Teams should define the use case SLA before selecting infrastructure.

Why Does Identity Resolution Matter For Real Time Profiles?

Identity resolution determines whether events from web, mobile, store, CRM, loyalty, and support systems are linked to the same customer. Without it, the profile service may serve fragmented profiles that miss key behaviors. Real time identity resolution matters because downstream systems need the unified profile during the current interaction, not after the next nightly identity job.

What Belongs In The Hot Profile Store?

The hot profile store should contain only data needed for real time serving. That includes current consent status, active suppression flags, lifecycle stage, active segment membership, loyalty state, recent session behavior, current cart state, pre computed ML scores, and next best action outputs. Full event history, analytics data, model training data, and compliance archives belong in the cold store.

When Should You Build A Custom Real Time Customer Profile Service?

A custom build is justified when the business has a sub 500 millisecond Profile API requirement, packaged or composable tools cannot meet the requirement, the engineering team can operate the streaming and serving stack, and the data volume or customization need justifies owned infrastructure. If the most demanding use case works on a minutes based cadence, a composable CDP or warehouse native architecture may be a better fit.

What Tools Are Commonly Used To Build A Real Time Customer Profile Service?

Common tools include Kafka, Confluent Cloud, Redpanda, or AWS Kinesis for streaming ingestion; Apache Flink, Spark Structured Streaming, RisingWave, or Kafka Streams for processing; Neo4j, Amazon Neptune, PostgreSQL, dbt, or custom logic for identity resolution; Redis, DynamoDB, or Cassandra for hot storage; Snowflake, Databricks, or BigQuery for cold storage; and Go, Node.js, REST, GraphQL, or gRPC for the Profile API layer.

Can Stable Kernel Help Build A Real Time Customer Profile Service?

Yes. Stable Kernel helps enterprise teams design and build real time customer profile services by mapping use cases to latency tiers, designing the five layer architecture, selecting the right tools, implementing identity resolution, defining the hot and cold profile store split, building the Profile API, and establishing the governance and observability needed to keep the service reliable in production.