Services

Resources

Company

How temporal versioning prevents configuration drift in stream processing

How temporal versioning prevents configuration drift in stream processing

How temporal versioning prevents configuration drift in stream processing

How temporal versioning prevents configuration drift in stream processing

An administrator changed a policy engine rule in a multi-tenant stream-processing system we were building with a client at One2N. The change propagated correctly, and services started using it within seconds.

That was exactly the problem.

Events processed before 2 PM used the older policy while events processed after 2 PM used the new one, even when both belonged to the same logical processing window. Nothing failed operationally, yet the result of processing had started depending on when an event happened to execute. Every service was up to date, yet the output was inconsistent. That is configuration drift.

This article covers how we addressed that by introducing effective-time configuration, window-level versioning and explicit configuration resolution, and what that changed across CDC, MongoDB, service boundaries, rollover, recovery and replay.

How mutable configuration broke deterministic stream processing

The platform processed millions of messages per day across multiple tenants. Each tenant had its own policies, quotas, processing rules and local operating windows. Those configuration values were part of the actual processing logic, so two services handling the same event needed to agree on which policy applied.

The original design was straightforward: an administrator updates a policy, the configuration changes, and services read the latest configuration. When an engineer changed a policy value, services picked it up immediately and processed incoming messages with the new ruleset. For a long time, that model was reasonable.

The gotcha was a function of its intended behaviour. Consider a tenant whose processing window runs for one logical day.

Timeline of one processing window where events at 13:59 and 14:01 are evaluated with different policy versions after a 14:00 configuration change
Fig 1: A 14:00 policy change splits one logical processing window in two. Event A at 13:59 is evaluated with v17, Event B at 14:01 with v18, and both lookups are “correct”.

At 1:59 PM, an event may be evaluated with policy v17. At 2:01 PM, another event belonging to the same processing window may be evaluated with v18.

Every individual configuration lookup succeeds. Neither service has read stale data. Yet the processing window no longer has a stable set of rules, and that produced consequences beyond the immediate policy decision.

During incident investigation, explaining why an event received a particular result meant reconstructing configuration state around its processing timestamp. Replaying the same event later could produce a different result because the latest policy had changed. Two events that logically belonged together could also be evaluated under different assumptions.

The architecture had allowed processing time to influence policy selection implicitly.

Temporal versioning: separating capture time from effective time

We needed one property to hold throughout a processing window:

Work associated with the same processing window must resolve to that window’s configuration version.

That invariant sounds simple, but it changes the meaning of “current configuration”. Suppose an administrator updates a quota at 2 PM. There are two relevant moments in that change:

  1. When the system learns about the new value.

  2. When the new value becomes valid for event processing.

Those moments had previously been identical. Pressing Save effectively meant “make this policy active now”. The idea is close to what Martin Fowler describes as bitemporal history, where record time and actual (effective) time are separate axes.

We separated them. Outside rollover, the curator forwarded configuration changes promptly with a nextWindowToken. Downstream consumers used that token to update a future versioned collection, while work carrying the current window’s token continued to read its existing collection. Only during the pre-rollover freeze did the curator temporarily buffer incoming changes in MongoDB.

Diagram separating configuration capture time from effective time: a 14:00 change is captured by CDC, written to the next window's collection, and applied only at the window boundary
Fig 2: The 14:00 change is captured within seconds and written to the W+1 collection. Work carrying W’s token keeps reading version W until the boundary.

The system could still capture and propagate changes quickly. The current processing window continued using its associated policy version. A future window could incorporate the new configuration, subject to the pre-rollover cutoff described below.

That gave configuration an effective time. Fast propagation no longer meant immediate activation for current-window work. MongoDB staging only held changes during the rollover freeze.

Using window tokens to version processing windows

Once configuration could have different authoring and activation times, we needed a stable way to identify which version belonged to which work.

We used a version identifier for each tenant processing window. Internally, we called it a window token. The token is an opaque identifier, not necessarily a sequential number. W, W+1 and W+2 below are labels for consecutive windows, not literal token values.

Diagram mapping a tenant's processing windows to opaque window tokens and versioned configuration collections, with a work event resolving configuration through its token
Fig 3: Each tenant window maps to an opaque token, and each token identifies one configuration version. A work event carrying W’s token always resolves to version W.

Services depend on the association, so the token can be any opaque value. A token identifies a configuration version; it is not itself the frozen snapshot. If an event carries the token for window W, configuration resolution leads to the version associated with W.

A policy change while W is active does not modify the collection selected by work carrying W’s token. During normal operation it is forwarded to the version identified by the next token; during the pre-rollover freeze it is buffered for replay to a further future version.

This moved configuration selection away from a wall-clock lookup and towards an explicit processing relationship:

# Before: the clock decides
event -> configuration available right now

# After: the event's processing identity decides
event -> original window token -

Resolving temporal configuration across tenant time zones

The platform was multi-tenant and timezone-aware. One tenant could cross into its next logical processing day while another still had several hours remaining in its current one.

We did not want individual services implementing their own interpretation of a tenant’s active window. Even small differences in timezone conversion or rollover logic could put two services on different configuration versions.

Tenant timezone determines when the curator schedules that tenant’s freeze and rollover. It should not be necessary for every downstream service to recalculate an event’s window from its own wall clock. Instead, the work carries a windowToken, and a downstream service uses that token to resolve the corresponding versioned collection. The collection-resolution logic belongs in the shared configuration path rather than in each business query.

Wall-clock time still matters. It helps determine when a tenant moves into its next processing window. Once an event has a processing identity, however, downstream services should not keep asking the clock which configuration to use. They can resolve it from the token carried with that work. That becomes especially important when the event is processed again hours or days later.

Separating the configuration control plane from the stream-processing data plane

Versioning configuration introduced lifecycle operations that did not belong inside every streaming service. Something had to consume updates, route them to future versions, buffer changes around rollover and coordinate the creation of the next version. We separated those responsibilities from event processing.

Architecture diagram of the configuration control plane (Debezium, Kafka, config curator, quota manager, MongoDB stage store) and the stream-processing data plane of token-aware services
Fig 4: The config curator owns the configuration lifecycle. Changes are forwarded to a future version during normal operation, and the MongoDB stage store is only used during the rollover freeze.

The config curator formed the control plane. It owned the lifecycle of configuration: consuming changes, tracking tenant-local rollover schedules, routing ordinary updates to a future version, buffering updates during the freeze and coordinating bootstrap.

Streaming services formed the data plane. They used the token on incoming work to select a configuration version, and their config-sync consumers wrote curator updates to the version specified by nextWindowToken. The data plane therefore had both token-directed configuration writes and token-directed processing reads.

That separation gave us a useful architectural rule:

The component processing an event should not also decide when a new policy becomes valid.

Keeping those responsibilities apart made the consistency model much easier to reason about. We wrote more about asking these questions before drawing components in why your architecture should start with questions, not boxes.

Using CDC for configuration capture without immediate activation

Change Data Capture with Debezium and Kafka was already useful because configuration updates needed to move quickly through the platform. An administrator might update a policy at 2 PM, and we wanted the control plane to learn about that update promptly. Capture latency is worth monitoring on its own. One parameter once stalled our Postgres replication with Debezium for three weeks.

After CDC observed an update:

  1. CDC captures the policy change.

  2. Outside a freeze, the curator forwards it with a nextWindowToken header.

  3. A downstream consumer uses that header to update the future versioned collection.

  4. Work for the applicable future window uses that version. Work carrying the current window token keeps using its own.

There is a separate branch during the pre-rollover freeze. Instead of forwarding immediately, the curator stores the incoming message in a per-tenant MongoDB stage store. After rollover, a freeze-ended event triggers replay of those messages with the new next-window token.

Capture latency is how quickly the control plane knows about a change. The effective window is which work is allowed to use that change. Neither is the same thing as the temporary buffer used during rollover.

Moving safely from window W to W+1

Keeping current-window processing on a stable version solves the mid-window update problem, but it still leaves a sensitive transition at the tenant’s next boundary.

Before that boundary, work associated with W reads version W, while normal configuration updates are forwarded to the collection for W+1. The curator stores both a current and a next token. In its bootstrap logic, the previously stored next token becomes the new current token.

The tenant’s freeze is scheduled ahead of the boundary; the checked-in default offset is 60 minutes. Once frozen, incoming changes go to MongoDB stage storage rather than the prepared W+1 collection. At rollover, the curator:

  1. Obtains another token from the quota manager.

  2. Calls downstream services to clone W+1 into W+2.

  3. Rotates its current/next token pair.

  4. Ends the freeze. A freeze-ended event then replays staged changes with W+2 as their next-window token, and normal forwarding resumes towards W+2.

Tenant rollover lifecycle from window W to W+1: normal operation, pre-rollover freeze, rollover with clone and token rotation, and freeze end with replay of buffered changes
Fig 5: The freeze starts 60 minutes before the boundary. At rollover the token pair rotates and buffered changes replay into the newly created W+2.

These are scheduled tenant-local transitions, not a guarantee that every service switches at precisely the same millisecond. The scheduler periodically checks for due jobs. Nor does “tomorrow” apply to every update made before midnight: a change received during the pre-rollover freeze is replayed to the newly created future version rather than inserted into the version being finalized. That cutoff is part of the effective-time contract.

The curator asks configured downstream services to clone the prepared version into the newly assigned future version using an endpoint of this shape:

POST /v1/tenants/{tenantName}/bootstrap-config
Content-Type: application/json

{
  "currentWindowToken": "token-for-W-plus-1",
  "nextWindowToken": "token-for-W-plus-2"

Here currentWindowToken names the clone source and nextWindowToken names its target. In the curator’s stored token pair before rollover, that source was called the next token. The names describe the bootstrap transition, not two different clocks.

The curator retries failed calls, but a retry should not be described as an unconditional no-op. Some downstream implementations drop an existing target collection before cloning it again. Safe retry behaviour depends on ordering and on whether anything has written to that target in the meantime.

Dynamic MongoDB collections for versioned configuration

With ordinary configuration storage, an entity can map to a stable MongoDB collection and a conventional repository abstraction works well. After versioning, the target collection depended on runtime context. For example, a downstream service with a surveillance-policies base collection selects a versioned collection of the form:

The tenant determines the database, the configuration type determines the base collection and the token determines the version suffix. The initial bootstrap token is a special case that uses the unsuffixed base collection.

Diagram of token-aware MongoDB collection resolution: tenant selects the database, config type the base collection and window token the version suffix, via MongoTemplate
Fig 6: Tenant, configuration type and window token together resolve the physical collection, which is then queried through MongoTemplate.

That made a statically bound MongoRepository less natural for the versioned access path. Downstream services used token-aware collection resolution and MongoTemplate where choosing the correct physical configuration version was part of processing correctness.

We were not trying to replace Spring Data repositories across the application. The choice only concerned the paths where the target collection changes with the processing context. We have also written about how enabling Spring Kafka batch listeners backfired in an event pipeline.

A query against the wrong collection could still return valid-looking configuration data. It would be valid for the wrong window.

Recovering configuration rollovers through reconciliation

The rollover spans Kafka, the curator, HTTP calls, MongoDB and multiple pieces of persisted state. Failures can therefore occur halfway through a transition:

  • The curator may restart after some downstream collections have been cloned but before the transition finishes.

  • An HTTP request may time out after the receiving service has already performed the clone.

  • Previously consumed events may appear again during recovery.

The curator uses retries and outbox events to help recover from partial failures, but this is not one atomic transaction across Kafka, MongoDB and every downstream service. It rotates tokens and ends the freeze during bootstrap; staged-message replay follows a freeze-ended event. The replay service republishes staged messages to Kafka and then deletes them from the Mongo stage store, which does not make the two systems one transaction. “Bootstrap completed” does not, by itself, prove that every staged change has been applied downstream.

That changes the questions we ask in monitoring. A successful scheduled rollover job says little by itself. Monitoring needs to show whether the expected downstream versions are available and the buffered changes have drained:

  • Has each downstream bootstrap finished?

  • Is the tenant still frozen?

  • Are staged messages draining?

  • Which current and next tokens are recorded?

  • Are downstream versions available and consistent with those tokens?

These are operational checks to build and validate, not a claim that a single reconciliation = running field or atomic activation gate already answers them.

Temporal versioning for deterministic event replay

The strongest benefit appeared when we looked at event replay. Without versioning, an unchanged event combined with the latest configuration at replay time can produce a different policy decision. After window versioning, the window token became part of the event’s processing context:

Event M + Tenant A + original window token W + policy version for

If the same event needs to be investigated later, the system can use the original token to resolve the same policy version again, provided the event retains that token and the historical collection still exists.

Comparison of event replay against latest configuration versus replay resolving the retained configuration version through the original window token
Fig 7: At replay time, the event’s token selects its configuration collection. Old versions have to be retained for this to work.

Temporal versioning makes the configuration input reproducible. It does not, by itself, guarantee the same overall result if application code, external data or other inputs have changed. Retention of old configuration versions is part of the replay contract, not just storage housekeeping.

That helped incident investigation because an engineer could identify the policy version associated with the event. It helped audits for the same reason. It also meant a configuration update no longer silently changed the meaning of work carrying a current-window token.

The window token had effectively made policy version part of processing identity.

Operational trade-offs of temporal configuration versioning

We would not use this pattern for every configuration value. It has real costs:

  • Versioned snapshots require storage, and old versions need a retention policy.

  • Rollovers need orchestration, and failures during transition need recovery.

  • Freeze-time updates need durable staging.

  • Tenant-local window boundaries add scheduling and operational behaviour.

Our implementation staged freeze-time changes in MongoDB. If we were designing it again, we would evaluate whether keeping that state in a compacted Kafka topic would simplify recovery and keep more of the configuration lifecycle inside the event pipeline. Changing the buffer alone would not solve cross-system cutover or idempotency. For governing Kafka itself as code, see how we ran GitOps for 78 Kafka clusters.

There are also cutoff timings to consider. The pre-rollover freeze offset gives the prepared version a stable cutoff before the boundary; in the checked-in configuration it is 60 minutes. Changes received during the freeze are buffered and replayed to a later version. We would evaluate the offset against observed bootstrap and replay behaviour rather than assume one duration is ideal for every tenant.

Those costs are justified when configuration changes the semantics of work. Quotas, compliance policies, entitlement decisions, routing rules, transformations and scheduled processing rules are good examples.

When temporal configuration versioning is the right consistency model

Two practical tests help decide:

  1. If this event is replayed tomorrow, does it need the policy version associated with its original processing window?

  2. Can two events belonging to the same logical window safely observe different versions of the policy?

When either answer creates a correctness problem, an unversioned “latest configuration” lookup deserves closer scrutiny.

Looking back, the original system already had an effective-time model. It was implicit. The moment an administrator saved a change became the moment that configuration affected processing. As long as policy changes did not interfere with work already underway, that behaviour was easy to live with. The 2 PM update exposed where the assumption stopped holding.

The principle we carried forward:

When configuration participates in the result of processing, its effective time belongs in the data model.

FAQ: temporal configuration versioning

What is configuration drift in stream processing?

It is when events that belong to the same logical processing window are evaluated under different configuration versions, usually because services always read the latest configuration and a change lands mid-window. Every lookup is correct, but results depend on when an event happened to execute.

What is effective-time configuration?

A model where each configuration change records both when it was made and which work it applies to. Changes are captured immediately and written to a future version, which only becomes active at a defined window boundary.

How does a window token make event replay deterministic?

Each event carries the opaque token of the processing window it belongs to. Services resolve configuration from that token rather than from the clock, so replaying the event later selects the same configuration version, as long as that version is retained.

Should every configuration value be versioned this way?

No. Versioning adds storage, rollover orchestration and recovery work. Reserve it for configuration that changes the semantics of processing, such as quotas, compliance policies, entitlements and routing rules.

If you are running into the same problems with Kafka pipelines, CDC or multi-tenant consistency, take a look at our backend engineering services or browse our case studies.

An administrator changed a policy engine rule in a multi-tenant stream-processing system we were building with a client at One2N. The change propagated correctly, and services started using it within seconds.

That was exactly the problem.

Events processed before 2 PM used the older policy while events processed after 2 PM used the new one, even when both belonged to the same logical processing window. Nothing failed operationally, yet the result of processing had started depending on when an event happened to execute. Every service was up to date, yet the output was inconsistent. That is configuration drift.

This article covers how we addressed that by introducing effective-time configuration, window-level versioning and explicit configuration resolution, and what that changed across CDC, MongoDB, service boundaries, rollover, recovery and replay.

How mutable configuration broke deterministic stream processing

The platform processed millions of messages per day across multiple tenants. Each tenant had its own policies, quotas, processing rules and local operating windows. Those configuration values were part of the actual processing logic, so two services handling the same event needed to agree on which policy applied.

The original design was straightforward: an administrator updates a policy, the configuration changes, and services read the latest configuration. When an engineer changed a policy value, services picked it up immediately and processed incoming messages with the new ruleset. For a long time, that model was reasonable.

The gotcha was a function of its intended behaviour. Consider a tenant whose processing window runs for one logical day.

Timeline of one processing window where events at 13:59 and 14:01 are evaluated with different policy versions after a 14:00 configuration change
Fig 1: A 14:00 policy change splits one logical processing window in two. Event A at 13:59 is evaluated with v17, Event B at 14:01 with v18, and both lookups are “correct”.

At 1:59 PM, an event may be evaluated with policy v17. At 2:01 PM, another event belonging to the same processing window may be evaluated with v18.

Every individual configuration lookup succeeds. Neither service has read stale data. Yet the processing window no longer has a stable set of rules, and that produced consequences beyond the immediate policy decision.

During incident investigation, explaining why an event received a particular result meant reconstructing configuration state around its processing timestamp. Replaying the same event later could produce a different result because the latest policy had changed. Two events that logically belonged together could also be evaluated under different assumptions.

The architecture had allowed processing time to influence policy selection implicitly.

Temporal versioning: separating capture time from effective time

We needed one property to hold throughout a processing window:

Work associated with the same processing window must resolve to that window’s configuration version.

That invariant sounds simple, but it changes the meaning of “current configuration”. Suppose an administrator updates a quota at 2 PM. There are two relevant moments in that change:

  1. When the system learns about the new value.

  2. When the new value becomes valid for event processing.

Those moments had previously been identical. Pressing Save effectively meant “make this policy active now”. The idea is close to what Martin Fowler describes as bitemporal history, where record time and actual (effective) time are separate axes.

We separated them. Outside rollover, the curator forwarded configuration changes promptly with a nextWindowToken. Downstream consumers used that token to update a future versioned collection, while work carrying the current window’s token continued to read its existing collection. Only during the pre-rollover freeze did the curator temporarily buffer incoming changes in MongoDB.

Diagram separating configuration capture time from effective time: a 14:00 change is captured by CDC, written to the next window's collection, and applied only at the window boundary
Fig 2: The 14:00 change is captured within seconds and written to the W+1 collection. Work carrying W’s token keeps reading version W until the boundary.

The system could still capture and propagate changes quickly. The current processing window continued using its associated policy version. A future window could incorporate the new configuration, subject to the pre-rollover cutoff described below.

That gave configuration an effective time. Fast propagation no longer meant immediate activation for current-window work. MongoDB staging only held changes during the rollover freeze.

Using window tokens to version processing windows

Once configuration could have different authoring and activation times, we needed a stable way to identify which version belonged to which work.

We used a version identifier for each tenant processing window. Internally, we called it a window token. The token is an opaque identifier, not necessarily a sequential number. W, W+1 and W+2 below are labels for consecutive windows, not literal token values.

Diagram mapping a tenant's processing windows to opaque window tokens and versioned configuration collections, with a work event resolving configuration through its token
Fig 3: Each tenant window maps to an opaque token, and each token identifies one configuration version. A work event carrying W’s token always resolves to version W.

Services depend on the association, so the token can be any opaque value. A token identifies a configuration version; it is not itself the frozen snapshot. If an event carries the token for window W, configuration resolution leads to the version associated with W.

A policy change while W is active does not modify the collection selected by work carrying W’s token. During normal operation it is forwarded to the version identified by the next token; during the pre-rollover freeze it is buffered for replay to a further future version.

This moved configuration selection away from a wall-clock lookup and towards an explicit processing relationship:

# Before: the clock decides
event -> configuration available right now

# After: the event's processing identity decides
event -> original window token -

Resolving temporal configuration across tenant time zones

The platform was multi-tenant and timezone-aware. One tenant could cross into its next logical processing day while another still had several hours remaining in its current one.

We did not want individual services implementing their own interpretation of a tenant’s active window. Even small differences in timezone conversion or rollover logic could put two services on different configuration versions.

Tenant timezone determines when the curator schedules that tenant’s freeze and rollover. It should not be necessary for every downstream service to recalculate an event’s window from its own wall clock. Instead, the work carries a windowToken, and a downstream service uses that token to resolve the corresponding versioned collection. The collection-resolution logic belongs in the shared configuration path rather than in each business query.

Wall-clock time still matters. It helps determine when a tenant moves into its next processing window. Once an event has a processing identity, however, downstream services should not keep asking the clock which configuration to use. They can resolve it from the token carried with that work. That becomes especially important when the event is processed again hours or days later.

Separating the configuration control plane from the stream-processing data plane

Versioning configuration introduced lifecycle operations that did not belong inside every streaming service. Something had to consume updates, route them to future versions, buffer changes around rollover and coordinate the creation of the next version. We separated those responsibilities from event processing.

Architecture diagram of the configuration control plane (Debezium, Kafka, config curator, quota manager, MongoDB stage store) and the stream-processing data plane of token-aware services
Fig 4: The config curator owns the configuration lifecycle. Changes are forwarded to a future version during normal operation, and the MongoDB stage store is only used during the rollover freeze.

The config curator formed the control plane. It owned the lifecycle of configuration: consuming changes, tracking tenant-local rollover schedules, routing ordinary updates to a future version, buffering updates during the freeze and coordinating bootstrap.

Streaming services formed the data plane. They used the token on incoming work to select a configuration version, and their config-sync consumers wrote curator updates to the version specified by nextWindowToken. The data plane therefore had both token-directed configuration writes and token-directed processing reads.

That separation gave us a useful architectural rule:

The component processing an event should not also decide when a new policy becomes valid.

Keeping those responsibilities apart made the consistency model much easier to reason about. We wrote more about asking these questions before drawing components in why your architecture should start with questions, not boxes.

Using CDC for configuration capture without immediate activation

Change Data Capture with Debezium and Kafka was already useful because configuration updates needed to move quickly through the platform. An administrator might update a policy at 2 PM, and we wanted the control plane to learn about that update promptly. Capture latency is worth monitoring on its own. One parameter once stalled our Postgres replication with Debezium for three weeks.

After CDC observed an update:

  1. CDC captures the policy change.

  2. Outside a freeze, the curator forwards it with a nextWindowToken header.

  3. A downstream consumer uses that header to update the future versioned collection.

  4. Work for the applicable future window uses that version. Work carrying the current window token keeps using its own.

There is a separate branch during the pre-rollover freeze. Instead of forwarding immediately, the curator stores the incoming message in a per-tenant MongoDB stage store. After rollover, a freeze-ended event triggers replay of those messages with the new next-window token.

Capture latency is how quickly the control plane knows about a change. The effective window is which work is allowed to use that change. Neither is the same thing as the temporary buffer used during rollover.

Moving safely from window W to W+1

Keeping current-window processing on a stable version solves the mid-window update problem, but it still leaves a sensitive transition at the tenant’s next boundary.

Before that boundary, work associated with W reads version W, while normal configuration updates are forwarded to the collection for W+1. The curator stores both a current and a next token. In its bootstrap logic, the previously stored next token becomes the new current token.

The tenant’s freeze is scheduled ahead of the boundary; the checked-in default offset is 60 minutes. Once frozen, incoming changes go to MongoDB stage storage rather than the prepared W+1 collection. At rollover, the curator:

  1. Obtains another token from the quota manager.

  2. Calls downstream services to clone W+1 into W+2.

  3. Rotates its current/next token pair.

  4. Ends the freeze. A freeze-ended event then replays staged changes with W+2 as their next-window token, and normal forwarding resumes towards W+2.

Tenant rollover lifecycle from window W to W+1: normal operation, pre-rollover freeze, rollover with clone and token rotation, and freeze end with replay of buffered changes
Fig 5: The freeze starts 60 minutes before the boundary. At rollover the token pair rotates and buffered changes replay into the newly created W+2.

These are scheduled tenant-local transitions, not a guarantee that every service switches at precisely the same millisecond. The scheduler periodically checks for due jobs. Nor does “tomorrow” apply to every update made before midnight: a change received during the pre-rollover freeze is replayed to the newly created future version rather than inserted into the version being finalized. That cutoff is part of the effective-time contract.

The curator asks configured downstream services to clone the prepared version into the newly assigned future version using an endpoint of this shape:

POST /v1/tenants/{tenantName}/bootstrap-config
Content-Type: application/json

{
  "currentWindowToken": "token-for-W-plus-1",
  "nextWindowToken": "token-for-W-plus-2"

Here currentWindowToken names the clone source and nextWindowToken names its target. In the curator’s stored token pair before rollover, that source was called the next token. The names describe the bootstrap transition, not two different clocks.

The curator retries failed calls, but a retry should not be described as an unconditional no-op. Some downstream implementations drop an existing target collection before cloning it again. Safe retry behaviour depends on ordering and on whether anything has written to that target in the meantime.

Dynamic MongoDB collections for versioned configuration

With ordinary configuration storage, an entity can map to a stable MongoDB collection and a conventional repository abstraction works well. After versioning, the target collection depended on runtime context. For example, a downstream service with a surveillance-policies base collection selects a versioned collection of the form:

The tenant determines the database, the configuration type determines the base collection and the token determines the version suffix. The initial bootstrap token is a special case that uses the unsuffixed base collection.

Diagram of token-aware MongoDB collection resolution: tenant selects the database, config type the base collection and window token the version suffix, via MongoTemplate
Fig 6: Tenant, configuration type and window token together resolve the physical collection, which is then queried through MongoTemplate.

That made a statically bound MongoRepository less natural for the versioned access path. Downstream services used token-aware collection resolution and MongoTemplate where choosing the correct physical configuration version was part of processing correctness.

We were not trying to replace Spring Data repositories across the application. The choice only concerned the paths where the target collection changes with the processing context. We have also written about how enabling Spring Kafka batch listeners backfired in an event pipeline.

A query against the wrong collection could still return valid-looking configuration data. It would be valid for the wrong window.

Recovering configuration rollovers through reconciliation

The rollover spans Kafka, the curator, HTTP calls, MongoDB and multiple pieces of persisted state. Failures can therefore occur halfway through a transition:

  • The curator may restart after some downstream collections have been cloned but before the transition finishes.

  • An HTTP request may time out after the receiving service has already performed the clone.

  • Previously consumed events may appear again during recovery.

The curator uses retries and outbox events to help recover from partial failures, but this is not one atomic transaction across Kafka, MongoDB and every downstream service. It rotates tokens and ends the freeze during bootstrap; staged-message replay follows a freeze-ended event. The replay service republishes staged messages to Kafka and then deletes them from the Mongo stage store, which does not make the two systems one transaction. “Bootstrap completed” does not, by itself, prove that every staged change has been applied downstream.

That changes the questions we ask in monitoring. A successful scheduled rollover job says little by itself. Monitoring needs to show whether the expected downstream versions are available and the buffered changes have drained:

  • Has each downstream bootstrap finished?

  • Is the tenant still frozen?

  • Are staged messages draining?

  • Which current and next tokens are recorded?

  • Are downstream versions available and consistent with those tokens?

These are operational checks to build and validate, not a claim that a single reconciliation = running field or atomic activation gate already answers them.

Temporal versioning for deterministic event replay

The strongest benefit appeared when we looked at event replay. Without versioning, an unchanged event combined with the latest configuration at replay time can produce a different policy decision. After window versioning, the window token became part of the event’s processing context:

Event M + Tenant A + original window token W + policy version for

If the same event needs to be investigated later, the system can use the original token to resolve the same policy version again, provided the event retains that token and the historical collection still exists.

Comparison of event replay against latest configuration versus replay resolving the retained configuration version through the original window token
Fig 7: At replay time, the event’s token selects its configuration collection. Old versions have to be retained for this to work.

Temporal versioning makes the configuration input reproducible. It does not, by itself, guarantee the same overall result if application code, external data or other inputs have changed. Retention of old configuration versions is part of the replay contract, not just storage housekeeping.

That helped incident investigation because an engineer could identify the policy version associated with the event. It helped audits for the same reason. It also meant a configuration update no longer silently changed the meaning of work carrying a current-window token.

The window token had effectively made policy version part of processing identity.

Operational trade-offs of temporal configuration versioning

We would not use this pattern for every configuration value. It has real costs:

  • Versioned snapshots require storage, and old versions need a retention policy.

  • Rollovers need orchestration, and failures during transition need recovery.

  • Freeze-time updates need durable staging.

  • Tenant-local window boundaries add scheduling and operational behaviour.

Our implementation staged freeze-time changes in MongoDB. If we were designing it again, we would evaluate whether keeping that state in a compacted Kafka topic would simplify recovery and keep more of the configuration lifecycle inside the event pipeline. Changing the buffer alone would not solve cross-system cutover or idempotency. For governing Kafka itself as code, see how we ran GitOps for 78 Kafka clusters.

There are also cutoff timings to consider. The pre-rollover freeze offset gives the prepared version a stable cutoff before the boundary; in the checked-in configuration it is 60 minutes. Changes received during the freeze are buffered and replayed to a later version. We would evaluate the offset against observed bootstrap and replay behaviour rather than assume one duration is ideal for every tenant.

Those costs are justified when configuration changes the semantics of work. Quotas, compliance policies, entitlement decisions, routing rules, transformations and scheduled processing rules are good examples.

When temporal configuration versioning is the right consistency model

Two practical tests help decide:

  1. If this event is replayed tomorrow, does it need the policy version associated with its original processing window?

  2. Can two events belonging to the same logical window safely observe different versions of the policy?

When either answer creates a correctness problem, an unversioned “latest configuration” lookup deserves closer scrutiny.

Looking back, the original system already had an effective-time model. It was implicit. The moment an administrator saved a change became the moment that configuration affected processing. As long as policy changes did not interfere with work already underway, that behaviour was easy to live with. The 2 PM update exposed where the assumption stopped holding.

The principle we carried forward:

When configuration participates in the result of processing, its effective time belongs in the data model.

FAQ: temporal configuration versioning

What is configuration drift in stream processing?

It is when events that belong to the same logical processing window are evaluated under different configuration versions, usually because services always read the latest configuration and a change lands mid-window. Every lookup is correct, but results depend on when an event happened to execute.

What is effective-time configuration?

A model where each configuration change records both when it was made and which work it applies to. Changes are captured immediately and written to a future version, which only becomes active at a defined window boundary.

How does a window token make event replay deterministic?

Each event carries the opaque token of the processing window it belongs to. Services resolve configuration from that token rather than from the clock, so replaying the event later selects the same configuration version, as long as that version is retained.

Should every configuration value be versioned this way?

No. Versioning adds storage, rollover orchestration and recovery work. Reserve it for configuration that changes the semantics of processing, such as quotas, compliance policies, entitlements and routing rules.

If you are running into the same problems with Kafka pipelines, CDC or multi-tenant consistency, take a look at our backend engineering services or browse our case studies.

An administrator changed a policy engine rule in a multi-tenant stream-processing system we were building with a client at One2N. The change propagated correctly, and services started using it within seconds.

That was exactly the problem.

Events processed before 2 PM used the older policy while events processed after 2 PM used the new one, even when both belonged to the same logical processing window. Nothing failed operationally, yet the result of processing had started depending on when an event happened to execute. Every service was up to date, yet the output was inconsistent. That is configuration drift.

This article covers how we addressed that by introducing effective-time configuration, window-level versioning and explicit configuration resolution, and what that changed across CDC, MongoDB, service boundaries, rollover, recovery and replay.

How mutable configuration broke deterministic stream processing

The platform processed millions of messages per day across multiple tenants. Each tenant had its own policies, quotas, processing rules and local operating windows. Those configuration values were part of the actual processing logic, so two services handling the same event needed to agree on which policy applied.

The original design was straightforward: an administrator updates a policy, the configuration changes, and services read the latest configuration. When an engineer changed a policy value, services picked it up immediately and processed incoming messages with the new ruleset. For a long time, that model was reasonable.

The gotcha was a function of its intended behaviour. Consider a tenant whose processing window runs for one logical day.

Timeline of one processing window where events at 13:59 and 14:01 are evaluated with different policy versions after a 14:00 configuration change
Fig 1: A 14:00 policy change splits one logical processing window in two. Event A at 13:59 is evaluated with v17, Event B at 14:01 with v18, and both lookups are “correct”.

At 1:59 PM, an event may be evaluated with policy v17. At 2:01 PM, another event belonging to the same processing window may be evaluated with v18.

Every individual configuration lookup succeeds. Neither service has read stale data. Yet the processing window no longer has a stable set of rules, and that produced consequences beyond the immediate policy decision.

During incident investigation, explaining why an event received a particular result meant reconstructing configuration state around its processing timestamp. Replaying the same event later could produce a different result because the latest policy had changed. Two events that logically belonged together could also be evaluated under different assumptions.

The architecture had allowed processing time to influence policy selection implicitly.

Temporal versioning: separating capture time from effective time

We needed one property to hold throughout a processing window:

Work associated with the same processing window must resolve to that window’s configuration version.

That invariant sounds simple, but it changes the meaning of “current configuration”. Suppose an administrator updates a quota at 2 PM. There are two relevant moments in that change:

  1. When the system learns about the new value.

  2. When the new value becomes valid for event processing.

Those moments had previously been identical. Pressing Save effectively meant “make this policy active now”. The idea is close to what Martin Fowler describes as bitemporal history, where record time and actual (effective) time are separate axes.

We separated them. Outside rollover, the curator forwarded configuration changes promptly with a nextWindowToken. Downstream consumers used that token to update a future versioned collection, while work carrying the current window’s token continued to read its existing collection. Only during the pre-rollover freeze did the curator temporarily buffer incoming changes in MongoDB.

Diagram separating configuration capture time from effective time: a 14:00 change is captured by CDC, written to the next window's collection, and applied only at the window boundary
Fig 2: The 14:00 change is captured within seconds and written to the W+1 collection. Work carrying W’s token keeps reading version W until the boundary.

The system could still capture and propagate changes quickly. The current processing window continued using its associated policy version. A future window could incorporate the new configuration, subject to the pre-rollover cutoff described below.

That gave configuration an effective time. Fast propagation no longer meant immediate activation for current-window work. MongoDB staging only held changes during the rollover freeze.

Using window tokens to version processing windows

Once configuration could have different authoring and activation times, we needed a stable way to identify which version belonged to which work.

We used a version identifier for each tenant processing window. Internally, we called it a window token. The token is an opaque identifier, not necessarily a sequential number. W, W+1 and W+2 below are labels for consecutive windows, not literal token values.

Diagram mapping a tenant's processing windows to opaque window tokens and versioned configuration collections, with a work event resolving configuration through its token
Fig 3: Each tenant window maps to an opaque token, and each token identifies one configuration version. A work event carrying W’s token always resolves to version W.

Services depend on the association, so the token can be any opaque value. A token identifies a configuration version; it is not itself the frozen snapshot. If an event carries the token for window W, configuration resolution leads to the version associated with W.

A policy change while W is active does not modify the collection selected by work carrying W’s token. During normal operation it is forwarded to the version identified by the next token; during the pre-rollover freeze it is buffered for replay to a further future version.

This moved configuration selection away from a wall-clock lookup and towards an explicit processing relationship:

# Before: the clock decides
event -> configuration available right now

# After: the event's processing identity decides
event -> original window token -

Resolving temporal configuration across tenant time zones

The platform was multi-tenant and timezone-aware. One tenant could cross into its next logical processing day while another still had several hours remaining in its current one.

We did not want individual services implementing their own interpretation of a tenant’s active window. Even small differences in timezone conversion or rollover logic could put two services on different configuration versions.

Tenant timezone determines when the curator schedules that tenant’s freeze and rollover. It should not be necessary for every downstream service to recalculate an event’s window from its own wall clock. Instead, the work carries a windowToken, and a downstream service uses that token to resolve the corresponding versioned collection. The collection-resolution logic belongs in the shared configuration path rather than in each business query.

Wall-clock time still matters. It helps determine when a tenant moves into its next processing window. Once an event has a processing identity, however, downstream services should not keep asking the clock which configuration to use. They can resolve it from the token carried with that work. That becomes especially important when the event is processed again hours or days later.

Separating the configuration control plane from the stream-processing data plane

Versioning configuration introduced lifecycle operations that did not belong inside every streaming service. Something had to consume updates, route them to future versions, buffer changes around rollover and coordinate the creation of the next version. We separated those responsibilities from event processing.

Architecture diagram of the configuration control plane (Debezium, Kafka, config curator, quota manager, MongoDB stage store) and the stream-processing data plane of token-aware services
Fig 4: The config curator owns the configuration lifecycle. Changes are forwarded to a future version during normal operation, and the MongoDB stage store is only used during the rollover freeze.

The config curator formed the control plane. It owned the lifecycle of configuration: consuming changes, tracking tenant-local rollover schedules, routing ordinary updates to a future version, buffering updates during the freeze and coordinating bootstrap.

Streaming services formed the data plane. They used the token on incoming work to select a configuration version, and their config-sync consumers wrote curator updates to the version specified by nextWindowToken. The data plane therefore had both token-directed configuration writes and token-directed processing reads.

That separation gave us a useful architectural rule:

The component processing an event should not also decide when a new policy becomes valid.

Keeping those responsibilities apart made the consistency model much easier to reason about. We wrote more about asking these questions before drawing components in why your architecture should start with questions, not boxes.

Using CDC for configuration capture without immediate activation

Change Data Capture with Debezium and Kafka was already useful because configuration updates needed to move quickly through the platform. An administrator might update a policy at 2 PM, and we wanted the control plane to learn about that update promptly. Capture latency is worth monitoring on its own. One parameter once stalled our Postgres replication with Debezium for three weeks.

After CDC observed an update:

  1. CDC captures the policy change.

  2. Outside a freeze, the curator forwards it with a nextWindowToken header.

  3. A downstream consumer uses that header to update the future versioned collection.

  4. Work for the applicable future window uses that version. Work carrying the current window token keeps using its own.

There is a separate branch during the pre-rollover freeze. Instead of forwarding immediately, the curator stores the incoming message in a per-tenant MongoDB stage store. After rollover, a freeze-ended event triggers replay of those messages with the new next-window token.

Capture latency is how quickly the control plane knows about a change. The effective window is which work is allowed to use that change. Neither is the same thing as the temporary buffer used during rollover.

Moving safely from window W to W+1

Keeping current-window processing on a stable version solves the mid-window update problem, but it still leaves a sensitive transition at the tenant’s next boundary.

Before that boundary, work associated with W reads version W, while normal configuration updates are forwarded to the collection for W+1. The curator stores both a current and a next token. In its bootstrap logic, the previously stored next token becomes the new current token.

The tenant’s freeze is scheduled ahead of the boundary; the checked-in default offset is 60 minutes. Once frozen, incoming changes go to MongoDB stage storage rather than the prepared W+1 collection. At rollover, the curator:

  1. Obtains another token from the quota manager.

  2. Calls downstream services to clone W+1 into W+2.

  3. Rotates its current/next token pair.

  4. Ends the freeze. A freeze-ended event then replays staged changes with W+2 as their next-window token, and normal forwarding resumes towards W+2.

Tenant rollover lifecycle from window W to W+1: normal operation, pre-rollover freeze, rollover with clone and token rotation, and freeze end with replay of buffered changes
Fig 5: The freeze starts 60 minutes before the boundary. At rollover the token pair rotates and buffered changes replay into the newly created W+2.

These are scheduled tenant-local transitions, not a guarantee that every service switches at precisely the same millisecond. The scheduler periodically checks for due jobs. Nor does “tomorrow” apply to every update made before midnight: a change received during the pre-rollover freeze is replayed to the newly created future version rather than inserted into the version being finalized. That cutoff is part of the effective-time contract.

The curator asks configured downstream services to clone the prepared version into the newly assigned future version using an endpoint of this shape:

POST /v1/tenants/{tenantName}/bootstrap-config
Content-Type: application/json

{
  "currentWindowToken": "token-for-W-plus-1",
  "nextWindowToken": "token-for-W-plus-2"

Here currentWindowToken names the clone source and nextWindowToken names its target. In the curator’s stored token pair before rollover, that source was called the next token. The names describe the bootstrap transition, not two different clocks.

The curator retries failed calls, but a retry should not be described as an unconditional no-op. Some downstream implementations drop an existing target collection before cloning it again. Safe retry behaviour depends on ordering and on whether anything has written to that target in the meantime.

Dynamic MongoDB collections for versioned configuration

With ordinary configuration storage, an entity can map to a stable MongoDB collection and a conventional repository abstraction works well. After versioning, the target collection depended on runtime context. For example, a downstream service with a surveillance-policies base collection selects a versioned collection of the form:

The tenant determines the database, the configuration type determines the base collection and the token determines the version suffix. The initial bootstrap token is a special case that uses the unsuffixed base collection.

Diagram of token-aware MongoDB collection resolution: tenant selects the database, config type the base collection and window token the version suffix, via MongoTemplate
Fig 6: Tenant, configuration type and window token together resolve the physical collection, which is then queried through MongoTemplate.

That made a statically bound MongoRepository less natural for the versioned access path. Downstream services used token-aware collection resolution and MongoTemplate where choosing the correct physical configuration version was part of processing correctness.

We were not trying to replace Spring Data repositories across the application. The choice only concerned the paths where the target collection changes with the processing context. We have also written about how enabling Spring Kafka batch listeners backfired in an event pipeline.

A query against the wrong collection could still return valid-looking configuration data. It would be valid for the wrong window.

Recovering configuration rollovers through reconciliation

The rollover spans Kafka, the curator, HTTP calls, MongoDB and multiple pieces of persisted state. Failures can therefore occur halfway through a transition:

  • The curator may restart after some downstream collections have been cloned but before the transition finishes.

  • An HTTP request may time out after the receiving service has already performed the clone.

  • Previously consumed events may appear again during recovery.

The curator uses retries and outbox events to help recover from partial failures, but this is not one atomic transaction across Kafka, MongoDB and every downstream service. It rotates tokens and ends the freeze during bootstrap; staged-message replay follows a freeze-ended event. The replay service republishes staged messages to Kafka and then deletes them from the Mongo stage store, which does not make the two systems one transaction. “Bootstrap completed” does not, by itself, prove that every staged change has been applied downstream.

That changes the questions we ask in monitoring. A successful scheduled rollover job says little by itself. Monitoring needs to show whether the expected downstream versions are available and the buffered changes have drained:

  • Has each downstream bootstrap finished?

  • Is the tenant still frozen?

  • Are staged messages draining?

  • Which current and next tokens are recorded?

  • Are downstream versions available and consistent with those tokens?

These are operational checks to build and validate, not a claim that a single reconciliation = running field or atomic activation gate already answers them.

Temporal versioning for deterministic event replay

The strongest benefit appeared when we looked at event replay. Without versioning, an unchanged event combined with the latest configuration at replay time can produce a different policy decision. After window versioning, the window token became part of the event’s processing context:

Event M + Tenant A + original window token W + policy version for

If the same event needs to be investigated later, the system can use the original token to resolve the same policy version again, provided the event retains that token and the historical collection still exists.

Comparison of event replay against latest configuration versus replay resolving the retained configuration version through the original window token
Fig 7: At replay time, the event’s token selects its configuration collection. Old versions have to be retained for this to work.

Temporal versioning makes the configuration input reproducible. It does not, by itself, guarantee the same overall result if application code, external data or other inputs have changed. Retention of old configuration versions is part of the replay contract, not just storage housekeeping.

That helped incident investigation because an engineer could identify the policy version associated with the event. It helped audits for the same reason. It also meant a configuration update no longer silently changed the meaning of work carrying a current-window token.

The window token had effectively made policy version part of processing identity.

Operational trade-offs of temporal configuration versioning

We would not use this pattern for every configuration value. It has real costs:

  • Versioned snapshots require storage, and old versions need a retention policy.

  • Rollovers need orchestration, and failures during transition need recovery.

  • Freeze-time updates need durable staging.

  • Tenant-local window boundaries add scheduling and operational behaviour.

Our implementation staged freeze-time changes in MongoDB. If we were designing it again, we would evaluate whether keeping that state in a compacted Kafka topic would simplify recovery and keep more of the configuration lifecycle inside the event pipeline. Changing the buffer alone would not solve cross-system cutover or idempotency. For governing Kafka itself as code, see how we ran GitOps for 78 Kafka clusters.

There are also cutoff timings to consider. The pre-rollover freeze offset gives the prepared version a stable cutoff before the boundary; in the checked-in configuration it is 60 minutes. Changes received during the freeze are buffered and replayed to a later version. We would evaluate the offset against observed bootstrap and replay behaviour rather than assume one duration is ideal for every tenant.

Those costs are justified when configuration changes the semantics of work. Quotas, compliance policies, entitlement decisions, routing rules, transformations and scheduled processing rules are good examples.

When temporal configuration versioning is the right consistency model

Two practical tests help decide:

  1. If this event is replayed tomorrow, does it need the policy version associated with its original processing window?

  2. Can two events belonging to the same logical window safely observe different versions of the policy?

When either answer creates a correctness problem, an unversioned “latest configuration” lookup deserves closer scrutiny.

Looking back, the original system already had an effective-time model. It was implicit. The moment an administrator saved a change became the moment that configuration affected processing. As long as policy changes did not interfere with work already underway, that behaviour was easy to live with. The 2 PM update exposed where the assumption stopped holding.

The principle we carried forward:

When configuration participates in the result of processing, its effective time belongs in the data model.

FAQ: temporal configuration versioning

What is configuration drift in stream processing?

It is when events that belong to the same logical processing window are evaluated under different configuration versions, usually because services always read the latest configuration and a change lands mid-window. Every lookup is correct, but results depend on when an event happened to execute.

What is effective-time configuration?

A model where each configuration change records both when it was made and which work it applies to. Changes are captured immediately and written to a future version, which only becomes active at a defined window boundary.

How does a window token make event replay deterministic?

Each event carries the opaque token of the processing window it belongs to. Services resolve configuration from that token rather than from the clock, so replaying the event later selects the same configuration version, as long as that version is retained.

Should every configuration value be versioned this way?

No. Versioning adds storage, rollover orchestration and recovery work. Reserve it for configuration that changes the semantics of processing, such as quotas, compliance policies, entitlements and routing rules.

If you are running into the same problems with Kafka pipelines, CDC or multi-tenant consistency, take a look at our backend engineering services or browse our case studies.

An administrator changed a policy engine rule in a multi-tenant stream-processing system we were building with a client at One2N. The change propagated correctly, and services started using it within seconds.

That was exactly the problem.

Events processed before 2 PM used the older policy while events processed after 2 PM used the new one, even when both belonged to the same logical processing window. Nothing failed operationally, yet the result of processing had started depending on when an event happened to execute. Every service was up to date, yet the output was inconsistent. That is configuration drift.

This article covers how we addressed that by introducing effective-time configuration, window-level versioning and explicit configuration resolution, and what that changed across CDC, MongoDB, service boundaries, rollover, recovery and replay.

How mutable configuration broke deterministic stream processing

The platform processed millions of messages per day across multiple tenants. Each tenant had its own policies, quotas, processing rules and local operating windows. Those configuration values were part of the actual processing logic, so two services handling the same event needed to agree on which policy applied.

The original design was straightforward: an administrator updates a policy, the configuration changes, and services read the latest configuration. When an engineer changed a policy value, services picked it up immediately and processed incoming messages with the new ruleset. For a long time, that model was reasonable.

The gotcha was a function of its intended behaviour. Consider a tenant whose processing window runs for one logical day.

Timeline of one processing window where events at 13:59 and 14:01 are evaluated with different policy versions after a 14:00 configuration change
Fig 1: A 14:00 policy change splits one logical processing window in two. Event A at 13:59 is evaluated with v17, Event B at 14:01 with v18, and both lookups are “correct”.

At 1:59 PM, an event may be evaluated with policy v17. At 2:01 PM, another event belonging to the same processing window may be evaluated with v18.

Every individual configuration lookup succeeds. Neither service has read stale data. Yet the processing window no longer has a stable set of rules, and that produced consequences beyond the immediate policy decision.

During incident investigation, explaining why an event received a particular result meant reconstructing configuration state around its processing timestamp. Replaying the same event later could produce a different result because the latest policy had changed. Two events that logically belonged together could also be evaluated under different assumptions.

The architecture had allowed processing time to influence policy selection implicitly.

Temporal versioning: separating capture time from effective time

We needed one property to hold throughout a processing window:

Work associated with the same processing window must resolve to that window’s configuration version.

That invariant sounds simple, but it changes the meaning of “current configuration”. Suppose an administrator updates a quota at 2 PM. There are two relevant moments in that change:

  1. When the system learns about the new value.

  2. When the new value becomes valid for event processing.

Those moments had previously been identical. Pressing Save effectively meant “make this policy active now”. The idea is close to what Martin Fowler describes as bitemporal history, where record time and actual (effective) time are separate axes.

We separated them. Outside rollover, the curator forwarded configuration changes promptly with a nextWindowToken. Downstream consumers used that token to update a future versioned collection, while work carrying the current window’s token continued to read its existing collection. Only during the pre-rollover freeze did the curator temporarily buffer incoming changes in MongoDB.

Diagram separating configuration capture time from effective time: a 14:00 change is captured by CDC, written to the next window's collection, and applied only at the window boundary
Fig 2: The 14:00 change is captured within seconds and written to the W+1 collection. Work carrying W’s token keeps reading version W until the boundary.

The system could still capture and propagate changes quickly. The current processing window continued using its associated policy version. A future window could incorporate the new configuration, subject to the pre-rollover cutoff described below.

That gave configuration an effective time. Fast propagation no longer meant immediate activation for current-window work. MongoDB staging only held changes during the rollover freeze.

Using window tokens to version processing windows

Once configuration could have different authoring and activation times, we needed a stable way to identify which version belonged to which work.

We used a version identifier for each tenant processing window. Internally, we called it a window token. The token is an opaque identifier, not necessarily a sequential number. W, W+1 and W+2 below are labels for consecutive windows, not literal token values.

Diagram mapping a tenant's processing windows to opaque window tokens and versioned configuration collections, with a work event resolving configuration through its token
Fig 3: Each tenant window maps to an opaque token, and each token identifies one configuration version. A work event carrying W’s token always resolves to version W.

Services depend on the association, so the token can be any opaque value. A token identifies a configuration version; it is not itself the frozen snapshot. If an event carries the token for window W, configuration resolution leads to the version associated with W.

A policy change while W is active does not modify the collection selected by work carrying W’s token. During normal operation it is forwarded to the version identified by the next token; during the pre-rollover freeze it is buffered for replay to a further future version.

This moved configuration selection away from a wall-clock lookup and towards an explicit processing relationship:

# Before: the clock decides
event -> configuration available right now

# After: the event's processing identity decides
event -> original window token -

Resolving temporal configuration across tenant time zones

The platform was multi-tenant and timezone-aware. One tenant could cross into its next logical processing day while another still had several hours remaining in its current one.

We did not want individual services implementing their own interpretation of a tenant’s active window. Even small differences in timezone conversion or rollover logic could put two services on different configuration versions.

Tenant timezone determines when the curator schedules that tenant’s freeze and rollover. It should not be necessary for every downstream service to recalculate an event’s window from its own wall clock. Instead, the work carries a windowToken, and a downstream service uses that token to resolve the corresponding versioned collection. The collection-resolution logic belongs in the shared configuration path rather than in each business query.

Wall-clock time still matters. It helps determine when a tenant moves into its next processing window. Once an event has a processing identity, however, downstream services should not keep asking the clock which configuration to use. They can resolve it from the token carried with that work. That becomes especially important when the event is processed again hours or days later.

Separating the configuration control plane from the stream-processing data plane

Versioning configuration introduced lifecycle operations that did not belong inside every streaming service. Something had to consume updates, route them to future versions, buffer changes around rollover and coordinate the creation of the next version. We separated those responsibilities from event processing.

Architecture diagram of the configuration control plane (Debezium, Kafka, config curator, quota manager, MongoDB stage store) and the stream-processing data plane of token-aware services
Fig 4: The config curator owns the configuration lifecycle. Changes are forwarded to a future version during normal operation, and the MongoDB stage store is only used during the rollover freeze.

The config curator formed the control plane. It owned the lifecycle of configuration: consuming changes, tracking tenant-local rollover schedules, routing ordinary updates to a future version, buffering updates during the freeze and coordinating bootstrap.

Streaming services formed the data plane. They used the token on incoming work to select a configuration version, and their config-sync consumers wrote curator updates to the version specified by nextWindowToken. The data plane therefore had both token-directed configuration writes and token-directed processing reads.

That separation gave us a useful architectural rule:

The component processing an event should not also decide when a new policy becomes valid.

Keeping those responsibilities apart made the consistency model much easier to reason about. We wrote more about asking these questions before drawing components in why your architecture should start with questions, not boxes.

Using CDC for configuration capture without immediate activation

Change Data Capture with Debezium and Kafka was already useful because configuration updates needed to move quickly through the platform. An administrator might update a policy at 2 PM, and we wanted the control plane to learn about that update promptly. Capture latency is worth monitoring on its own. One parameter once stalled our Postgres replication with Debezium for three weeks.

After CDC observed an update:

  1. CDC captures the policy change.

  2. Outside a freeze, the curator forwards it with a nextWindowToken header.

  3. A downstream consumer uses that header to update the future versioned collection.

  4. Work for the applicable future window uses that version. Work carrying the current window token keeps using its own.

There is a separate branch during the pre-rollover freeze. Instead of forwarding immediately, the curator stores the incoming message in a per-tenant MongoDB stage store. After rollover, a freeze-ended event triggers replay of those messages with the new next-window token.

Capture latency is how quickly the control plane knows about a change. The effective window is which work is allowed to use that change. Neither is the same thing as the temporary buffer used during rollover.

Moving safely from window W to W+1

Keeping current-window processing on a stable version solves the mid-window update problem, but it still leaves a sensitive transition at the tenant’s next boundary.

Before that boundary, work associated with W reads version W, while normal configuration updates are forwarded to the collection for W+1. The curator stores both a current and a next token. In its bootstrap logic, the previously stored next token becomes the new current token.

The tenant’s freeze is scheduled ahead of the boundary; the checked-in default offset is 60 minutes. Once frozen, incoming changes go to MongoDB stage storage rather than the prepared W+1 collection. At rollover, the curator:

  1. Obtains another token from the quota manager.

  2. Calls downstream services to clone W+1 into W+2.

  3. Rotates its current/next token pair.

  4. Ends the freeze. A freeze-ended event then replays staged changes with W+2 as their next-window token, and normal forwarding resumes towards W+2.

Tenant rollover lifecycle from window W to W+1: normal operation, pre-rollover freeze, rollover with clone and token rotation, and freeze end with replay of buffered changes
Fig 5: The freeze starts 60 minutes before the boundary. At rollover the token pair rotates and buffered changes replay into the newly created W+2.

These are scheduled tenant-local transitions, not a guarantee that every service switches at precisely the same millisecond. The scheduler periodically checks for due jobs. Nor does “tomorrow” apply to every update made before midnight: a change received during the pre-rollover freeze is replayed to the newly created future version rather than inserted into the version being finalized. That cutoff is part of the effective-time contract.

The curator asks configured downstream services to clone the prepared version into the newly assigned future version using an endpoint of this shape:

POST /v1/tenants/{tenantName}/bootstrap-config
Content-Type: application/json

{
  "currentWindowToken": "token-for-W-plus-1",
  "nextWindowToken": "token-for-W-plus-2"

Here currentWindowToken names the clone source and nextWindowToken names its target. In the curator’s stored token pair before rollover, that source was called the next token. The names describe the bootstrap transition, not two different clocks.

The curator retries failed calls, but a retry should not be described as an unconditional no-op. Some downstream implementations drop an existing target collection before cloning it again. Safe retry behaviour depends on ordering and on whether anything has written to that target in the meantime.

Dynamic MongoDB collections for versioned configuration

With ordinary configuration storage, an entity can map to a stable MongoDB collection and a conventional repository abstraction works well. After versioning, the target collection depended on runtime context. For example, a downstream service with a surveillance-policies base collection selects a versioned collection of the form:

The tenant determines the database, the configuration type determines the base collection and the token determines the version suffix. The initial bootstrap token is a special case that uses the unsuffixed base collection.

Diagram of token-aware MongoDB collection resolution: tenant selects the database, config type the base collection and window token the version suffix, via MongoTemplate
Fig 6: Tenant, configuration type and window token together resolve the physical collection, which is then queried through MongoTemplate.

That made a statically bound MongoRepository less natural for the versioned access path. Downstream services used token-aware collection resolution and MongoTemplate where choosing the correct physical configuration version was part of processing correctness.

We were not trying to replace Spring Data repositories across the application. The choice only concerned the paths where the target collection changes with the processing context. We have also written about how enabling Spring Kafka batch listeners backfired in an event pipeline.

A query against the wrong collection could still return valid-looking configuration data. It would be valid for the wrong window.

Recovering configuration rollovers through reconciliation

The rollover spans Kafka, the curator, HTTP calls, MongoDB and multiple pieces of persisted state. Failures can therefore occur halfway through a transition:

  • The curator may restart after some downstream collections have been cloned but before the transition finishes.

  • An HTTP request may time out after the receiving service has already performed the clone.

  • Previously consumed events may appear again during recovery.

The curator uses retries and outbox events to help recover from partial failures, but this is not one atomic transaction across Kafka, MongoDB and every downstream service. It rotates tokens and ends the freeze during bootstrap; staged-message replay follows a freeze-ended event. The replay service republishes staged messages to Kafka and then deletes them from the Mongo stage store, which does not make the two systems one transaction. “Bootstrap completed” does not, by itself, prove that every staged change has been applied downstream.

That changes the questions we ask in monitoring. A successful scheduled rollover job says little by itself. Monitoring needs to show whether the expected downstream versions are available and the buffered changes have drained:

  • Has each downstream bootstrap finished?

  • Is the tenant still frozen?

  • Are staged messages draining?

  • Which current and next tokens are recorded?

  • Are downstream versions available and consistent with those tokens?

These are operational checks to build and validate, not a claim that a single reconciliation = running field or atomic activation gate already answers them.

Temporal versioning for deterministic event replay

The strongest benefit appeared when we looked at event replay. Without versioning, an unchanged event combined with the latest configuration at replay time can produce a different policy decision. After window versioning, the window token became part of the event’s processing context:

Event M + Tenant A + original window token W + policy version for

If the same event needs to be investigated later, the system can use the original token to resolve the same policy version again, provided the event retains that token and the historical collection still exists.

Comparison of event replay against latest configuration versus replay resolving the retained configuration version through the original window token
Fig 7: At replay time, the event’s token selects its configuration collection. Old versions have to be retained for this to work.

Temporal versioning makes the configuration input reproducible. It does not, by itself, guarantee the same overall result if application code, external data or other inputs have changed. Retention of old configuration versions is part of the replay contract, not just storage housekeeping.

That helped incident investigation because an engineer could identify the policy version associated with the event. It helped audits for the same reason. It also meant a configuration update no longer silently changed the meaning of work carrying a current-window token.

The window token had effectively made policy version part of processing identity.

Operational trade-offs of temporal configuration versioning

We would not use this pattern for every configuration value. It has real costs:

  • Versioned snapshots require storage, and old versions need a retention policy.

  • Rollovers need orchestration, and failures during transition need recovery.

  • Freeze-time updates need durable staging.

  • Tenant-local window boundaries add scheduling and operational behaviour.

Our implementation staged freeze-time changes in MongoDB. If we were designing it again, we would evaluate whether keeping that state in a compacted Kafka topic would simplify recovery and keep more of the configuration lifecycle inside the event pipeline. Changing the buffer alone would not solve cross-system cutover or idempotency. For governing Kafka itself as code, see how we ran GitOps for 78 Kafka clusters.

There are also cutoff timings to consider. The pre-rollover freeze offset gives the prepared version a stable cutoff before the boundary; in the checked-in configuration it is 60 minutes. Changes received during the freeze are buffered and replayed to a later version. We would evaluate the offset against observed bootstrap and replay behaviour rather than assume one duration is ideal for every tenant.

Those costs are justified when configuration changes the semantics of work. Quotas, compliance policies, entitlement decisions, routing rules, transformations and scheduled processing rules are good examples.

When temporal configuration versioning is the right consistency model

Two practical tests help decide:

  1. If this event is replayed tomorrow, does it need the policy version associated with its original processing window?

  2. Can two events belonging to the same logical window safely observe different versions of the policy?

When either answer creates a correctness problem, an unversioned “latest configuration” lookup deserves closer scrutiny.

Looking back, the original system already had an effective-time model. It was implicit. The moment an administrator saved a change became the moment that configuration affected processing. As long as policy changes did not interfere with work already underway, that behaviour was easy to live with. The 2 PM update exposed where the assumption stopped holding.

The principle we carried forward:

When configuration participates in the result of processing, its effective time belongs in the data model.

FAQ: temporal configuration versioning

What is configuration drift in stream processing?

It is when events that belong to the same logical processing window are evaluated under different configuration versions, usually because services always read the latest configuration and a change lands mid-window. Every lookup is correct, but results depend on when an event happened to execute.

What is effective-time configuration?

A model where each configuration change records both when it was made and which work it applies to. Changes are captured immediately and written to a future version, which only becomes active at a defined window boundary.

How does a window token make event replay deterministic?

Each event carries the opaque token of the processing window it belongs to. Services resolve configuration from that token rather than from the clock, so replaying the event later selects the same configuration version, as long as that version is retained.

Should every configuration value be versioned this way?

No. Versioning adds storage, rollover orchestration and recovery work. Reserve it for configuration that changes the semantics of processing, such as quotas, compliance policies, entitlements and routing rules.

If you are running into the same problems with Kafka pipelines, CDC or multi-tenant consistency, take a look at our backend engineering services or browse our case studies.

An administrator changed a policy engine rule in a multi-tenant stream-processing system we were building with a client at One2N. The change propagated correctly, and services started using it within seconds.

That was exactly the problem.

Events processed before 2 PM used the older policy while events processed after 2 PM used the new one, even when both belonged to the same logical processing window. Nothing failed operationally, yet the result of processing had started depending on when an event happened to execute. Every service was up to date, yet the output was inconsistent. That is configuration drift.

This article covers how we addressed that by introducing effective-time configuration, window-level versioning and explicit configuration resolution, and what that changed across CDC, MongoDB, service boundaries, rollover, recovery and replay.

How mutable configuration broke deterministic stream processing

The platform processed millions of messages per day across multiple tenants. Each tenant had its own policies, quotas, processing rules and local operating windows. Those configuration values were part of the actual processing logic, so two services handling the same event needed to agree on which policy applied.

The original design was straightforward: an administrator updates a policy, the configuration changes, and services read the latest configuration. When an engineer changed a policy value, services picked it up immediately and processed incoming messages with the new ruleset. For a long time, that model was reasonable.

The gotcha was a function of its intended behaviour. Consider a tenant whose processing window runs for one logical day.

Timeline of one processing window where events at 13:59 and 14:01 are evaluated with different policy versions after a 14:00 configuration change
Fig 1: A 14:00 policy change splits one logical processing window in two. Event A at 13:59 is evaluated with v17, Event B at 14:01 with v18, and both lookups are “correct”.

At 1:59 PM, an event may be evaluated with policy v17. At 2:01 PM, another event belonging to the same processing window may be evaluated with v18.

Every individual configuration lookup succeeds. Neither service has read stale data. Yet the processing window no longer has a stable set of rules, and that produced consequences beyond the immediate policy decision.

During incident investigation, explaining why an event received a particular result meant reconstructing configuration state around its processing timestamp. Replaying the same event later could produce a different result because the latest policy had changed. Two events that logically belonged together could also be evaluated under different assumptions.

The architecture had allowed processing time to influence policy selection implicitly.

Temporal versioning: separating capture time from effective time

We needed one property to hold throughout a processing window:

Work associated with the same processing window must resolve to that window’s configuration version.

That invariant sounds simple, but it changes the meaning of “current configuration”. Suppose an administrator updates a quota at 2 PM. There are two relevant moments in that change:

  1. When the system learns about the new value.

  2. When the new value becomes valid for event processing.

Those moments had previously been identical. Pressing Save effectively meant “make this policy active now”. The idea is close to what Martin Fowler describes as bitemporal history, where record time and actual (effective) time are separate axes.

We separated them. Outside rollover, the curator forwarded configuration changes promptly with a nextWindowToken. Downstream consumers used that token to update a future versioned collection, while work carrying the current window’s token continued to read its existing collection. Only during the pre-rollover freeze did the curator temporarily buffer incoming changes in MongoDB.

Diagram separating configuration capture time from effective time: a 14:00 change is captured by CDC, written to the next window's collection, and applied only at the window boundary
Fig 2: The 14:00 change is captured within seconds and written to the W+1 collection. Work carrying W’s token keeps reading version W until the boundary.

The system could still capture and propagate changes quickly. The current processing window continued using its associated policy version. A future window could incorporate the new configuration, subject to the pre-rollover cutoff described below.

That gave configuration an effective time. Fast propagation no longer meant immediate activation for current-window work. MongoDB staging only held changes during the rollover freeze.

Using window tokens to version processing windows

Once configuration could have different authoring and activation times, we needed a stable way to identify which version belonged to which work.

We used a version identifier for each tenant processing window. Internally, we called it a window token. The token is an opaque identifier, not necessarily a sequential number. W, W+1 and W+2 below are labels for consecutive windows, not literal token values.

Diagram mapping a tenant's processing windows to opaque window tokens and versioned configuration collections, with a work event resolving configuration through its token
Fig 3: Each tenant window maps to an opaque token, and each token identifies one configuration version. A work event carrying W’s token always resolves to version W.

Services depend on the association, so the token can be any opaque value. A token identifies a configuration version; it is not itself the frozen snapshot. If an event carries the token for window W, configuration resolution leads to the version associated with W.

A policy change while W is active does not modify the collection selected by work carrying W’s token. During normal operation it is forwarded to the version identified by the next token; during the pre-rollover freeze it is buffered for replay to a further future version.

This moved configuration selection away from a wall-clock lookup and towards an explicit processing relationship:

# Before: the clock decides
event -> configuration available right now

# After: the event's processing identity decides
event -> original window token -

Resolving temporal configuration across tenant time zones

The platform was multi-tenant and timezone-aware. One tenant could cross into its next logical processing day while another still had several hours remaining in its current one.

We did not want individual services implementing their own interpretation of a tenant’s active window. Even small differences in timezone conversion or rollover logic could put two services on different configuration versions.

Tenant timezone determines when the curator schedules that tenant’s freeze and rollover. It should not be necessary for every downstream service to recalculate an event’s window from its own wall clock. Instead, the work carries a windowToken, and a downstream service uses that token to resolve the corresponding versioned collection. The collection-resolution logic belongs in the shared configuration path rather than in each business query.

Wall-clock time still matters. It helps determine when a tenant moves into its next processing window. Once an event has a processing identity, however, downstream services should not keep asking the clock which configuration to use. They can resolve it from the token carried with that work. That becomes especially important when the event is processed again hours or days later.

Separating the configuration control plane from the stream-processing data plane

Versioning configuration introduced lifecycle operations that did not belong inside every streaming service. Something had to consume updates, route them to future versions, buffer changes around rollover and coordinate the creation of the next version. We separated those responsibilities from event processing.

Architecture diagram of the configuration control plane (Debezium, Kafka, config curator, quota manager, MongoDB stage store) and the stream-processing data plane of token-aware services
Fig 4: The config curator owns the configuration lifecycle. Changes are forwarded to a future version during normal operation, and the MongoDB stage store is only used during the rollover freeze.

The config curator formed the control plane. It owned the lifecycle of configuration: consuming changes, tracking tenant-local rollover schedules, routing ordinary updates to a future version, buffering updates during the freeze and coordinating bootstrap.

Streaming services formed the data plane. They used the token on incoming work to select a configuration version, and their config-sync consumers wrote curator updates to the version specified by nextWindowToken. The data plane therefore had both token-directed configuration writes and token-directed processing reads.

That separation gave us a useful architectural rule:

The component processing an event should not also decide when a new policy becomes valid.

Keeping those responsibilities apart made the consistency model much easier to reason about. We wrote more about asking these questions before drawing components in why your architecture should start with questions, not boxes.

Using CDC for configuration capture without immediate activation

Change Data Capture with Debezium and Kafka was already useful because configuration updates needed to move quickly through the platform. An administrator might update a policy at 2 PM, and we wanted the control plane to learn about that update promptly. Capture latency is worth monitoring on its own. One parameter once stalled our Postgres replication with Debezium for three weeks.

After CDC observed an update:

  1. CDC captures the policy change.

  2. Outside a freeze, the curator forwards it with a nextWindowToken header.

  3. A downstream consumer uses that header to update the future versioned collection.

  4. Work for the applicable future window uses that version. Work carrying the current window token keeps using its own.

There is a separate branch during the pre-rollover freeze. Instead of forwarding immediately, the curator stores the incoming message in a per-tenant MongoDB stage store. After rollover, a freeze-ended event triggers replay of those messages with the new next-window token.

Capture latency is how quickly the control plane knows about a change. The effective window is which work is allowed to use that change. Neither is the same thing as the temporary buffer used during rollover.

Moving safely from window W to W+1

Keeping current-window processing on a stable version solves the mid-window update problem, but it still leaves a sensitive transition at the tenant’s next boundary.

Before that boundary, work associated with W reads version W, while normal configuration updates are forwarded to the collection for W+1. The curator stores both a current and a next token. In its bootstrap logic, the previously stored next token becomes the new current token.

The tenant’s freeze is scheduled ahead of the boundary; the checked-in default offset is 60 minutes. Once frozen, incoming changes go to MongoDB stage storage rather than the prepared W+1 collection. At rollover, the curator:

  1. Obtains another token from the quota manager.

  2. Calls downstream services to clone W+1 into W+2.

  3. Rotates its current/next token pair.

  4. Ends the freeze. A freeze-ended event then replays staged changes with W+2 as their next-window token, and normal forwarding resumes towards W+2.

Tenant rollover lifecycle from window W to W+1: normal operation, pre-rollover freeze, rollover with clone and token rotation, and freeze end with replay of buffered changes
Fig 5: The freeze starts 60 minutes before the boundary. At rollover the token pair rotates and buffered changes replay into the newly created W+2.

These are scheduled tenant-local transitions, not a guarantee that every service switches at precisely the same millisecond. The scheduler periodically checks for due jobs. Nor does “tomorrow” apply to every update made before midnight: a change received during the pre-rollover freeze is replayed to the newly created future version rather than inserted into the version being finalized. That cutoff is part of the effective-time contract.

The curator asks configured downstream services to clone the prepared version into the newly assigned future version using an endpoint of this shape:

POST /v1/tenants/{tenantName}/bootstrap-config
Content-Type: application/json

{
  "currentWindowToken": "token-for-W-plus-1",
  "nextWindowToken": "token-for-W-plus-2"

Here currentWindowToken names the clone source and nextWindowToken names its target. In the curator’s stored token pair before rollover, that source was called the next token. The names describe the bootstrap transition, not two different clocks.

The curator retries failed calls, but a retry should not be described as an unconditional no-op. Some downstream implementations drop an existing target collection before cloning it again. Safe retry behaviour depends on ordering and on whether anything has written to that target in the meantime.

Dynamic MongoDB collections for versioned configuration

With ordinary configuration storage, an entity can map to a stable MongoDB collection and a conventional repository abstraction works well. After versioning, the target collection depended on runtime context. For example, a downstream service with a surveillance-policies base collection selects a versioned collection of the form:

The tenant determines the database, the configuration type determines the base collection and the token determines the version suffix. The initial bootstrap token is a special case that uses the unsuffixed base collection.

Diagram of token-aware MongoDB collection resolution: tenant selects the database, config type the base collection and window token the version suffix, via MongoTemplate
Fig 6: Tenant, configuration type and window token together resolve the physical collection, which is then queried through MongoTemplate.

That made a statically bound MongoRepository less natural for the versioned access path. Downstream services used token-aware collection resolution and MongoTemplate where choosing the correct physical configuration version was part of processing correctness.

We were not trying to replace Spring Data repositories across the application. The choice only concerned the paths where the target collection changes with the processing context. We have also written about how enabling Spring Kafka batch listeners backfired in an event pipeline.

A query against the wrong collection could still return valid-looking configuration data. It would be valid for the wrong window.

Recovering configuration rollovers through reconciliation

The rollover spans Kafka, the curator, HTTP calls, MongoDB and multiple pieces of persisted state. Failures can therefore occur halfway through a transition:

  • The curator may restart after some downstream collections have been cloned but before the transition finishes.

  • An HTTP request may time out after the receiving service has already performed the clone.

  • Previously consumed events may appear again during recovery.

The curator uses retries and outbox events to help recover from partial failures, but this is not one atomic transaction across Kafka, MongoDB and every downstream service. It rotates tokens and ends the freeze during bootstrap; staged-message replay follows a freeze-ended event. The replay service republishes staged messages to Kafka and then deletes them from the Mongo stage store, which does not make the two systems one transaction. “Bootstrap completed” does not, by itself, prove that every staged change has been applied downstream.

That changes the questions we ask in monitoring. A successful scheduled rollover job says little by itself. Monitoring needs to show whether the expected downstream versions are available and the buffered changes have drained:

  • Has each downstream bootstrap finished?

  • Is the tenant still frozen?

  • Are staged messages draining?

  • Which current and next tokens are recorded?

  • Are downstream versions available and consistent with those tokens?

These are operational checks to build and validate, not a claim that a single reconciliation = running field or atomic activation gate already answers them.

Temporal versioning for deterministic event replay

The strongest benefit appeared when we looked at event replay. Without versioning, an unchanged event combined with the latest configuration at replay time can produce a different policy decision. After window versioning, the window token became part of the event’s processing context:

Event M + Tenant A + original window token W + policy version for

If the same event needs to be investigated later, the system can use the original token to resolve the same policy version again, provided the event retains that token and the historical collection still exists.

Comparison of event replay against latest configuration versus replay resolving the retained configuration version through the original window token
Fig 7: At replay time, the event’s token selects its configuration collection. Old versions have to be retained for this to work.

Temporal versioning makes the configuration input reproducible. It does not, by itself, guarantee the same overall result if application code, external data or other inputs have changed. Retention of old configuration versions is part of the replay contract, not just storage housekeeping.

That helped incident investigation because an engineer could identify the policy version associated with the event. It helped audits for the same reason. It also meant a configuration update no longer silently changed the meaning of work carrying a current-window token.

The window token had effectively made policy version part of processing identity.

Operational trade-offs of temporal configuration versioning

We would not use this pattern for every configuration value. It has real costs:

  • Versioned snapshots require storage, and old versions need a retention policy.

  • Rollovers need orchestration, and failures during transition need recovery.

  • Freeze-time updates need durable staging.

  • Tenant-local window boundaries add scheduling and operational behaviour.

Our implementation staged freeze-time changes in MongoDB. If we were designing it again, we would evaluate whether keeping that state in a compacted Kafka topic would simplify recovery and keep more of the configuration lifecycle inside the event pipeline. Changing the buffer alone would not solve cross-system cutover or idempotency. For governing Kafka itself as code, see how we ran GitOps for 78 Kafka clusters.

There are also cutoff timings to consider. The pre-rollover freeze offset gives the prepared version a stable cutoff before the boundary; in the checked-in configuration it is 60 minutes. Changes received during the freeze are buffered and replayed to a later version. We would evaluate the offset against observed bootstrap and replay behaviour rather than assume one duration is ideal for every tenant.

Those costs are justified when configuration changes the semantics of work. Quotas, compliance policies, entitlement decisions, routing rules, transformations and scheduled processing rules are good examples.

When temporal configuration versioning is the right consistency model

Two practical tests help decide:

  1. If this event is replayed tomorrow, does it need the policy version associated with its original processing window?

  2. Can two events belonging to the same logical window safely observe different versions of the policy?

When either answer creates a correctness problem, an unversioned “latest configuration” lookup deserves closer scrutiny.

Looking back, the original system already had an effective-time model. It was implicit. The moment an administrator saved a change became the moment that configuration affected processing. As long as policy changes did not interfere with work already underway, that behaviour was easy to live with. The 2 PM update exposed where the assumption stopped holding.

The principle we carried forward:

When configuration participates in the result of processing, its effective time belongs in the data model.

FAQ: temporal configuration versioning

What is configuration drift in stream processing?

It is when events that belong to the same logical processing window are evaluated under different configuration versions, usually because services always read the latest configuration and a change lands mid-window. Every lookup is correct, but results depend on when an event happened to execute.

What is effective-time configuration?

A model where each configuration change records both when it was made and which work it applies to. Changes are captured immediately and written to a future version, which only becomes active at a defined window boundary.

How does a window token make event replay deterministic?

Each event carries the opaque token of the processing window it belongs to. Services resolve configuration from that token rather than from the clock, so replaying the event later selects the same configuration version, as long as that version is retained.

Should every configuration value be versioned this way?

No. Versioning adds storage, rollover orchestration and recovery work. Reserve it for configuration that changes the semantics of processing, such as quotas, compliance policies, entitlements and routing rules.

If you are running into the same problems with Kafka pipelines, CDC or multi-tenant consistency, take a look at our backend engineering services or browse our case studies.

Share
Share
On this page
Section
On this page

Continue reading.

Subscribe for more such content

Get the latest in software engineering best practices straight to your inbox. Subscribe now!

Subscribe for more such content

Get the latest in software engineering best practices straight to your inbox. Subscribe now!

Subscribe for more such content

Get the latest in software engineering best practices straight to your inbox. Subscribe now!