Parottasalna

AI, Backend Engineering & Architecture Guides

Learning Notes #31 – Outbox Pattern | Cloud Pattern

Today, i came across https://www.milanjovanovic.tech/blog/outbox-pattern-for-reliable-microservices-messaging which explains about Outbox pattern, which enables atleast-once delivery mechanism. In this blog, i jot down notes on outbox pattern for better understanding.

In distributed systems, ensuring data consistency and reliable communication across multiple services is a significant challenge. The Outbox Pattern is a proven architectural solution to this problem, helping developers manage data consistency, especially when dealing with events, messaging systems, or external APIs.

The Need for the Outbox Pattern

In systems where services need to perform database updates and send messages to other services simultaneously, achieving atomicity (ensuring both actions succeed or fail together) can be complex.

Common Challenges:

  1. Two-Phase Commit Complexity: Implementing distributed transactions using two-phase commit (2PC) is often complicated, resource-intensive, and impacts performance.
  2. Partial Failures: If a database operation succeeds but the subsequent message delivery fails (or vice versa), the system can enter an inconsistent state.
  3. High Latency: Synchronous operations, where services wait for confirmations of message delivery, can significantly slow down the system.
  4. Scalability Issues: Traditional approaches may not scale effectively in high-throughput environments.

The Outbox Pattern addresses these issues by separating the concerns of data persistence and message delivery, ensuring both are handled reliably and independently.

Consider an example, where we need to send an email notification when a user is created, along with populating in message queues.

But if either of the flow is failed, then its not consistent, like below.

How it works ?

The Outbox Pattern decouples database operations from message delivery by introducing an intermediary step: the Outbox Table.

1. Transactionally Persist Data and Events: When a service performs a database operation, it simultaneously writes an “outbox” record in the same transaction. This outbox record contains details of the message/event to be sent.

2. Outbox Table: This is a dedicated table in the database that acts as a staging area for messages or events.

3. Message Dispatcher: A separate process (or thread) periodically reads the outbox table, sends the messages to the target systems, and marks them as processed.

4. Idempotency: Ensure that messages sent from the outbox table are idempotent to handle retries gracefully.

    Advantages of the Outbox Pattern

    1. Atomicity: By including the outbox record in the same transaction as the database update, you ensure both succeed or fail together.
    2. Retry Mechanism: Messages that fail to send can remain in the outbox table for subsequent retry attempts.
    3. Scalability: The outbox table and dispatcher can be scaled independently to handle higher loads.
    4. Reduced Complexity: Avoids the need for distributed transactions or complex coordination mechanisms.
    5. Flexibility: Works seamlessly with various message brokers (e.g., RabbitMQ, Kafka, AWS SQS) or direct API calls.

    Implementing the Outbox Pattern

    Step 1: Design the Outbox Table

    Define an outbox table with the following fields:

    • ID: Unique identifier for each record.
    • Payload: The data or event to be sent (e.g., JSON).
    • Status: Indicates whether the message is pending, sent, or failed.
    • Created At: Timestamp for when the record was created.
    • Retries: Number of attempts made to send the message.

    Step 2: Modify Your Application Logic

    Update your service to:

    • Perform the main database operation (e.g., order creation).
    • Write a corresponding outbox record in the same transaction.

    Step 3: Implement the Dispatcher

    The dispatcher reads pending records from the outbox table, sends messages to the appropriate target (e.g., a message broker or API), and updates the status of the record. Consider using a job queue or a background worker to run the dispatcher.

    Step 4: Ensure Idempotency

    Design the dispatcher to handle duplicate sends without side effects. For example, include unique message IDs to prevent duplicate processing by the receiver.

    Use Cases

    1. Event-Driven Architectures: Reliable delivery of events to a message broker like Kafka or RabbitMQ.
    2. Microservices Communication: Decoupled communication between microservices using message queues.
    3. Third-Party Integrations: Safely sending data to external APIs or services without risking data loss.
    4. Audit Logs: Persisting and delivering audit log entries reliably.

    Best Practices

    1. Monitoring and Alerts: Monitor the outbox table for unprocessed or failed records and alert on anomalies.
    2. Batch Processing: Optimize the dispatcher to process records in batches for better performance.
    3. Dead Letter Queue (DLQ): Move records that repeatedly fail to send to a DLQ for further analysis and manual intervention.
    4. Cleanup Policies: Periodically archive or delete processed records to keep the outbox table manageable.

    Challenges and Considerations

    1. Table Growth: The outbox table can grow quickly. Implement mechanisms to archive or clean up old records.
    2. Dispatcher Reliability: Ensure the dispatcher process is robust and handles failures gracefully.
    3. Database Load: Reading from the outbox table can add load to the database. Optimize queries and consider using read replicas.

    Discover more from Parottasalna

    Subscribe now to keep reading and get access to the full archive.

    Continue reading