InsightsAutomation5 min read

Making Automated Workflows Safe to Rerun

The hard part of automation is not the first successful run. It is recovering safely after a timeout, crash, or partially completed operation without duplicating business effects.

Making Automated Workflows Safe to Rerun

Define failure boundaries before adding retries

An enterprise workflow is rarely one atomic operation. A typical flow might read a customer from a CRM, create an ERP order, upload an attachment, notify a user through LINE, and write the result to a data platform. Any step can succeed, fail, or time out. A timeout is especially dangerous because it only tells the caller that no response arrived; the receiving system may already have completed the request. Restarting the whole workflow can therefore create a second order, send another notification, or leave two systems with conflicting states.

Before choosing a retry count, list every external side effect and identify where it becomes committed. Reads are usually safe to repeat. Creating orders, reserving inventory, sending messages, and initiating payments need explicit protection. Give every workflow run a stable identifier and track each step separately with states such as pending, running, completed, failed, outcome unknown, and compensated. The outcome-unknown state matters: treating every network timeout as a definite failure is a common cause of duplicate work.

Retry policy should follow the failure type. Temporary connectivity errors, rate limits, and unavailable services can be retried with exponential backoff and jitter. Invalid fields, missing authorization, and business-rule conflicts require correction rather than repetition. When a downstream result is uncertain, query it using the original business reference before deciding to submit the command again.

Checkpoints must preserve recoverable facts

A checkpoint is more than a log entry saying that a step finished. It must allow a different worker, with none of the original in-memory context, to resume from the right boundary. Useful checkpoints align with business milestones: input validated, ERP order created, document uploaded, or notification accepted into a delivery queue. Coarse checkpoints make recovery repeat too much work; very fine checkpoints add database writes, state complexity, and cleanup overhead. A practical rule is to checkpoint around work that is expensive, slow, or externally visible.

Each checkpoint should answer several concrete questions:

  • Execution identity:Which workflow instance, tenant, source event, and step does this record represent?
  • Input version:Which data and configuration were used, so a rerun does not silently process newer content?
  • Processing state:Has the step started, completed, failed, become uncertain, or been compensated?
  • External result:What order number, object location, message identifier, or lookup reference was returned?
  • Diagnostic context:How many attempts were made, what failed, when did it happen, and which trace links the calls?

Whenever possible, update the checkpoint and the related business data in one database transaction. When an event must be published after a database commit, use an outbox: write both the business change and a pending event in the same transaction, then let a separate publisher deliver it. This does not create exactly-once delivery. It turns the problem into manageable at-least-once delivery, provided consumers can absorb duplicates idempotently.

Idempotency means recognizing the same intent

An idempotent operation can receive the same business command more than once while producing the effect of one execution. The main design decision is the idempotency key. A newly generated request ID does not protect a rerun because every attempt looks new. Use stable business identity instead, such as source system plus source order number, tenant plus event ID, or workflow ID plus step name. Enforce that identity with a database unique constraint, and return the previously stored result when the same request arrives again.

If an external API supports idempotency keys, every retry must reuse the original key. If it does not, query for an existing resource before creating one, or keep a durable local mapping between the request and the external result. Message consumers can use an inbox table: claim the event ID uniquely, perform the business update, and mark the event complete in the same transaction. A simple read-before-write check is insufficient under concurrency because two workers may both observe that no record exists. Unique constraints, conditional writes, or carefully scoped locks provide the real protection.

Idempotency records also need a retention policy. Deleting them too early allows delayed redelivery to bypass the protection; keeping everything forever increases storage and index costs. Base the retention window on the longest realistic upstream retry period, audit requirements, and the lifetime of the business object. If the same key arrives with different content, reject it as a conflict. Silently accepting changed payloads makes later investigation almost impossible.

Irreversible steps need compensation and guardrails

A workflow spanning ERP, CRM, cloud services, and messaging platforms cannot normally share one database transaction. Instead of pretending every action can be rolled back, define a compensating action for each committed step. An order may be cancelled, reserved stock released, or an incorrect notification followed by a correction. Compensation is a new business operation, not time travel. It needs its own idempotency key, authorization checks, audit trail, retry policy, and failure handling.

Not every failure should trigger automatic compensation. If an order has entered manual review, goods have shipped, or cancellation creates more operational risk than pausing, route the case to a person. Classify each workflow step as safe to repeat, requiring status lookup, automatically compensable, or requiring manual handling. This gives operators a pre-agreed recovery path instead of forcing them to invent one during an incident.

Before release, test recovery as deliberately as the successful path:

  • Stop the worker immediately before and after every external call, then resume from the stored checkpoint.
  • Deliver the same event repeatedly and verify that no second business object or notification appears.
  • Start two identical jobs concurrently and verify that constraints and locking handle the race.
  • Simulate a downstream success whose response is lost, and confirm that the workflow queries before resubmitting.
  • Make the compensating action fail, then verify that it retries, alerts, and preserves enough context for an operator.

The rerun interface itself also needs controls. Restrict access, show which steps will execute, offer a dry-run preview where practical, and require an operator to record the reason. Monitoring should present the workflow instance, current state, last successful checkpoint, and external references together rather than leaving operators to reconstruct them from scattered logs. For integrations spanning several enterprise systems, these recovery semantics and operating procedures are as important as the happy-path implementation.

Get started

Have a project like this?

Tell us your industry, current systems and budget range. We reply within two working days and offer a free 30-minute consultation.

Chat on LINE