Article: Implementing Durable Workflows on Postgres Without an External Orchestrator
| Source: InfoQ AI/ML
Tags: PostgreSQL, Temporal, durable-workflows, distributed-systems, incident-response, Kubernetes
Detailed engineering guide shows how SELECT FOR UPDATE SKIP LOCKED, primary-key idempotency constraints, and a lease-and-sweeper pattern turn a Postgres table into a durable workflow engine — eliminating Temporal or Step Functions as a separate stateful dependency.
Details
Written by the team behind Kestrel Workflows (an automation platform for incident response and cloud provisioning), this article makes a concrete case for Postgres as a self-sufficient durable execution substrate.\n\nThree building blocks carry the load. First, SELECT FOR UPDATE SKIP LOCKED turns a table into a concurrent work queue: each row is claimed by exactly one worker with no broker, leader election, or external lock service required. Second, step checkpoints are stored under a primary-key constraint (workflow_id + step_id), so a crashed worker that reruns a step reads the prior result instead of repeating a side effect — idempotency enforced by the database, not application code. Third, crash recovery uses a lease model: workers heartbeat their claimed rows, and a periodic sweeper re-enqueues any execution whose lease expired.\n\nObservability collapses to a plain SQL query. The entire reliability and security surface reduces to one dependency — the database already in the critical path. The article includes working code and schema examples.\n\nThe authors acknowledge when this pattern is inappropriate: extremely high throughput (millions of events/second), workflows that need versioned history replay at scale, or teams lacking Postgres operational expertise. For most enterprise automation scenarios, they argue the complexity tradeoff favors the approach.