Retry a temporary failure with a limit
Taskboard's SMTP connection times out while sending an invitation. Retrying may help because the service could recover. Retrying an invalid recipient address indefinitely does not repair the address.
A retry policy describes which failures deserve another attempt, when the attempt becomes due, and when the worker stops trying automatically.
Waiting is durable state
The worker records an attempt count, a safe error description, and a future availableAt time. It releases the claim. The job remains in Postgres while the process does other work or restarts.
Sleeping inside one worker keeps the waiting period in memory. A database timestamp survives a restart and lets another worker perform the next attempt.
Exponential backoff increases the delay after repeated failures. A cap prevents delays from growing without bound. Jitter adds variation so many workers do not retry at exactly the same instant.
const capMs = Math.min(60_000, 1_000 * 2 ** attempt);
const delayMs = Math.floor(Math.random() * capMs);
const availableAt = new Date(Date.now() + delayMs);
These numbers illustrate the policy. Taskboard's worker uses a different schedule: five seconds multiplied by two to the attempt count, capped at one hour, plus up to nine seconds of jitter. It stops after eight attempts. An invalid typed job payload fails immediately. Other handler errors currently use the bounded retry policy, so richer SMTP failure classification remains an extension.
A timeout is an uncertain outcome
Suppose SMTP accepts the message, but the response disappears. Taskboard observes a timeout even though delivery progressed. The next attempt may send a duplicate. Backoff reduces pressure on the service. It does not remove that ambiguity.
For operations with provider-supported idempotency keys, the worker should reuse the same key across retries. Generating a fresh key turns a retry into a new provider operation. SMTP has no general equivalent guarantee.
After the attempt limit, the job stays available for diagnosis rather than disappearing. A human or a controlled recovery process can inspect the cause before replaying it. Monitoring should expose overdue and exhausted work; a running worker process alone does not mean jobs are succeeding.
Read the retry state in Read api/src/jobs.ts. Database-only transactions have a separate short retry policy for serialization failures and deadlocks in Read api/src/db.ts. An external send must not move into that retried transaction callback.
Retry temporary failures, schedule retries durably, and bound automatic attempts. A timeout may mean that the external action already happened.
Why not retry immediately until SMTP responds?
Immediate retries can increase load on a failing service and keep one job monopolizing a worker. A delay gives the service recovery time and leaves capacity for other work.