Exhausted jobs need a recovery decision
The email provider rejects every attempt for an invitation. Taskboard eventually reaches its automatic retry limit. Deleting the job would hide unfinished work. Retrying forever would consume worker capacity without fixing the cause.
An exhausted job remains a durable record of failure. Some systems move these records into a dead-letter queue. Taskboard can retain them in the jobs table with the attempt count and last error. The underlying purpose is the same: preserve work that needs diagnosis.
Read the failure in context
A useful failure record identifies the job type, creation time, attempts, scheduling state, and a safe error summary. Taskboard stores the error's class name rather than its potentially sensitive message. Worker logs summarize batches; the stored job ID identifies one record for inspection and replay. A richer per-attempt log is an extension. The record must avoid dumping secret-bearing email bodies or raw credentials.
An operator distinguishes several causes. A provider outage may justify replay after recovery. A bad recipient may need corrected domain data. A code defect may require a deploy before replay. A job referencing a deleted workspace may need cancellation rather than another send.
Here is an illustrative recovery decision:
Job: invitation email
Attempts: limit reached
Cause: SMTP authentication failed
Action: fix worker credentials, then schedule a controlled replay
Taskboard's worker accepts --retry followed by a failed job ID. The recovery function updates only rows with a terminal failure timestamp, clears failure and attempts, and makes the job due now. This condition keeps ordinary active jobs outside the recovery operation.
export async function retryJob(
database: Database,
jobId: string,
): Promise<boolean> {
const rows = await database.db
.update(jobs)
.set({
failedAt: null,
attempts: 0,
availableAt: new Date(),
lastError: null,
leaseToken: null,
leaseExpiresAt: null,
})
.where(and(eq(jobs.id, jobId), isNotNull(jobs.failedAt)))
.returning({ id: jobs.id });
return rows.length === 1;
}
Updating a leased job arbitrarily can race with a worker that still has valid ownership.
Replay can repeat an external action
An exhausted job might have succeeded externally on an earlier attempt whose response was lost. Replaying it can send another message. The operator needs that fact when choosing recovery, especially for payments or irreversible actions.
Operational monitoring should measure unfinished work by age, not only by row count. One overdue verification email can matter even when the queue is small. The reference worker records failures; a production notification service and recovery UI are extensions unless the code explicitly provides them.
Inspect attempts and error handling in Read api/src/jobs.ts.
Stopping automatic retries does not finish the obligation. Keep the failure visible and make replay a deliberate operation with the same concurrency protections.
Why not reset every failed job's attempt count each night?
That policy repeatedly retries permanent failures and hides how long they have remained unresolved. Diagnose the cause and choose a scoped recovery action instead.