Follow a failure across processes
An owner reports that Maya never received an invitation. The API process is running, but that fact does not explain whether the invitation committed, the job was claimed, or SMTP rejected the message.
Observability means exposing enough evidence to explain the system's behavior. Taskboard begins with structured logs, request IDs, health endpoints, and aggregate HTTP and job metrics. Durable job records provide a separate view of background progress.
Connect the records
A request ID identifies one HTTP attempt. An invitation ID identifies the domain record. A job ID identifies background work that may have several attempts. These identifiers have different lifetimes.
{
"requestId": "request-id",
"method": "POST",
"path": "/api/workspaces/workspace-id/invitations",
"status": 201
}
The structure lets a logging tool filter fields without parsing prose. The real logger's exact fields are in the app module. Logs must exclude cookies, passwords, reset tokens, and raw email bodies. Even URLs can contain sensitive query parameters, so redaction needs more thought than hiding JSON passwords.
The error response contains a request ID the user can share. Internal logs can include diagnostic context while the response returns a safe message. Exposing stack traces to clients leaks implementation details without helping them recover.
Logs and metrics answer different questions
A log describes one occurrence. Taskboard's request log records duration and route templates. Its metrics expose HTTP request and server-error counters plus pending and failed job counts. HTTP counters are local to each process and reset on restart; job counts come from Postgres. Labels must remain bounded. A task ID or user email creates an unbounded series and belongs in restricted logs rather than metric labels.
Useful production extensions include latency histograms, job backlog age, database pool usage, and provider failure rates. The /metrics endpoint has no built-in operator authentication, so production routing must restrict access. Metrics collection does not itself create an alert; another system must evaluate conditions and notify an operator.
Tracing is another extension. A trace follows spans across the API, database, worker, and provider calls. Carrying context in a job lets an operator connect asynchronous work to its origin, while each retry remains distinguishable.
Read request logging and metrics in Read api/src/app.ts and background outcomes in Read api/src/jobs.ts.
Process uptime does not prove successful work. Connect request evidence with durable domain and job records, while keeping secrets out of diagnostic output.
Why not use the workspace ID as every metric label?
Each workspace would create separate time series. As tenants grow, that becomes expensive and can expose tenant identifiers. Aggregate metrics and restricted logs serve different purposes.