Skip to content
Job Contractstable

Async write and job contract

The rules for work moved off the request path: 202 means accepted not applied, jobs carry IDs and re-load state, commit-then-ack backed by idempotency, explicit ordering and coordination, bounded payloads, declared timeouts, and queues you can actually see.

Moving work off the request path trades immediacy for throughput, and that trade has a contract. This spec defines it: what an acknowledgement means, what a job may carry, how work commits relative to the queue, how it is ordered, timed, and observed, and how it fails loudly. Break any clause and you get stale reads, lost writes, or duplicated effects.

Conformance language follows RFC 2119: MUST / MUST NOT are enforceable requirements, SHOULD / SHOULD NOT are strong recommendations with legitimate exceptions, and MAY marks a genuine choice.

Contract

Acknowledgement semantics

A 202 means the work was accepted onto the queue, not that it has been applied. The client must not render the submitted value as settled truth; for data the user will act on, it confirms by reading back from the system of record. Fields whose wrongness is expensive are written synchronously, not queued.

ResponseMeansClient must
202 AcceptedEnqueued, not yet appliedShow pending; confirm via a read
200 + bodyApplied synchronouslyTrust the returned value
4xxRejected, never enqueuedSurface the error
What each acknowledgement guarantees

Contract

Transactional boundaries

A state change and the job that follows from it must not diverge. On the producer side, a job MUST NOT be enqueued before the transaction that justifies it commits — enqueue on commit, with a transactional outbox as the robust form — so a rolled-back transaction never leaves a job acting on a change that didn't happen. The residual failure (a crash after commit, before enqueue) is a lost-job risk covered by the outbox or a reconciliation sweep, not a phantom-job risk.

On the consumer side, commit-then-ack MUST be used for any job that mutates critical state: the job commits its work to the system of record, then acknowledges the message. A crash after commit but before ack causes redelivery — which is exactly why every such job is idempotent. Ack-then-commit MUST NOT be used for critical writes — a crash in the gap loses the work with the message already gone. It MAY be used only for at-most-once-tolerant work such as idempotent logging or metrics, where dropping the occasional message is acceptable, and only with explicit design-review sign-off.

A job with several effects MUST make partial failure safe — either wrap the effects so a failure rolls them back, or make each effect idempotent and separately retryable so a re-run completes only the ones that didn't land. Where a downstream effect cannot be rolled back (a sent email, an external charge), use a compensating action rather than pretending the rollback was clean. Idempotency is the safety net beneath all of it: it is what makes the redelivery commit-then-ack permits harmless.

Commit before enqueue; commit before ack
// Producer: enqueue only after the state it depends on has committed.
DB::transaction(function () use ($order) {
    $order->markPaid();
    // Outbox row committed in the same transaction; a relay enqueues post-commit.
    Outbox::record(new ChargeSettled($order->id));
});

// Consumer: commit the work, THEN ack.
// A crash before the ack => redelivery => idempotent no-op.

Enqueue MUST NOT precede the commit it depends on; a consumer that mutates critical state MUST commit before it acks (ack-first is permitted only for at-most-once-tolerant logging/metrics, by exception); and every job MUST be idempotent so the redelivery this permits is harmless.

Definition

Job payload rule

A job carries the identifiers it needs to find its data, never a hydrated model. It re-loads current state at execution time, so it acts on what is true when it runs — not on a snapshot frozen at dispatch, which under backlog may be minutes stale. Re-loading also forces the job to handle a record that was deleted in between.

IDs in the constructor; re-hydrate in handle()
class SendInvoice implements ShouldQueue
{
    public function __construct(public int $tenantId, public int $invoiceId) {}

    public function handle(): void
    {
        $invoice = Invoice::find($this->invoiceId);   // current state
        if ($invoice === null) return;                // superseded — don't act
        Mail::to($invoice->customer)->queue(new InvoiceIssued($invoice));
    }
}

Contract

Job payload limits

Carrying identifiers rather than data keeps a payload small by construction, but it MUST also stay within the queue provider's hard message-size limit — SQS, for example, caps a message at 256 KB — because a payload that exceeds it fails to enqueue, often silently at the edge.

When a job genuinely needs a large input, the input goes to object storage and the payload carries a reference — a key, not the bytes — which the job fetches at execution time. A job that must act on many records carries a bounded batch of IDs (or a query descriptor), not the hydrated rows, and very large sets are chunked across multiple jobs rather than forced into one oversized message.

PayloadRule
Small (IDs, keys)Inline in the message
Large input (file, blob)Store in object storage; carry the reference
Many recordsBounded batch of IDs, or chunk into multiple jobs
Over the provider limitMUST NOT enqueue inline — pass by reference or chunk
Payload sizing

The message payload MUST NOT exceed the queue provider's size limit; large data is passed by reference, never inline.

Contract

Idempotency and retry

Queues deliver at least once, so every job must be safe to run more than once — idempotent on a stable key. Transient failures retry with exponential backoff; they never retry immediately into a struggling dependency.

Idempotency and transactional boundaries are complementary, not redundant: idempotency makes re-running a job safe, transactional boundaries ensure partial work is rolled back or compensated, and together they guarantee a job either completes fully or can be retried without harm.

PropertyRule
IdempotencyEffect keyed on a stable ID; re-run = no-op
RetriesBounded (`tries`), with exponential backoff
OrderingNot assumed; jobs tolerate out-of-order delivery
Poison jobsLand in the failed store after max tries
Retry contract

Contract

Ordering and dependencies

Ordering is not assumed by default — jobs tolerate out-of-order delivery. When a specific order is genuinely required it MUST be enforced explicitly, not left to the queue: attach a sequence id or version and apply a change only if it is newer than the state the job finds (a stale, out-of-order job becomes a no-op), or serialise the ordered work onto a single per-entity FIFO key so order is preserved within that key.

Dependencies between jobs are modelled explicitly. Fan-out dispatches N independent children from a parent. Fan-in — do X only after all N finish — MUST be coordinated by a completion counter or a coordinator job that children signal on completion, never by a fixed delay hoping they're done; the coordinator owns the “all done” decision and is itself idempotent.

Scheduled and delayed work uses the queue's native delay / visibility timeout or a scheduler, under the same contract: a delayed job still carries IDs, re-loads state on execution, and is idempotent. The delay changes when it runs, not what it may assume.

NeedMechanism
Strict orderSequence id / version check, or a per-entity FIFO key
Fan-outParent dispatches N independent children
Fan-inCompletion counter or coordinator job — never a fixed delay
Scheduled / delayedNative delay / visibility timeout; same job contract
Coordination patterns

Contract

Timeouts and cancellation

Every job type MUST declare a timeout appropriate to its work — a job that runs unbounded ties up a worker and hides a stuck dependency. The default is short (for example 30 seconds) and raised deliberately per type for known-slow work, never left at the platform maximum.

On timeout a job is treated as a failure: it retries with backoff up to its retry budget and, on exhaustion, lands in the dead-letter store like any other poison job. Because a timed-out job may have partially run, the commit-then-ack and idempotency rules are what make the retry safe.

Long-running jobs SHOULD observe a cancellation signal — a context/cancellation token checked at safe points — so a shutdown, a superseded job, or an operator cancel stops work promptly rather than being hard-killed mid-effect. A job that cannot be cancelled cooperatively MUST at least be safe to hard-kill and re-run, which idempotency already provides.

EventBehaviour
Default timeoutShort per type (e.g. 30s); raised deliberately, not maxed
On timeoutFail → retry with backoff → dead-letter on exhaustion
CancellationCooperative via a context signal, checked at safe points
Hard killSafe by idempotency; re-run completes or no-ops
Timeout and cancellation behaviour

Definition

Dead-letter and replay

A job that exhausts its retries is recorded in a failed-jobs store and pages someone — a silently vanished job is worse than a loudly failed one. A replay path keyed on the job's identity lets you re-drive it safely once fixed, because idempotency makes replay a no-op if it already succeeded.

Contract

Observability

Async work is invisible unless it is measured. Each queue MUST emit at least queue depth (backlog), per-job processing time, and success / failure / retry rates. These are what distinguish “healthy and fast” from “silently backing up” — a growing queue depth is the earliest warning that consumers are losing to producers.

Alerting is on thresholds, not on individual failures: page when queue depth exceeds a per-queue bound (backlog draining slower than it fills), when p95 processing time crosses the job's budget, or when the failure rate exceeds a small percentage over a window. A single failure is expected; a rising failure rate is an incident.

SLIs and SLOs are defined per job type, because a payment-settlement job and a thumbnail-generation job share nothing. Each type declares its own latency and success-rate objectives, and the dead-letter rate is an explicit SLI — a job type whose poison rate climbs is failing its contract even when nothing throws on the request path.

SignalAlert when
Queue depth / backlogExceeds the per-queue bound (draining slower than filling)
Processing time (p95)Crosses the job type's latency budget
Failure / retry rateExceeds a small % over a rolling window
Dead-letter rateAny sustained rise — a per-type SLI in its own right
Minimum signals and alert thresholds

Invariants this spec guarantees

  • A 202 MUST NOT cause the client to assert a write as applied; critical values are confirmed by a read or written synchronously.
  • Jobs carry identifiers and re-load current state — never a snapshot frozen at dispatch — and payloads MUST stay within the provider's size limit, passing large inputs by reference.
  • Every job MUST be idempotent on a stable key and safe under at-least-once delivery.
  • A job MUST NOT be enqueued before the transaction it depends on commits, and a consumer that mutates critical state MUST commit before acking (ack-first is permitted only for at-most-once-tolerant logging/metrics, by exception).
  • Required ordering MUST be enforced explicitly (sequence/version or a per-entity FIFO key); fan-in is coordinated by a counter or coordinator, never a fixed delay.
  • Each job type MUST declare a timeout; on timeout it fails, retries with backoff, and dead-letters on exhaustion.
  • A job that exhausts its retries MUST be visible and alertable; each queue emits depth, processing time, and failure/retry rates with per-type SLOs.

Revision history: revised on 24 August 2026 to add transactional boundaries (commit-then-ack), payload-size limits with by-reference storage, explicit ordering and fan-out/fan-in coordination, per-type timeouts and cooperative cancellation, and observability (queue depth, processing time, failure rates, per-type SLOs). Conformance language follows RFC 2119.

Want this specified for your system?

We turn definitions like these into the actual schema, policies, and contracts your system runs on. Fixed scope, fixed price, defined delivery date.

Request a Fixed-Scope Architecture Blueprint