Moving work off the request path trades immediacy for throughput, and that trade has a contract. This spec defines it: what an acknowledgement means, what a job may carry, how work commits relative to the queue, how it is ordered, timed, and observed, and how it fails loudly. Break any clause and you get stale reads, lost writes, or duplicated effects.
Conformance language follows RFC 2119: MUST / MUST NOT are enforceable requirements, SHOULD / SHOULD NOT are strong recommendations with legitimate exceptions, and MAY marks a genuine choice.
Contract
Acknowledgement semantics
A 202 means the work was accepted onto the queue, not that it has been applied. The client must not render the submitted value as settled truth; for data the user will act on, it confirms by reading back from the system of record. Fields whose wrongness is expensive are written synchronously, not queued.
| Response | Means | Client must |
|---|---|---|
| 202 Accepted | Enqueued, not yet applied | Show pending; confirm via a read |
| 200 + body | Applied synchronously | Trust the returned value |
| 4xx | Rejected, never enqueued | Surface the error |
Contract
Transactional boundaries
A state change and the job that follows from it must not diverge. On the producer side, a job MUST NOT be enqueued before the transaction that justifies it commits — enqueue on commit, with a transactional outbox as the robust form — so a rolled-back transaction never leaves a job acting on a change that didn't happen. The residual failure (a crash after commit, before enqueue) is a lost-job risk covered by the outbox or a reconciliation sweep, not a phantom-job risk.
On the consumer side, commit-then-ack MUST be used for any job that mutates critical state: the job commits its work to the system of record, then acknowledges the message. A crash after commit but before ack causes redelivery — which is exactly why every such job is idempotent. Ack-then-commit MUST NOT be used for critical writes — a crash in the gap loses the work with the message already gone. It MAY be used only for at-most-once-tolerant work such as idempotent logging or metrics, where dropping the occasional message is acceptable, and only with explicit design-review sign-off.
A job with several effects MUST make partial failure safe — either wrap the effects so a failure rolls them back, or make each effect idempotent and separately retryable so a re-run completes only the ones that didn't land. Where a downstream effect cannot be rolled back (a sent email, an external charge), use a compensating action rather than pretending the rollback was clean. Idempotency is the safety net beneath all of it: it is what makes the redelivery commit-then-ack permits harmless.
// Producer: enqueue only after the state it depends on has committed.
DB::transaction(function () use ($order) {
$order->markPaid();
// Outbox row committed in the same transaction; a relay enqueues post-commit.
Outbox::record(new ChargeSettled($order->id));
});
// Consumer: commit the work, THEN ack.
// A crash before the ack => redelivery => idempotent no-op.Enqueue MUST NOT precede the commit it depends on; a consumer that mutates critical state MUST commit before it acks (ack-first is permitted only for at-most-once-tolerant logging/metrics, by exception); and every job MUST be idempotent so the redelivery this permits is harmless.
Definition
Job payload rule
A job carries the identifiers it needs to find its data, never a hydrated model. It re-loads current state at execution time, so it acts on what is true when it runs — not on a snapshot frozen at dispatch, which under backlog may be minutes stale. Re-loading also forces the job to handle a record that was deleted in between.
class SendInvoice implements ShouldQueue
{
public function __construct(public int $tenantId, public int $invoiceId) {}
public function handle(): void
{
$invoice = Invoice::find($this->invoiceId); // current state
if ($invoice === null) return; // superseded — don't act
Mail::to($invoice->customer)->queue(new InvoiceIssued($invoice));
}
}Contract
Job payload limits
Carrying identifiers rather than data keeps a payload small by construction, but it MUST also stay within the queue provider's hard message-size limit — SQS, for example, caps a message at 256 KB — because a payload that exceeds it fails to enqueue, often silently at the edge.
When a job genuinely needs a large input, the input goes to object storage and the payload carries a reference — a key, not the bytes — which the job fetches at execution time. A job that must act on many records carries a bounded batch of IDs (or a query descriptor), not the hydrated rows, and very large sets are chunked across multiple jobs rather than forced into one oversized message.
| Payload | Rule |
|---|---|
| Small (IDs, keys) | Inline in the message |
| Large input (file, blob) | Store in object storage; carry the reference |
| Many records | Bounded batch of IDs, or chunk into multiple jobs |
| Over the provider limit | MUST NOT enqueue inline — pass by reference or chunk |
The message payload MUST NOT exceed the queue provider's size limit; large data is passed by reference, never inline.
Contract
Idempotency and retry
Queues deliver at least once, so every job must be safe to run more than once — idempotent on a stable key. Transient failures retry with exponential backoff; they never retry immediately into a struggling dependency.
Idempotency and transactional boundaries are complementary, not redundant: idempotency makes re-running a job safe, transactional boundaries ensure partial work is rolled back or compensated, and together they guarantee a job either completes fully or can be retried without harm.
| Property | Rule |
|---|---|
| Idempotency | Effect keyed on a stable ID; re-run = no-op |
| Retries | Bounded (`tries`), with exponential backoff |
| Ordering | Not assumed; jobs tolerate out-of-order delivery |
| Poison jobs | Land in the failed store after max tries |
Contract
Ordering and dependencies
Ordering is not assumed by default — jobs tolerate out-of-order delivery. When a specific order is genuinely required it MUST be enforced explicitly, not left to the queue: attach a sequence id or version and apply a change only if it is newer than the state the job finds (a stale, out-of-order job becomes a no-op), or serialise the ordered work onto a single per-entity FIFO key so order is preserved within that key.
Dependencies between jobs are modelled explicitly. Fan-out dispatches N independent children from a parent. Fan-in — do X only after all N finish — MUST be coordinated by a completion counter or a coordinator job that children signal on completion, never by a fixed delay hoping they're done; the coordinator owns the “all done” decision and is itself idempotent.
Scheduled and delayed work uses the queue's native delay / visibility timeout or a scheduler, under the same contract: a delayed job still carries IDs, re-loads state on execution, and is idempotent. The delay changes when it runs, not what it may assume.
| Need | Mechanism |
|---|---|
| Strict order | Sequence id / version check, or a per-entity FIFO key |
| Fan-out | Parent dispatches N independent children |
| Fan-in | Completion counter or coordinator job — never a fixed delay |
| Scheduled / delayed | Native delay / visibility timeout; same job contract |
Contract
Timeouts and cancellation
Every job type MUST declare a timeout appropriate to its work — a job that runs unbounded ties up a worker and hides a stuck dependency. The default is short (for example 30 seconds) and raised deliberately per type for known-slow work, never left at the platform maximum.
On timeout a job is treated as a failure: it retries with backoff up to its retry budget and, on exhaustion, lands in the dead-letter store like any other poison job. Because a timed-out job may have partially run, the commit-then-ack and idempotency rules are what make the retry safe.
Long-running jobs SHOULD observe a cancellation signal — a context/cancellation token checked at safe points — so a shutdown, a superseded job, or an operator cancel stops work promptly rather than being hard-killed mid-effect. A job that cannot be cancelled cooperatively MUST at least be safe to hard-kill and re-run, which idempotency already provides.
| Event | Behaviour |
|---|---|
| Default timeout | Short per type (e.g. 30s); raised deliberately, not maxed |
| On timeout | Fail → retry with backoff → dead-letter on exhaustion |
| Cancellation | Cooperative via a context signal, checked at safe points |
| Hard kill | Safe by idempotency; re-run completes or no-ops |
Definition
Dead-letter and replay
A job that exhausts its retries is recorded in a failed-jobs store and pages someone — a silently vanished job is worse than a loudly failed one. A replay path keyed on the job's identity lets you re-drive it safely once fixed, because idempotency makes replay a no-op if it already succeeded.
Contract
Observability
Async work is invisible unless it is measured. Each queue MUST emit at least queue depth (backlog), per-job processing time, and success / failure / retry rates. These are what distinguish “healthy and fast” from “silently backing up” — a growing queue depth is the earliest warning that consumers are losing to producers.
Alerting is on thresholds, not on individual failures: page when queue depth exceeds a per-queue bound (backlog draining slower than it fills), when p95 processing time crosses the job's budget, or when the failure rate exceeds a small percentage over a window. A single failure is expected; a rising failure rate is an incident.
SLIs and SLOs are defined per job type, because a payment-settlement job and a thumbnail-generation job share nothing. Each type declares its own latency and success-rate objectives, and the dead-letter rate is an explicit SLI — a job type whose poison rate climbs is failing its contract even when nothing throws on the request path.
| Signal | Alert when |
|---|---|
| Queue depth / backlog | Exceeds the per-queue bound (draining slower than filling) |
| Processing time (p95) | Crosses the job type's latency budget |
| Failure / retry rate | Exceeds a small % over a rolling window |
| Dead-letter rate | Any sustained rise — a per-type SLI in its own right |
Invariants this spec guarantees
- A 202 MUST NOT cause the client to assert a write as applied; critical values are confirmed by a read or written synchronously.
- Jobs carry identifiers and re-load current state — never a snapshot frozen at dispatch — and payloads MUST stay within the provider's size limit, passing large inputs by reference.
- Every job MUST be idempotent on a stable key and safe under at-least-once delivery.
- A job MUST NOT be enqueued before the transaction it depends on commits, and a consumer that mutates critical state MUST commit before acking (ack-first is permitted only for at-most-once-tolerant logging/metrics, by exception).
- Required ordering MUST be enforced explicitly (sequence/version or a per-entity FIFO key); fan-in is coordinated by a counter or coordinator, never a fixed delay.
- Each job type MUST declare a timeout; on timeout it fails, retries with backoff, and dead-letters on exhaustion.
- A job that exhausts its retries MUST be visible and alertable; each queue emits depth, processing time, and failure/retry rates with per-type SLOs.
Revision history: revised on 24 August 2026 to add transactional boundaries (commit-then-ack), payload-size limits with by-reference storage, explicit ordering and fan-out/fan-in coordination, per-type timeouts and cooperative cancellation, and observability (queue depth, processing time, failure rates, per-type SLOs). Conformance language follows RFC 2119.