A webhook is not a durable queue
The receiver accepted the request. Did the review happen?
In the PR reviewer walkthrough, I called out a limitation in the automatic dispatch path: when every worker slot is occupied, the receiver logs concurrency_limited and responds with HTTP 202. It does not queue that review for later.
That behavior is visible in the code. This is not a story about a measured production outage, and I do not have a count of missed reviews to report. It is a design gap worth understanding before expanding the automation.
The next version needs to distinguish accepting work from starting it.

What the current path guarantees
The receiver checks capacity before appending a job record. If it gets a slot, it deduplicates the request using the repository, PR number, and head SHA, then dispatches a worker. If capacity is unavailable, it records the event without creating a pending review job.
A concurrency limit protects the machine from starting too much work at once. That is useful. But it does not provide a backlog, and a successful HTTP response does not create an obligation that the service knows how to fulfill later.
The job records are JSONL. Writing them to disk gives an operator evidence, but a record of a dispatch attempt is not the same thing as recoverable pending work. Recovery also needs a component that finds unfinished jobs and decides what to do with them.
Returning an error instead of 202 would make the admission failure more visible, but it would not solve recovery by itself. The sender’s actual redelivery behavior matters. I would not build an eventual-execution guarantee on an assumed retry.
Accept first, dispatch separately
For this small service, my preferred next design is a durable job table and a dispatcher. SQLite is a reasonable starting point on one host. A managed queue becomes more attractive if workers span machines or the service needs stronger operational isolation.
For supported events admitted for review, the receiver would:
- Validate the request and select the review revision.
- Insert or find the job under a unique revision key.
- Commit that record before acknowledging acceptance.
Worker capacity would then control how quickly the dispatcher drains the backlog, rather than whether accepted work survives a busy moment.
This is a proposed design, not a description of the existing webhook queue. The repository already uses SQLite for orchestration state; that alone does not mean incoming jobs have this admission and recovery contract.
A database commit also has a scope. It can protect against a receiver restart under an appropriate durability configuration. It is not a promise to survive losing the only machine and its disk.
Give running work an owner and an expiry
A job marked running forever after a worker crash is barely better than a missing job.
The dispatcher needs an atomic claim so two workers do not pick up the same pending job. One practical design records a lease owner, a lease expiry, an attempt count, and the next eligible retry time. A live worker renews its lease; an expired lease makes the job eligible for investigation or retry.
The worker must prove it still owns the claim when updating state. Otherwise, a slow old attempt can overwrite the result of a newer attempt after its lease expires. Lease expiry means ownership expired, not that the old process necessarily stopped.
Retries should be bounded and delayed. A temporary provider error may recover. An invalid result that repeats on every attempt should eventually become a visible terminal failure, not an endless source of model calls.
Completion includes the external write
Publishing a review introduces a separate ambiguity. GitHub might accept the comment just before the worker loses its connection. The worker sees an error, but the comment exists.
Retrying the whole operation blindly can duplicate the comment. Marking the job complete before posting creates the opposite risk: local state says success while GitHub has nothing.
A stable publication marker and a reconciliation step can help determine whether the intended comment already exists. The reviewer already uses comment markers for deduplication. Those markers are useful building blocks, not an atomic transaction across the database and GitHub.
A robust design still needs to handle concurrent attempts, uncertain responses, and stale lease holders. I would describe the goal as recoverable processing with idempotent effects where possible, rather than casually promise exactly-once execution.
Some pending reviews should never run
A backlog creates another decision: what if the PR receives a new commit before a waiting review starts?
For this workflow, I would normally supersede pending reviews of older heads. That is an explicit outcome, not silently dropping the job. Before publishing, the worker should check again whether its result still applies to the current revision.
That policy avoids spending model calls on a change nobody plans to merge. Other workloads may require every event to run; a PR reviewer usually cares about the latest candidate revision.
The useful operational questions then become concrete: how old is the oldest pending job, which claims expired, which jobs exhausted retries, and which publications need reconciliation?
Those are also test cases. Fill the worker slots. Crash after claiming. Simulate a successful comment followed by a lost response. Advance the PR head before publication. The next post will focus on testing these workflow decisions without needing a live model for every test.