Structured agent outputs: reviewer, fixer, arbiter
In the PR reviewer walkthrough, I described the service around the model call. This post looks at what comes back from the model.
A reviewer saying “mostly fine, but the error handling worries me” gives a person something to investigate. It does not give an automated workflow a clear next action. Is that a blocker? A suggestion? A request for more context?
The service needs a contract it can validate before acting on the answer.

Two review roles, one result format
The implementation defines a specification reviewer and a quality reviewer. The first focuses on requested behavior and scope. The second focuses on correctness, tests, reliability, and maintainability.
They return the same basic structure. Here is a synthetic example of a minimal approval result, not output from a real review:
{
"schema_version": 1,
"role": "quality_reviewer",
"verdict": "APPROVE",
"confidence": "medium",
"blocking_findings": [],
"non_blocking_findings": [],
"suggested_tests": [],
"needs_human_reason": null
}
A request for changes must include at least one blocking finding. Each finding needs an ID, severity, relative file path, positive line number, summary, rationale, and suggested fix. A request for human attention needs a reason.
The specification reviewer cannot discover requirements that were never provided. Its prompt includes the PR description and diff, so the quality of that input still limits the review.
Also, separate roles do not imply parallel execution. The current runner invokes them sequentially. The distinction is about their jobs, not a claim about concurrency or statistical independence.
Parse before deciding
The result loader parses JSON and validates required fields, recognized roles and verdicts, finding shapes, and path syntax. It rejects extra fields inside finding objects. The runner also checks that the returned role matches the reviewer it requested.
Those checks catch operational mistakes before a result reaches the next stage. A missing finding is different from an empty list. A malformed path should not be passed to a component that might later interpret it as a file to edit.
There are limits. A syntactically safe relative path might name a file that does not exist. A positive line number might point beyond the end of a file. A rationale can be well formed and wrong.
The current validator is not a complete semantic consistency checker either. For example, it does not reject an APPROVE result merely because its blocking-findings list is nonempty. A stricter contract should reject that contradiction rather than leave downstream code to resolve it.
That is the distinction I care about: validating an interface makes failures easier to handle. It does not validate the model’s reasoning.
The arbiter is ordinary code
Earlier in this series I used “arbiter” for the component that decides what happens next. In the implementation inspected for this post, that decision is a Python reducer, not another model call.
A reviewer asking for a human causes escalation. Blocking findings produce a fix decision while the iteration budget remains, or escalation when it is exhausted. Approval requires both reviewer roles, without duplicate role results, plus a positive checks flag.
There is an important caveat: the reviewer-only entrypoint currently passes checks_passed=True to that reducer. It does not establish that by running a test suite there. An automated approval therefore must not be described as evidence that tests passed.
The deterministic reducer is preferable to another model when the decision is already expressible as a small policy. It makes the rules inspectable and testable without asking a third agent to reinterpret the first two.
A fix decision is not permission to push
The reducer can return fix, but the surrounding execution path determines whether fixing is allowed. The documented webhook path remains reviewer-only; it does not provide the command and worktree dependencies needed for branch mutation.
The repository contains separate fixer primitives, including disabled-by-default configuration and verdicts such as PATCH_APPLIED, NO_SAFE_FIX, and NEEDS_HUMAN. Their presence is not evidence of an enabled automatic repair service.
A future rollout needs to verify actual changes and command results rather than trust a patch claim in JSON. It also needs a clear answer to what happens when the branch moves while the fixer is working.
For now, the useful boundary is straightforward: agents produce review artifacts, code validates and routes them, and humans retain authority over changes the workflow is not equipped to make safely.