Agentic workflows need boring guardrails
A bot that leaves a convincing review on a pull request is easy to appreciate. A bot that leaves the same review twice, reviews an obsolete commit, or silently fails is less useful.
Those problems do not disappear with a better prompt. They belong to the workflow around the model.
In the Hermes runtime post, I described the tools available to an agent. Before connecting those tools to recurring work, I want a clear answer to a different question: what limits the damage when the agent is wrong?

Start with a smaller job
For PR automation, reviewing a change and modifying a branch are separate permissions.
The repository for my Hermes-backed reviewer now includes an opt-in orchestration path with SQLite state and structured artifacts. The webhook path is reviewer-only. It does not mutate PR branches, and the fixer loop is not enabled in that path.
That distinction is deliberate. A review can still mislead a person, leak information, or waste time. But letting the same workflow write code creates additional failure modes. It needs a separate rollout decision, not an enthusiastic sentence in the reviewer prompt.
My preferred starting point is a reviewer that produces evidence for a human. Automatic fixes can wait until there is a narrow policy for which changes qualify and how they will be verified.
A prompt is not a permission boundary
“Do not merge” is a useful instruction. Credentials that cannot merge provide a stronger boundary.
The same reasoning applies to filesystem access and secrets. If a review job only needs a diff and selected source files, giving it broad access to an operational machine makes the consequences of a mistake much larger.
Repository content is also input, not authority. A comment in a source file telling the reviewer to ignore its instructions must not become a new instruction. If reviewing a PR involves executing its code, that needs an isolated environment without privileged credentials. A signed webhook does not make the submitted code trustworthy.
These are design requirements, not a claim that a “reviewer-only” label automatically provides a sandbox.
Know which change the answer describes
A review should identify the repository, PR, and head commit it examined. The service uses this logical identity:
repo#pr@head_sha
The revision matters. The same PR can contain different code a few minutes later. Duplicate delivery and a new revision are different events, and the system needs to tell them apart.
Before publishing a result, a workflow should check whether it still applies to the current revision. A finding can be accurate about an old diff and misleading on the current one.
Stable identity also gives an operator something to search for when a job fails. “The agent ran earlier” is not enough to investigate.
Validate the result before acting on it
Separate reviewers can examine behavior and code quality, but adding reviewers does not create correctness by consensus. They can share the same blind spot.
What helps operationally is requiring an explicit result: a verdict, findings tied to source locations, and enough reasoning for someone to check the claim. Structured output lets ordinary code reject missing fields or unexpected verdicts before taking the next action.
Valid JSON is still only valid JSON. It does not prove that a finding is true. The source location must exist, the reasoning must fit the code, and any claimed test run needs actual execution evidence.
I would rather surface a malformed result as a failed review than quietly interpret it as approval.
Make stopping ordinary
Timeouts, a failed model call, and ambiguous findings are normal outcomes to design for. A workflow needs a visible way to say that it could not finish.
The reviewer documents worker timeouts, failure comments, and commit statuses for the orchestration path. Those mechanisms make failures visible; they do not guarantee recovery.
For any future fixer loop, I want explicit limits on attempts and changes, with human review for authentication, secrets, destructive operations, and other high-risk work. A second agent should not be able to grant the first one broader authority.
Autonomy is the liability. The goal is leverage. For this project, that means a bounded review with a traceable result before it means a bot that keeps editing until it likes its own answer.