Operating an agent workflow: status, evidence, and failure visibility
A review never appeared on the PR. Where do you start?
The model transcript might eventually help, but it is rarely the first thing I want. First I need to know whether the service admitted the work, started a worker, received a result, and attempted to publish it.
In the testing post, I separated workflow tests from model evaluation. Operations needs a similar separation: process execution, review decisions, and externally visible results are different pieces of evidence.
The reviewer repository has a small operations CLI to make some of that evidence easier to inspect. It is useful, but its limits matter as much as its commands.

Start with the revision, not the conversation
The logical identity of a review is the repository, PR number, and head SHA. That identity connects dispatch records to orchestration state.
The operations tool reads the worker run log and SQLite records, merges them by revision identity, and sorts the results by their latest recorded timestamp. It can show the worker’s process ID, return code, timeout flag, orchestration status, agent-result count, and log location when those fields are available.
From the repository root, the entry points are:
python3 scripts/pr_review_ops.py recent --limit 10
python3 scripts/pr_review_ops.py failures --limit 20
python3 scripts/pr_review_ops.py health --json
These commands are an index into the investigation, not a replacement for the underlying evidence. A row can tell me where to look next without dumping a whole model conversation into the terminal.
The current merge produces a revision-level view. If there were multiple attempts for the same revision, that view is not a complete attempt history. The raw event records remain important when reconstructing the sequence.
Keep process outcomes separate from review decisions
A nonzero exit code or timeout is a process problem. A valid review requesting changes can be a successful execution with an unfavorable review decision. A request for human attention may be exactly the right outcome.
The CLI’s failure filter includes failed states, preflight failures, stale-head cancellations, timeouts, and nonzero return codes. It does not classify needs_human as a failure merely because of that status.
That distinction is defensible, but it creates an operational trap: an empty failures result does not mean nobody needs to act. A triage view needs to consider human-attention states separately. A cancelled stale review may require no intervention, while a process that exited cleanly may still have left a question unanswered.
I want the interface to tell me what kind of outcome occurred rather than flatten everything into a green or red dot.
A health summary is not an end-to-end probe
The local health command counts known runs and failures, reports referenced logs that are missing, summarizes database tables, and inventories artifact storage. It does not contact GitHub, test the model provider, or establish that the receiver is listening.
It also returns ok: true when its data locations are missing and there are no known runs. I checked that behavior directly using temporary paths. The flag means the summary completed; it is not proof that the review service is healthy.
The JSONL loader skips malformed lines. That makes inspection tolerant of damaged records, but it can also hide missing evidence. A reliable monitor should distinguish “no failures recorded” from “the records could not all be read.”
For a stronger operational check, I would combine local state inspection with receiver readiness, data-source availability, and verification of an expected recent external result. Those are proposed checks, not capabilities proved by this summary command.
Keep detail behind the status
SQLite tracks orchestration state and events. Raw review artifacts preserve what the model returned. Worker logs help explain crashes or command failures. GitHub shows the result the developer actually sees.
Each has a different job. A successful local transition does not prove that GitHub accepted a comment, and an existing comment does not prove it belongs to the current head.
When those sources disagree, investigate the boundary instead of accepting whichever looks happiest. Was the response lost after an external write? Did the PR advance? Did a log disappear under retention? The goal is to explain a specific run, not collect more transcripts by default.
Retention is part of this design. Old logs and artifacts consume space and may contain private repository content. Cleanup should preserve enough evidence for the recovery window, and public summaries should never copy raw logs without reviewing them for secrets and sensitive data.
Test what the operational surface claims
For this post I ran the existing operations tests:
python3 -m pytest tests/test_pr_review_ops.py -q
The result was 4 passed. I also exercised the missing-data health summary and the needs_human classification directly. This was not a live deployment health check, and the test result does not establish that reviews are currently being published correctly.
The next improvement I would make is sharper semantics: distinguish summary completion from service health, count attention-required outcomes separately, and report unreadable input. Before adding a dashboard, I want each label to make a claim its evidence can support.