Testing agent workflows without calling a model
A reviewer returns an approval. Another returns a blocking finding. Should the workflow publish, stop, or attempt a fix?
That decision should not require a paid model call to test.
In the queue post, I described failure cases around accepting and recovering work. Before adding more infrastructure, I want tests that make the existing workflow’s decisions explicit. Much of that work is ordinary software testing with the model replaced at a narrow boundary.

Replace the model call, not the workflow
The reviewer runner accepts a callable named hermes_runner. Production supplies the real implementation. A test can supply a function returning a known JSON fixture.
The existing test does essentially this:
def fake_hermes(prompt):
if "role: spec_reviewer" in prompt:
return fixture_text("approve_spec.json")
if "role: quality_reviewer" in prompt:
return fixture_text("request_changes_quality.json")
raise AssertionError("unexpected reviewer prompt")
This is a shortened excerpt from the test setup, not a standalone program. The fixtures are deliberately controlled test inputs, not evidence that a model produced those answers.
The useful part is what stays real. The runner still builds prompts, parses results, checks the returned role, writes raw artifacts, and records parsed output in a temporary SQLite database.
The test then checks the roles and verdicts, verifies that the raw-output files exist, and checks that the runner made the expected calls. Replacing the whole runner with a mock would skip most of the behavior I actually want to exercise.
Test refusal paths as carefully as approval
The simplest policy tests operate directly on the deterministic reducer described in the structured-output post.
The suite covers approval with both reviewer roles and a positive checks flag. It also covers missing roles, duplicate roles, explicit requests for human attention, exhausted iteration budgets, failed checks, and empty reviewer results.
Those cases protect different invariants. Two approvals from the same role are not equivalent to one approval from each required role. An empty result list is not a quiet success.
The parser tests cover another boundary: missing fields, malformed findings, unexpected finding fields, unsafe paths, and invalid line numbers or severity values. A runner test supplies the wrong reviewer’s result and verifies that the role mismatch is rejected.
These tests do not need the model to cooperate. In fact, they are easier to write when it cannot.
Keep the database real and temporary
For small persistence tests, a temporary SQLite database is a useful middle ground. It exercises actual storage behavior without touching operational state.
The existing suite checks that creating the same run again returns the same identity, that status updates and agent results are recorded, and that artifact helpers produce readable files. These are useful checks, but their names should not imply more than their assertions establish.
For example, verifying that an artifact was written does not demonstrate survival through power loss. Testing a successful status update does not establish recovery from every partially completed transaction.
As the durable queue design develops, it needs separate tests for competing claims, expired leases, stale workers, and retry exhaustion. Those are proposed additions for that design, not capabilities proved by the current persistence tests.
What I ran for this post
I ran the existing parser, reviewer-runner, reducer, state-store, and artifact tests with this selection from the repository root:
python3 -m pytest tests/test_review_orchestrator.py -q \
-k 'load_reviewer_result or run_reviewers or reduce_review_decision or sqlite_store or artifact_helpers or artifact_path'
The result was 22 passed, 40 deselected. This was a targeted run, not the full repository suite or a production health check. The selected tests use fixtures and temporary storage; they do not need live model reviews.
The result tells me those particular assertions held against the inspected implementation. It does not tell me the reviewer finds real bugs reliably.
Keep a separate place for model evaluation
A live-model evaluation answers different questions: does the reviewer identify a known defect, invent a finding on a correct change, or admit that the requirements are incomplete?
I would use a curated set of changes with human-reviewed expected findings for that work. I would compare the substance of the findings rather than require identical prose. Model and prompt changes should be recorded so differences can be investigated instead of dismissed as randomness.
Even strong unit coverage can leave an integration gap. The reducer accepts a checks_passed argument, but the reviewer-only entrypoint currently supplies True without establishing test success there. A test proving that the reducer rejects False does not fix that wiring.
That is why I want both layers: deterministic tests for the workflow’s rules, and separate evaluations and integration checks for the evidence entering those rules. Passing one layer should never be reported as passing the other.