Agent artifacts need an expiration date
Raw model outputs are useful when a review fails. They are less useful when they become an indefinitely growing archive of repository content nobody has decided to retain.
In the operations post, I argued for preserving enough evidence to investigate a specific run. That raises another question: when should that evidence disappear?
The reviewer repository includes a retention utility. Its behavior is small enough to inspect, and its limitations illustrate why cleanup belongs in the workflow design rather than a forgotten housekeeping task.

Retain evidence for a reason
A review run can leave a webhook envelope, dispatch records, worker logs, raw model responses, parsed findings, and SQLite state. Those records have different purposes and should not automatically share one lifetime.
A raw response may contain source code or text copied from a private PR. A compact outcome record can preserve revision identity and the reason a run stopped without preserving every sentence the model produced.
For this service, I would define a normal investigation window for detailed artifacts, a longer period for minimal operational records where useful, and explicit preservation for unresolved incidents. The appropriate durations depend on recovery needs and data sensitivity. A default of thirty days is a convenience, not a universal privacy policy.
Retention also cannot compensate for collecting unnecessary secrets in the first place. Avoid logging credential values, restrict access to stored artifacts, and review anything copied into public posts or issue comments.
What the utility actually removes
The current command defaults to a dry-run. From the repository root, a preview is:
python3 scripts/pr_review_retention.py --days 30 --json
Adding --apply enables rewriting and deletion. I did not run cleanup against operational data for this article.
The implementation considers three groups:
- JSONL records older than the cutoff, using the latest recognized timestamp in each record.
- Matching worker log files, using file modification time.
- Top-level artifact directories whose names start with
run_, using the directory’s modification time.
Timestamp-less JSONL records and malformed lines are preserved. Directories outside the selected naming pattern are not part of that artifact cleanup. SQLite state is not pruned by this utility.
That is a concrete scope, not a claim that all runtime data has a thirty-day limit. Unknown timestamps remain unknown; keeping those records conservatively can also mean keeping them indefinitely.
File age is not workflow state
A directory’s modification time is not the completion time of its run. Updating an existing file inside it does not necessarily refresh the directory timestamp.
The utility does not consult the run’s status before deleting that directory. A failed investigation, an active run, and an old completed run can therefore need more context than an age check supplies.
Before automating retention, I would add a state-aware eligibility rule: protect active work, respect incident holds, and derive the retention clock from an explicit terminal outcome. That is a proposed improvement, not behavior the current command already provides.
Deleting raw artifacts while preserving database records can leave references to files that no longer exist. Sometimes that is the intended policy. The inspection interface should make the difference between “expired under policy” and “unexpectedly missing” visible instead of treating both as the same storage failure.
Cleanup can change deduplication behavior
Some JSONL files are more than diagnostic logs. The receiver’s append-once helper scans existing rows to decide whether a key has already been recorded. The retention utility considers every JSONL file in the selected runtime directory, including job records.
Remove an old row and that local file can no longer suppress a repeat of its key. Other protections may still prevent a duplicate external comment, but this particular deduplication memory has expired.
That means the retention window participates in the execution contract. If old deliveries can be replayed, decide whether rerunning them is acceptable. If it is not, preserve compact deduplication state separately from bulky review output.
The same applies to retry and publication state. An artifact can be disposable while the fact that an external effect already occurred still matters. Age alone does not tell the cleaner which information is safe to forget.
Make cleanup inspectable and coordinated
The JSONL rewrite takes an exclusive lock and replaces the file through a temporary file in the same directory. The receiver’s append helpers use the same lock-file convention. That coordination is important: replacing a file while another process writes through an old handle can otherwise lose records.
A lock only coordinates writers that honor it. It is also not a complete crash-durability guarantee. The utility’s log and artifact deletion paths do not gain active-run protection from the JSONL lock.
The dry-run is useful for inspecting scope before deletion. Its removed counts describe candidates even when deletion is disabled; check the mode and changed fields rather than interpreting the label as evidence that files disappeared.
For this post, the existing retention tests returned 5 passed. They exercise temporary data, including dry-run preservation, timestamp-based pruning, locking calls, and selected log and artifact deletion. They do not prove safe cleanup under every concurrent production failure.
I want retention to be predictable enough to explain before it runs: which evidence will disappear, which state must survive, and what protects unfinished work. Keeping everything forever avoids those decisions temporarily. It does not make the system easier to operate.