Files
cgalo5758 bea0ffc024 Add Codex CLI harness to agent runner
Support Codex CLI as a second agent harness alongside opencode, with
per-harness credential mounting, model routing, and reasoning effort
handling.
2026-09-21 02:42:53 -05:00

7.2 KiB

title, audience, summary
title audience summary
Agent Runner
developer
How models from other families review this repository: what contains them, how a task is shaped, and what a finding must carry before it counts.

Agent runner

A repeatable method for putting this repository in front of models that did not write it. Two lanes use it. The security lane audits a named slice of the code against a threat model. The ideation lane answers a design brief with a proposal and a mockup. Both run the same containment and the same rules about what comes back.

The runner is scripts/agent-runner/, which documents the commands and the environment variables. This page states the method.

The problem this solves

The model that built a flow cannot audit it. It knows which inputs it already validated and which branch it already considered, so the gap it left is the gap it cannot see. A model from another family lacks that memory, and it also carries different priors: published evaluations put the same frontier model far apart on different bug classes, so the family that misses a class is not the family that catches it.

The cost of that instrument is over-reporting. Models reading code unaided invent vulnerabilities, cite lines that do not say what they claim, and rate hardening advice as critical. Every rule below exists to keep that cost bounded.

Containment

The model never sees the working tree. The runner unpacks git archive HEAD into a disposable copy, mounts that copy as the container's working directory, and deletes it at teardown. The model may edit anything it likes; the edits go nowhere, and uncommitted work is not exposed.

The model container sits on a Docker network marked internal, which has no route off the machine. Its only exit is a proxy sidecar that denies every host except an allowlist of the model APIs. The arrangement is fail-closed: a misconfigured proxy costs the run its network instead of leaking the code.

One secret reaches a container: the sign-in file of the coding agent that lane runs. For opencode that is the auth file holding the model-provider keys, mounted read-only. For Codex CLI it is a copy of the ChatGPT token in a throwaway home directory, writable because codex rewrites the file when it refreshes the token, and removed with the lane. No other credential, and no live database, identity provider or payment surface, is in the run at all.

Shaping a task

A task names a small set of files, states a one-sentence threat model, and asks falsifiable questions that each end in a yes, a no, or an explicit "undetermined" with the code that settles it. The questions are the floor: the model then hunts freely inside the slice, in whatever category it finds something.

A task is not a bug-class checklist. A checklist inherits the reviewing model's blind spots, because a class the model scores badly on stays badly scored when it is listed, and it anchors the model away from whatever the list omits. One benchmark scored a frontier model at F1 0.000 on authentication failures while the same model scored 0.689 on SSRF; a checklist would have hidden that gap behind a line item marked done.

The system context is supplied verbatim rather than left to the model to assemble. Models are poor at gathering their own context, so the auth model, the request pipeline in its actual middleware order, and the trust boundaries are written out as fact, with file and line references the model can verify, and prepended to every task. Prescriptive about slice and stakes; open about bug class.

The ideation lane shapes a task the same way: the brief names the surface, what survives any redesign, and the constraints the repository enforces, and it does not name candidate solutions. A brief that lists the shapes it expects gets those shapes back.

Method rules

  • Union, never intersection. Findings from several models are combined, not filtered by agreement. Coverage is the reason to run more than one model, and consensus filtering throws that coverage away: a model has reported a real, later-confirmed finding on one pass and reported nothing on the same slice on the next, so requiring a second reporter would have discarded it.
  • No self-review inside a run. A model never re-checks or argues with its own findings in the same context. Self-refinement loops measurably raise the count of reported critical vulnerabilities without raising the count of real ones. Adjudication is a separate pass by a different reader, against the source.
  • Every claim is spot-checked before it counts. A finding is a claim about this repository, and it is checked against this repository, at the cited path and line, before it enters any report or ledger. Claims fail this check in practice: one model reported a committed node_modules directory that does not exist here.
  • Candidates are not findings. What a run produces is a candidate list. Only what survives adjudication and the spot check becomes an issue on the forge tracker.

The evidence toll

A finding that does not pay all four is deleted before the report is written.

  1. An exact path:line-range and a verbatim quote of the code. A model that cannot quote the line has not found the bug.
  2. The full path from attacker-controlled input to the dangerous operation, naming each function it passes through. "Probably reachable" is not a path.
  3. The preconditions: what the attacker must already hold, such as no session, a member session, an operator session, a webhook signing key, or control of DNS.
  4. A failure scenario with concrete inputs and state, and the wrong outcome they produce.

Severity and difficulty are rated on two independent axes rather than combined into a CVSS score, because the one number hides which of the two a reader disagrees with. Hardening and defence-in-depth are reported in a separate section from vulnerabilities, so advice cannot inflate the vulnerability count.

Security audits: the scanners run first

The deterministic scanners run before any model reads the code, offline, over the same snapshot; the model then triages their hits as one of its tasks. That order is deliberate. A model adjudicating scanner output removes most of the scanner's false positives at a small cost in recall, while a model reading code unaided over-reports. testing.md names the scanners and fixes the order.

Scanner output is a lead, not a finding. A hit confirmed in the source becomes a candidate and pays the evidence toll like any other; a hit the model judges false is reported as false, with the reason.

Definition of done for a review

  1. Every task in the lane has run for every model in the set, and the run manifest records the audited commit, the model ids and the per-lane budget.
  2. Each candidate finding has been adjudicated against the source by a reader who did not produce it.
  3. Each surviving finding has been spot-checked at its cited path and line.
  4. Survivors are filed as issues on the forge tracker; the rest are dropped, and a finding dropped for being unreproducible is recorded as such rather than silently deleted.
  5. teardown.sh has run, and its residue check is clean.