Latest Results
research: re-measure the recursion-depth loop through the wired engine and land what the round found (#2917)
## Summary
The first recursion-depth round measured through the product's own
engine assembly (issue #2913), the harness and product defects it
stopped on, and what the record now says.
**What the round measured** (`evals/recursion_depth/results/wired-r0/`,
`docs/reference/harness-round-log.md`,
`docs/design/recursive-decomposition.md`): one gated cap-1 cell on the
wired engine, every wiring finding read off the wire rather than off
config. 7 of 42 criteria satisfied with the program LIVE; 1 of 8 claimed
leaves survived; three merge attempts rejected by a roster reviewer;
50.5M tokens over 16 sessions. The verification register is answered per
question in the recording's README, and `results/root-causes/README.md`
carries a dated re-measured verdict per root cause: RC1 and RC4 are not
the product's, RC3 is (the assembly not acting on findings, rather than
nothing telling it), RC2 stands smaller. The round records none of the
issue's five design outcomes: the operator stopped the sweep after
attempt 5 rather than pay for the cap-1 repetitions, and the strategy
question that decision raises is #2916. This PR does **not** close
#2913: the cap-1 repetitions at the floor, the depth 2 to 4 go/no-go and
the design decision the issue asks for are not settled by one cell. Of
the issue's acceptance items, this PR delivers the wiring report with
every finding named, the treatment variables recorded in the row, the
register answered on the wire, the dossier figures re-derived, and the
never-run probes each with a result or a named reason; it leaves open
the repetition-floor decision (#2844) and the one-of-five design
outcome, which #2916 now owns.
**Harness fixes the round forced** (each one a stop point in the log):
leaf tasks filed with the host so the product's own review path runs per
leaf, with the leaf review typed in the product's vocabulary
(`TaskStatus` + `CompletionOracleVerdict`) and a parked leaf recorded as
parked; contract, merge and blind-pass sessions run as unfiled
`IN_PROGRESS` transients; the host's coordination pair published so the
peer-review gate is built and the smoke reads it; the embedder probed in
preflight and routed through the provider's litellm key; the shell tool
given the record store the build/test oracle reads; a run the host
already settled never resumed; the smallest between-arm factor a bucket
could detect reported beside the interval; a session digest script that
renders a recorded session turn by turn.
**Product fixes found on the wire**: a turn's mutating tool calls run in
the order the model issued them, and a batch stops at a fatal error or a
park (later calls are withheld with an error naming the pending approval
rather than run into a state nobody vouched for); a structured tool
argument a model sent as JSON text is decoded when the schema admits no
string; a tool result reaches the model unless an HTML document hid
something, and a rejected payload says so out loud (the DOM strip itself
now lives in `tools/html_parse_strip.py`, reading an element's parent
directly instead of through a ghost `getattr`); a reviewer's findings
travel back with the reworked task rather than the summary alone
(`engine/_review_rework_text.py`, shared by the oracle, red-team and
completion gates); the credential detector reads a secret assignment's
key as an identifier suffix and its value by shape; the L1 tool summary
names only identifier-shaped parameters and refuses an overlap; a
written file's parent directories are created and a symlinked write path
below the workspace is refused; a no-credential connection on litellm's
OpenAI-SDK route sends the placeholder key the SDK insists on; the shell
tool says which command shape records a test run; the verdict schema
names the field the rework brief reads and normalises a null-spelled
command.
**Docs**: the argument-decoding seam in `tools.md` and `conventions.md`,
the no-credential placeholder in `providers.md`, the two round-log
tracks stated in the harness log's intro, and the merge account
corrected in the design page.
Pre-reviewed by 18 agents; 41 findings addressed (one dropped as
invalid). One baseline row shrunk with approval
(`ghost_attribute_read_baseline.txt`, the retired `getparent` read).
## Test plan
- [x] `uv run ruff check` / `ruff format` clean on `src/ tests/ evals/
scripts/`
- [x] `make typecheck` clean (main and scripts daemons)
- [x] Unit suite green (44186 passed) plus the affected packages re-run
after the last batch (`tests/evals_spine`,
`tests/unit/{engine,providers,memory,tools}`)
- [x] Pre-push gate suite: 96/96 consolidated gates, markdownlint, vale,
lychee, import-linter, audits
- [ ] CI: integration, e2e, conformance and the full unit shards on the
pushed branch (the local push deferred the unit run at 1311 affected
files)
Refs #2913 (not closed; see above). Strategy follow-up: #2916. Latest Branches
0%
release-please--branches--main--components--synthorg 0%
research/remeasure-wired-loop 0%
refactor/engine-wiring-parity © 2026 CodSpeed Technology