Avatar for the Aureliolo user
Aureliolo
synthorg
BlogDocsChangelog

Performance History

Latest Results

chore(main): release 0.9.5
release-please--branches--main--components--synthorg
1 month ago
research: re-measure the recursion-depth loop through the wired engine and land what the round found (#2917) ## Summary The first recursion-depth round measured through the product's own engine assembly (issue #2913), the harness and product defects it stopped on, and what the record now says. **What the round measured** (`evals/recursion_depth/results/wired-r0/`, `docs/reference/harness-round-log.md`, `docs/design/recursive-decomposition.md`): one gated cap-1 cell on the wired engine, every wiring finding read off the wire rather than off config. 7 of 42 criteria satisfied with the program LIVE; 1 of 8 claimed leaves survived; three merge attempts rejected by a roster reviewer; 50.5M tokens over 16 sessions. The verification register is answered per question in the recording's README, and `results/root-causes/README.md` carries a dated re-measured verdict per root cause: RC1 and RC4 are not the product's, RC3 is (the assembly not acting on findings, rather than nothing telling it), RC2 stands smaller. The round records none of the issue's five design outcomes: the operator stopped the sweep after attempt 5 rather than pay for the cap-1 repetitions, and the strategy question that decision raises is #2916. This PR does **not** close #2913: the cap-1 repetitions at the floor, the depth 2 to 4 go/no-go and the design decision the issue asks for are not settled by one cell. Of the issue's acceptance items, this PR delivers the wiring report with every finding named, the treatment variables recorded in the row, the register answered on the wire, the dossier figures re-derived, and the never-run probes each with a result or a named reason; it leaves open the repetition-floor decision (#2844) and the one-of-five design outcome, which #2916 now owns. **Harness fixes the round forced** (each one a stop point in the log): leaf tasks filed with the host so the product's own review path runs per leaf, with the leaf review typed in the product's vocabulary (`TaskStatus` + `CompletionOracleVerdict`) and a parked leaf recorded as parked; contract, merge and blind-pass sessions run as unfiled `IN_PROGRESS` transients; the host's coordination pair published so the peer-review gate is built and the smoke reads it; the embedder probed in preflight and routed through the provider's litellm key; the shell tool given the record store the build/test oracle reads; a run the host already settled never resumed; the smallest between-arm factor a bucket could detect reported beside the interval; a session digest script that renders a recorded session turn by turn. **Product fixes found on the wire**: a turn's mutating tool calls run in the order the model issued them, and a batch stops at a fatal error or a park (later calls are withheld with an error naming the pending approval rather than run into a state nobody vouched for); a structured tool argument a model sent as JSON text is decoded when the schema admits no string; a tool result reaches the model unless an HTML document hid something, and a rejected payload says so out loud (the DOM strip itself now lives in `tools/html_parse_strip.py`, reading an element's parent directly instead of through a ghost `getattr`); a reviewer's findings travel back with the reworked task rather than the summary alone (`engine/_review_rework_text.py`, shared by the oracle, red-team and completion gates); the credential detector reads a secret assignment's key as an identifier suffix and its value by shape; the L1 tool summary names only identifier-shaped parameters and refuses an overlap; a written file's parent directories are created and a symlinked write path below the workspace is refused; a no-credential connection on litellm's OpenAI-SDK route sends the placeholder key the SDK insists on; the shell tool says which command shape records a test run; the verdict schema names the field the rework brief reads and normalises a null-spelled command. **Docs**: the argument-decoding seam in `tools.md` and `conventions.md`, the no-credential placeholder in `providers.md`, the two round-log tracks stated in the harness log's intro, and the merge account corrected in the design page. Pre-reviewed by 18 agents; 41 findings addressed (one dropped as invalid). One baseline row shrunk with approval (`ghost_attribute_read_baseline.txt`, the retired `getparent` read). ## Test plan - [x] `uv run ruff check` / `ruff format` clean on `src/ tests/ evals/ scripts/` - [x] `make typecheck` clean (main and scripts daemons) - [x] Unit suite green (44186 passed) plus the affected packages re-run after the last batch (`tests/evals_spine`, `tests/unit/{engine,providers,memory,tools}`) - [x] Pre-push gate suite: 96/96 consolidated gates, markdownlint, vale, lychee, import-linter, audits - [ ] CI: integration, e2e, conformance and the full unit shards on the pushed branch (the local push deferred the unit run at 1311 affected files) Refs #2913 (not closed; see above). Strategy follow-up: #2916.
main
1 month ago

Latest Branches

CodSpeed Performance Gauge
0%
chore(main): release 0.9.5#2909
1 month ago
4ef2981
release-please--branches--main--components--synthorg
CodSpeed Performance Gauge
0%
research: re-measure the recursion-depth loop through the wired engine and land what the round found#2917
1 month ago
81c1e8a
research/remeasure-wired-loop
CodSpeed Performance Gauge
0%
© 2026 CodSpeed Technology
Home Terms Privacy Docs