2026-08-14

S1 results — vocabulary agreement (H8), run s1-2026-08-14

Version: 1.1 Date: 2026-08-14 Status: measured results. Recommendation 1 RATIFIED by the operator 2026-08-14 ("ratify B plus tooling role and rerun the panel") — see the S1b section at the end for the rerun; ddd-hexagonal-v2 (B + tooling) is the ratified candidate vocabulary. Lexicon.tsv entry and any renames remain separate, later steps. Design: study-programme-handoff.md S1 (with S2 folded in as a confirmation arm, ratified 2026-08-14).

Method (what actually ran)

+ name-prefix strata, every-kth), no hand-picking. Full sample and per-crate facts: s1-vocabulary-agreement-votes-2026-08-14.tsv (360 votes, git-tracked).

dependencies. Panelists saw ONLY these facts — no repo access (one Read of the payload file each; anti-hallucination by construction).

context each; no agent classified under both vocabularies (h8 handoff:85-93).

roles, registered in architectural-role-vocabulary.tsv this run) × prose paragraph / typed fields presentation. Deviation from the h8 design, disclosed: the prose arm's ANSWER format was constrained to parseable name: role lines (not fully free text) so votes could be tabulated; the presentation contrast lives in the input (paragraph vs typed fields) and the answer typing (lines vs enum-constrained JSON).

provenance_class='model', confidence = modal fraction); 8 agreement metrics → run_measurement run_id s1-2026-08-14; DB corpus_sourcecode on 127.0.0.1:5432. Git copy of every vote is the source of truth (ADR 0041).

Agreement rates (the H8 metric)

CellFull-agreement rateMean modal agreementFleiss' κ
A · prose0.8670.9560.889
A · typed0.7670.9110.787
B · prose0.7330.9110.774
B · typed0.8330.9440.857

Pooled: A κ̄ = 0.838, B κ̄ = 0.816. Note: κ is if anything flattered for B (11 categories vs 8 lowers chance agreement), so the raw read is A ties or narrowly leads; the gap (~0.02) is within noise for n=30, 3 raters.

The archetype split — confirmed, unanimously

The h8 handoff's predicted split held exactly:

B: aggregate 6/6 across both B cells.

B: pipeline 5/6.

One A-term absorbed two referents that B distinguishes at near-unanimous agreement: the state-machine-with-invariant went to aggregate, the staged acquisition ladder went to pipeline. This is the single clearest result of the run: whatever vocabulary wins, archetype should retire into two roles.

Structure arm (S2 confirmation) — an interaction, not a main effect

Typed presentation did NOT uniformly raise agreement: it raised B (+0.083 κ) and lowered A (−0.102 κ). The precedent-based WorkerPacket ratification stands (this arm gates nothing), but the claim "typed presentation always increases consistency" is NOT confirmed by this sample — reported as found.

Where each vocabulary failed (disagreement cases, all in the votes TSV)

tools-dev, tools-merge-resolve split infrastructure/application/pipeline in B-prose (0.67 each), while A's support was unanimous on all four in both arms. B needs a tooling/support role or a stated infrastructure convention.

where A's domain_type was unanimous — A's coarser type-class hides a real ambiguity B exposes. <!-- NAMING-ALLOW: this is a measured agreement result; the class names below are the labels the two vocabularies assigned, quoted as scored -->

capability 0.67) — auth is genuinely boundary-ambiguous and would go to unclassified under the never-guess rule; no panelist used unclassified anywhere, which is itself a panel-behavior finding.

30/30 votes across both B cells, vs A's component/domain_module split in A-typed (three crates at 0.67).

Recommendation (operator ratifies; this run decides nothing)

1. Treat the headline as a statistical tie. The standing rule (harness-design-notes-2026-08-12.md:44-47: established term wins ties) then points to B — but B as-run has a measured gap (no tools role). Candidate: adopt B + one added tooling role, and rerun the panel on the amended B before any lexicon entry. 2. Retire archetype into aggregate + pipeline regardless of which vocabulary wins — the only near-unanimous cross-vocabulary result. 3. identity-auth-class ambiguity: the panel never used unclassified; future runs should force a confidence field per item so "never guess past the evidence" is checkable, not aspirational.

Rule 14 ledger (operator instruction, this session: document every ad-hoc step)

Each of these ran ad-hoc this session and is repeatable, therefore owed to a registered Rust tool (target: tools-org-knowledge corpus subcommands writing to run_measurement):

1. Crate-facts extraction + stratified sampling — scratchpad/s1_extract.py (392 crates → 30-crate sample + two payload presentations). 2. Agreement computation (full-agreement, modal, Fleiss' κ) — scratchpad/s1_agreement.py over the votes TSV. 3. Consensus/metrics SQL emission — scratchpad/s1_to_sql.pycrate_classification + run_measurement INSERTs. 4. Earlier this session, same debt class: adverb/adjective frequency counts and the suffixed-sprint-doc file count (WORD-CHOICE.md §5) — belong in corpus audit as named metrics (harness-design-notes-2026-08-12.md:203-206 already states this meta-rule).

The panel invocation itself (12 subagents, prompts in the session transcript) is the step-runner's job once it exists (study-programme S6 dependency).

---

S1b — rerun on the ratified ddd-hexagonal-v2 (B + tooling), run s1b-2026-08-14

Operator ratified recommendation 1 same day; 6 fresh sonnet panelists (3 per arm), same payloads, same method. Votes: s1b-vocabulary-v2-votes-2026-08-14.tsv (180 votes). Storage: 30 typed-arm consensus rows (ddd-hexagonal-v2) + 4 metrics, run_id s1b-2026-08-14.

CellFull-agreement rateMean modal agreementFleiss' κ
v2 · prose0.833 (was B 0.733)0.9440.854 (+0.080)
v2 · typed0.900 (was B 0.833)0.9670.914 (+0.057)

The amended B now wins outright, not by tie-break.

put all five tool crates in tooling; in v1 those five crates produced most of B's disagreements.

(bounded_context vs value_object, both arms — the one stable residual), platform-api (application vs infrastructure — a composition root that is also substrate; candidate for unclassified or an operator row), operations-approval-workflow (prose split aggregate/capability; typed arm aggregate 3/3 — the identity/occurrent line needs the typed fields to hold), infrastructure-acquire (capability vs pipeline, typed), identity-auth and infrastructure-dom-reduce (one dissent each, prose only).

runs. Recommendation 3 (a forced per-item confidence field) stands.

State after S1b: ddd-hexagonal-v2 is the ratified vocabulary, measured best on both arms. Lexicon entry landed same day (operator ratification; see the ratified-vocabulary section of docs/reference/lexicon-policy.tsv — losing coinages banned where substring-safe, roles listed as KEEP terms). Still deliberately NOT done: a CHECK constraint on architectural_role and any crate renames — each a separate operator decision.

---

The harness, 2026-09-01 (sprint 4.31)

Recommendation 3 and rule-14 ledger items 1–3 above are now code, in tools_corpus::vocabulary_panel (agreement, votes, sample, store).

Ledger items 1–3 are discharged. scratchpad/s1_extract.py, s1_agreement.py and s1_to_sql.py were gone from disk by the time the replacement was written, so the twelve metric values published above are the only surviving statement of what they computed. They are therefore the Rust implementation's acceptance test: agreement reproduces all four s1-2026-08-14 cells and both s1b-2026-08-14 cells to three decimals from the git-tracked vote TSVs. The numbers in the tables above are unchanged and were not recomputed — they were reproduced.

Recommendation 3, restated by what was built. The recommendation reads the 0/540 abstention count as panel behavior needing a forced confidence field. The mechanism built takes a stronger reading: 0/540 is not evidence the panel was certain, it is evidence the instrument could not record uncertainty. The vote format was four columns — cell / agent / crate / role — with nowhere to say that the facts did not decide the question, so a panelist who felt uncertain had exactly one way to express it, which was to pick a role. The rule lived in the prompt; nothing in the file could hold it.

So confidence is a required fifth column with no default, votes below the run's threshold are coerced to unclassified by the loader rather than by the panelist's memory, and declared abstentions are counted separately from coerced ones. abstention_rate and coerced_abstention_rate are stored per cell, and the threshold is stored with the run. Both published runs now report an abstention rate, and it is 0.000 — the same fact as before, in a field where its absence would have been visible.

Re-scored through the CLI (corpus vocabulary-panel --score --no-confidence) 2026-09-01, all eight S1 cells and both S1b cells return the published values.

Kappa is reported Undefined when one category carried every vote, never 1.0: a wholly abstaining panel produces exactly that shape, and it is not perfect agreement.

The run path exists as of the same day: corpus vocabulary-panel --sample draws the sample and writes both arms' payloads; --score reads a returned vote file; --score --write --run-id <id> --cell <cell> stores metrics and classifications. A sample over this workspace draws 30 of 416 crates.

One correction the review caught before anything was stored. Consensus was keyed on the crate alone while its documentation claimed per-cell scope. S1's vote file spans four cells across two vocabularies, so scoring it whole pooled votes cast under the retired vocabulary with votes cast under the ratified one and returned a single row — which would then have been written stamped ddd-hexagonal-v2. Every test fixture used one cell, so nothing failed. Consensus is now keyed on (cell, crate) and the writer refuses rows spanning cells rather than storing whichever lands last, since crate_classification's key has no cell in it. Which arm to store is now asked for (--cell), as it was a decision S1 made silently by only storing the typed arm.

---

H8 run, 2026-09-01 — what the panel was being shown

Two panels ran on S1b's own 30 crates, taken from its published votes (h8-2026-09-01-sample.txt) so the sample could not confound the comparison. Same design as S1b: 3 sonnet raters per cell, fresh context, one payload file each, no repo access. Confidence threshold 0.5. Votes: h8-2026-09-01-votes.tsv (180) and h8-2026-09-01-structural-votes.tsv (90).

CellPayloadFull agreementMean modalFleiss' κAbstention
s1b V2·typed (2026-08-14)name, path, description, ≤5 deps0.9000.9670.914not recordable
h8 V2·typed (thin)same four facts0.8330.9330.8480.244
h8 V3·structural+ public surface, + inbound dependents0.9000.9670.9210.100

The 0/540 was the instrument, and this measures how much

Given a field for it, the same design on the same crates abstains on 24.4% of typed-arm votes. S1b recorded 0.0% on these exact crates six weeks earlier. The panel's certainty had not changed; the vote file had gained a column.

Most of that abstention was starvation, not ambiguity

Adding what the crate exposes and who depends on it halves abstention (0.244 → 0.100) and raises agreement to the best cell of the whole study (κ 0.921, above S1b's 0.914). Consensus rows that came back unclassified fell from 8 of 30 to 3 of 30.

Both directions matter. More evidence did not merely raise confidence — it raised agreement and left a residue of genuine ambiguity that no amount of description would have settled. The thin payload's abstentions were mostly the panel reporting that it had been shown a name and asked about a structure.

That is the finding the four-column format could not have produced. A run with no abstention field reports 0.0% whether the evidence was overwhelming or absent, and S1/S1b read that silence as agreement.

Why the structural payload is not "more access"

The roles in ddd-hexagonal-v2 are structural claims: an aggregate owns invariants over children, a port is an interface something implements, infrastructure is what others stand on. Inbound dependents and public surface are the evidence those claims are about; a one-line description is not. The panelist still reads one file, still cannot reach the source, and the payload still names no source file — S1's anti-hallucination property is unchanged.

A limit of the store, found by running it twice

crate_classification's natural key is (crate, vocabulary_version, provenance_class), so two model runs over the same crates collide: the structural run's rows replaced the thin run's, and only provenance records which file produced what survives. Metrics are safe (run_measurement is keyed by run_id, and both runs are there), and the vote TSVs are the real record — but the database holds one model classification per crate, not a history of them. Recorded as a limit, not fixed: adding a run dimension to that key is an operator decision about what the table is for.

State

crate_classification held 489 rows over 426 distinct crates, 381 of them unclassified, before this run. 30 crates now carry a structural-arm model classification, 3 of them unclassified because the panel said so and the instrument could record it.

All research