diff --git a/.abcd/development/brief/04-surfaces/07-memory.md b/.abcd/development/brief/04-surfaces/07-memory.md index 3de141df5..73e36585d 100644 --- a/.abcd/development/brief/04-surfaces/07-memory.md +++ b/.abcd/development/brief/04-surfaces/07-memory.md @@ -148,5 +148,5 @@ meaningful. - [`05-internals/07-memory.md`](../05-internals/07-memory.md): substrate spec - [`05-internals/09-provenance-substrate.md`](../05-internals/09-provenance-substrate.md): provenance and licence subsystem -- [`../../intents/planned/itd-36-memory-unification.md`](../../intents/planned/itd-36-memory-unification.md): the full intent spec with acceptance criteria +- [`../../intents/shipped/itd-36-memory-unification.md`](../../intents/shipped/itd-36-memory-unification.md): the full intent spec with acceptance criteria - [`research/related-work.md § Karpathy LLM Wiki`](../../research/related-work.md#karpathy-llm-wiki--pattern-source-for-abcdmemory): pattern source diff --git a/.abcd/development/brief/06-delivery/03-out-of-scope.md b/.abcd/development/brief/06-delivery/03-out-of-scope.md index 185fa1c6a..48717f3d7 100644 --- a/.abcd/development/brief/06-delivery/03-out-of-scope.md +++ b/.abcd/development/brief/06-delivery/03-out-of-scope.md @@ -65,7 +65,6 @@ gate. That is what keeps "not hand-counted" true after the day it was written. - `itd-128` — One canonical YAML scalar resolver: every decoder delegates to one exported frontmatter helper - `itd-129` — Forge mirror as an opt-in adapter: one-way mirror-out, schema'd forge id, self-healing closures - `itd-78` — Intent-dependency graph: what to build first, even when it is something small -- `itd-82` — `abcd drain` fixes the open issues that need no decision through the implement machinery and hands the rest back by kind (a user moment to an intent draft, a trust rule to a decision record, the rest flagged with their home); all by default, capped by flag - `itd-83` — The review bar fires by itself, in every repo abcd manages - `itd-85` — Read-only repo-conformance audit - `itd-86` — Cold-reading surface: abcd reads its own design documents as a stranger would @@ -125,7 +124,7 @@ gate. That is what keeps "not hand-counted" true after the day it was written. - `itd-2609090746414083` — A lifeboat packs from a lab session home, the throwaway experiment's intention, harvest and bundle, with the same coverage honesty as a repository (refines itd-88 and adr-35; the non-git half, sequenced after the lab verb family) - `itd-2609180517121254` — every payload a host hands back from a delegated step names the model that produced it and the number of agents that ran, and the ingesting verb refuses one that does not - `itd-2609201916056194` — a delegated agent runs through a command-line model runner the operator chose (claude CLI, opencode/openrouter); the opt-in cli oracle rung -- `itd-2609211116005482` — `abcd build next` picks the readiest planned intent itself, writes on the record why and what would show the pick wrong, and hands it to the implement machinery; one by default, all by flag +- `itd-2609211913453478` — one glossary page maps the record families (intent, spec, bundle, phase, batch, issue, roadmap, release) and how they relate, and answers whether a phase is still the sequencing layer **Later-phase items with no intent id.** These four were written into the brief diff --git a/.abcd/development/decisions/adrs/0007-grill-skill-and-glossary.md b/.abcd/development/decisions/adrs/0007-grill-skill-and-glossary.md index f4a738e02..6ed6cc0cc 100644 --- a/.abcd/development/decisions/adrs/0007-grill-skill-and-glossary.md +++ b/.abcd/development/decisions/adrs/0007-grill-skill-and-glossary.md @@ -153,6 +153,6 @@ are written to PRD frontmatter. After freeze: ## Related -- [itd-27](../../intents/planned/itd-27-grill-skill-and-glossary.md) — source intent +- [itd-27](../../intents/superseded/itd-27-grill-skill-and-glossary.md) — source intent - [05-internals/01-agents.md](../../brief/05-internals/01-agents.md) — `intent-fidelity-reviewer` auditor contract - tests/fixtures/grill/ — fixture corpus diff --git a/.abcd/development/intents/disciplines/itd-37-modification-grammar.md b/.abcd/development/intents/disciplines/itd-37-modification-grammar.md index d5c359827..944625725 100644 --- a/.abcd/development/intents/disciplines/itd-37-modification-grammar.md +++ b/.abcd/development/intents/disciplines/itd-37-modification-grammar.md @@ -22,7 +22,7 @@ Every spec carries a `## Modification Grammar` section with three sub-headings, A fourth axis — **`### Ripple`** — captures the per-spec-external concerns absorbed from idea-3 (systems thinking) at adversarial review (chat `idea-3-itd-38-adversaria-31A06A`, MAJOR_RETHINK outcome): vocabulary delta (HARD-enforced via the [vocabulary-registration requirement in `02-constraints/04-naming.md`](../../brief/02-constraints/04-naming.md)), surface delta (new commands / sub-verbs / agents / disciplines), coupling delta (what this spec newly depends on; what newly depends on it). The `Ripple` axis stays under itd-37 (same retrieval key as modification grammar — domain) rather than spawning a separate discipline. -At spec completion, `principle-distiller` (the role-extended curator from [itd-36](../planned/itd-36-memory-unification.md)) extracts the `## Modification Grammar` section into two memory pages per [`05-internals/07-memory.md`](../../brief/05-internals/07-memory.md): +At spec completion, `principle-distiller` (the role-extended curator from [itd-36](../shipped/itd-36-memory-unification.md)) extracts the `## Modification Grammar` section into two memory pages per [`05-internals/07-memory.md`](../../brief/05-internals/07-memory.md): - **`spec_modification_grammar_.md`** — append-only, per-spec. Source class `spec_modification_grammar`. Lifecycle: append-only per [`05-internals/04-universal-patterns.md § 8`](../../brief/05-internals/04-universal-patterns.md#8-artefact-lifecycle-taxonomy). - **`modification_grammar_.md`** — compounding-curated, per-domain. `principle-distiller` runs a curator-pass that synthesises across the per-spec pages. Lifecycle: compounding-curated per the same taxonomy. @@ -111,7 +111,7 @@ _Empty. Populated by `intent-auditor` Role 1 (single-document fidelity per itd-1 - [`research/related-work.md § Naur 1985`](../../research/related-work.md#naur-1985--programming-as-theory-building) — full prior-art comparison. - [`itd-1-acceptance-gates.md`](itd-1-acceptance-gates.md) — companion discipline; this discipline's acceptance criteria conform to its Given-When-Then shape. - [`itd-5-prompt-quality-additions.md`](itd-5-prompt-quality-additions.md) — companion discipline; disciplines stack at three (itd-1 + itd-5 + itd-37). -- [`../planned/itd-36-memory-unification.md`](../planned/itd-36-memory-unification.md) — companion intent (standalone); ships alongside itd-37; provides the substrate for `spec_modification_grammar` and `modification_grammar` page classes. +- [`../shipped/itd-36-memory-unification.md`](../shipped/itd-36-memory-unification.md) — companion intent (standalone); ships alongside itd-37; provides the substrate for `spec_modification_grammar` and `modification_grammar` page classes. - [`05-internals/07-memory.md`](../../brief/05-internals/07-memory.md) — substrate spec for the curator agent's two-page-class extraction. - [`02-constraints/04-naming.md`](../../brief/02-constraints/04-naming.md) — vocabulary-registration requirement (HARD) and bare-command-as-render discipline; both companion rules to this discipline's `Ripple` axis. - [`01-product/03-mental-model.md § The Naurian gap`](../../brief/01-product/03-mental-model.md) — the framing this discipline closes. diff --git a/.abcd/development/intents/drafts/itd-2609211116005482-abcd-build-next-picks-the-readiest-planned-intent-itself-wri.md b/.abcd/development/intents/drafts/itd-2609211116005482-abcd-build-next-picks-the-readiest-planned-intent-itself-wri.md deleted file mode 100644 index 4f60e2f44..000000000 --- a/.abcd/development/intents/drafts/itd-2609211116005482-abcd-build-next-picks-the-readiest-planned-intent-itself-wri.md +++ /dev/null @@ -1,92 +0,0 @@ ---- -id: itd-2609211116005482 -slug: abcd-build-next-picks-the-readiest-planned-intent-itself-wri -spec_id: null -kind: standalone -suggested_kind: null -reclassification_history: [] -builds_on: [itd-2609201916151817, itd-2609201925079472] -related_intents: [itd-78, itd-82, itd-2609201916151817] -severity: major -impact: additive -origin: researcher-authored -production_mode: hand-written ---- - -# `abcd build next` picks the readiest planned intent itself, and writes down why - -## Press Release - -> **`abcd build next` chooses the next intent to build, says why on the record, and builds it.** It reads every planned intent whose gate says READY, scores each on facts its own record holds — how clear its acceptance criteria are, whether there is an obvious way to test it, how small the change it asks for looks — takes the readiest, oldest among equals, and writes onto that intent a grounds entry marked as the run's: who else was considered, why each lost, and what would show the pick wrong. Then it hands the intent to the same machinery `abcd build ` uses and takes it to delivered. One intent per run, because "next" means one; `--until-empty` or `--max ` keeps it going under the pace rule, which is how `abcd drain` reads by default, because "drain" means all. -> -> "I used to open the run file every morning and pick by feel, then forget by Friday why Tuesday's pick was that one," said a product thinker who had planned forty-eight intents and could build one a day. "Now the run picks, the intent says why in its own record, and when a pick turns out wrong the reason is there to argue with." - -## Why This Matters - -Forty-eight planned intents were READY in this repository on 2026-09-20, and the file that routes them into batches was written by hand, ordered by a person's sense of what comes first, with no line saying why. Every autonomous coding platform surveyed on 2026-09-21 (Copilot's coding agent, OpenHands, Jules, Devin, Cursor's cloud agents, Claude Code and Codex routines) is triggered by a human assigning a task or a label; none publishes a policy for choosing among a backlog, so a run that is meant to be unattended still begins with a human's choice. Two things the same survey found do have evidence: readiness is predictable from the record alone (static features of the task predict an agent's success at 0.76 to 0.84 AUC before any code is written), and small, tidy work merges far more often than large features (84.7 per cent for clean-ups against 64.5 per cent for features across 878 agent pull requests in one large repository). A pick made from those facts, written where a person can later check it, is what turns the run file's hand-ordered batches into something the binary does. - -The reason has to be checkable. Agent-written rationales, when audited, were factually wrong more than half the time in one repository's decision log; the survey found no evidence that recording a reason improves later review, only that an unrecorded one cannot be reviewed at all. So the record the run writes is computed, not composed: the candidate set, each candidate's score and the rule that placed it, and a falsifier the lane's own outcome can meet. - -## Mechanism - -We expect a pick made from the record's own facts, with its reason written as computed comparison and a falsifier, to choose intents that deliver at a higher rate than hand-ordering and to be correctable when it does not, because the facts that predict an agent's success are in the acceptance criteria, the test path and the footprint before a lane starts, and a written falsifier is what a later reader checks the outcome against; shown wrong if the readiest-scored intents deliver no more often than oldest-first over a run of ten or more picks, or if the falsifier on a failed pick did not name what actually failed. - -## Scope Conditions - -- Holds for a repository abcd manages whose planned intents carry acceptance criteria in Given-When-Then form and a linked spec, so the readiness score has fields to read; an intent with a placeholder criteria section is not a candidate. -- Holds at the scale the run file recorded: tens of READY intents, not thousands, so the whole candidate set is scored on every run and written into the reason. -- Holds where the intent store's grounds section accepts an entry the run writes and marks as its own beside a human's; the human's entry is never edited or displaced. -- The ordering is a heuristic with the evidence stated above; where `blocked_by` and `builds_on` edges exist, an intent with an unshipped blocker is not a candidate, and the edges are not otherwise ranked on (the dependency-graph draft, `itd-78`, is where a derived priority would come from, and this intent does not build it). - -## What's In Scope - -- **The candidate set**: every intent in `planned/` whose `abcd intent ready` verdict is READY, that is not held, that no peer holds, and whose `blocked_by` names nothing unshipped. -- **The score**, from the record alone, each component named in the reason: criteria clarity (count and shape of the Given-When-Then bullets), a test path (a spec that names the packages and the tests, or criteria a test can hold), the expected footprint (packages and surfaces the spec names), and the record's age; the weights are a declared configuration, not a prompt. -- **The reason**, written onto the chosen intent's `## Grounds` section as one entry marked as the run's (`pursued:` prefixed with the run's identity and date), carrying the candidates considered with their scores, the rule that placed the winner, the runner-up and why it lost, and the falsifier: the lane outcome that would show this pick wrong (more than the pace rule's fix rounds, a reviewer raising a design question, a footprint outside what the spec named). -- **The hand-over**: the chosen intent goes to the implement machinery exactly as `abcd build ` would send it; nothing in the lane differs. -- **The run count**: one by default; `--max ` and `--until-empty` continue under the pace rule, each further pick written the same way. -- **Refusals**: no READY intent (says so, writes nothing); every READY intent held or blocked (lists them); a tie the score cannot break beyond age (takes the oldest and says the tie was broken by age). - -## What's Out of Scope - -- The build itself, the lane, the validators, the landing: `itd-2609201916151817`. -- The pace, the ceiling and the pause: `itd-2609201925079472`. -- A derived priority over the dependency graph: `itd-78`; this intent filters on edges and does not rank on them. -- The issue ledger: `abcd drain` (`itd-82`) is the run over issues, and it hands over to the same machinery. -- Any estimate the agent makes of its own confidence as an input to the pick. - -## Decisions - -Ruled by the product thinker on 2026-09-21, in the interview that filed this draft: - -1. **The verb is `build`, for people; `implement` is the machinery, for agents.** `abcd build ` and `abcd build next` are what a person types; both hand over to `abcd implement`, whose step-level words (`step`, `receipt`) a driving host calls. The planned implement intent is amended to say so; which verbs the command list shows a person and which it shows an agent is a separate record (`iss-2609211119023345`). -2. **Readiest first, oldest among equals.** Not oldest-first, not most-important-first; the evidence for readiness as the predictor is the reason, and the order is declared a heuristic. -3. **The reason lives on the intent itself**, as a grounds entry marked as the run's, so opening the intent shows why it was picked and what would prove the pick wrong. -4. **One by default, many by flag**, and the defaults differ from `drain` on purpose: "next" reads one and "drain" reads all; the flags are the same on both. -5. **The reason is computed, never composed.** Candidates, scores, placing rule, falsifier; no prose the agent writes about its own choice. - -## Open Questions - -- **Whether a run-made grounds entry needs its own token.** `pursued:` with the run's identity in the text is the shape assumed here; if the vocabulary gains a token for a machine's conjecture, this intent takes it. -- **The score's weights.** Declared in configuration with a bundled default; the default is set from the first ten picks' outcomes, not chosen up front. - -## Acceptance Criteria - -- **Given** planned intents of which some are READY, some held, some blocked and some not READY, **when** `abcd build next` runs, **then** only the READY, unheld, unblocked ones are candidates, and the refusal for an empty candidate set names each excluded intent and the check that excluded it, writing nothing. -- **Given** two or more candidates, **when** the pick is made, **then** the chosen intent's `## Grounds` section gains exactly one entry marked as the run's, naming every candidate with its score, the rule that placed the winner, the runner-up and why it lost, and a falsifier, and no existing entry is changed. -- **Given** candidates whose scores tie, **when** the pick is made, **then** the oldest is taken and the entry says the tie was broken by age. -- **Given** a pick, **when** the lane starts, **then** it is the lane `abcd build ` would start for that intent, with the same brief, worktree, validators and landing. -- **Given** the default run count, **when** the picked intent is delivered or its lane stops, **then** the run reports and exits without picking again; **given** `--max ` or `--until-empty`, **then** it picks again under the pace rule and writes each further reason the same way. -- **Given** a delivered pick, **when** the falsifier's condition is met by the lane's outcome (a design question raised by a reviewer, a footprint outside the spec, fix rounds past the pace rule's count), **then** the run record names the pick as falsified and the intent's entry is not edited. -- **Given** `--json`, **when** any of the above renders, **then** the candidate set, scores, the reason and every refusal are in the payload. - -## Typed Links - -- **builds_on `itd-2609201916151817`** (the implement machinery, `abcd build `): the pick hands over to it and adds nothing to the lane. -- **builds_on `itd-2609201925079472`** (the pace): `--until-empty` and `--max` run under it. -- **refines `itd-78`** (the dependency graph): this intent filters on `blocked_by` and does not derive a priority; that draft is where a derived one would come from. -- **refines `itd-82`** (`abcd drain`): the sibling run over the issue ledger; the two share the flags and the hand-over and differ in their default count by the meaning of their names. - -## Audit Notes - -_Empty. Populated by intent-auditor when intent moves to shipped/._ diff --git a/.abcd/development/intents/drafts/itd-2609211913453478-one-page-in-the-glossary-maps-abcd-s-record-families-and-how.md b/.abcd/development/intents/drafts/itd-2609211913453478-one-page-in-the-glossary-maps-abcd-s-record-families-and-how.md new file mode 100644 index 000000000..a1c7849fd --- /dev/null +++ b/.abcd/development/intents/drafts/itd-2609211913453478-one-page-in-the-glossary-maps-abcd-s-record-families-and-how.md @@ -0,0 +1,57 @@ +--- +id: itd-2609211913453478 +slug: one-page-in-the-glossary-maps-abcd-s-record-families-and-how +spec_id: null +kind: standalone +suggested_kind: null +reclassification_history: [] +builds_on: [itd-34, itd-78] +related_intents: [itd-24, itd-42, itd-172] +severity: minor +impact: additive +origin: researcher-authored +production_mode: hand-written +--- + +# One page maps the record families and how they relate, and says whether a phase still is one + +## Press Release + +> One page in the glossary maps abcd's record families and how they relate: intent, spec, bundle, phase, batch, issue, roadmap and release each get one definition and one line saying what they group, what groups them, and which verb moves them. A product thinker reads it in five minutes and knows which word to use; bundle and batch gain entries; and the page answers whether a phase is still the sequencing layer now that bundles group intents that ship together and a run works in batches. + +## Why This Matters + +On 2026-09-21 the product thinker asked, mid-interview, whether abcd has an intent that sorts out its own vocabulary (intents, issues, phases, bundles, roadmap, specs) and how they relate, and whether phases are still wanted now that bundles exist. The answer was: a glossary exists (`.abcd/development/brief/glossary/`, with phase, intent, spec, roadmap, record and ledger defined in prose and adr-9 making the phase the product layer between the brief and the intent), but nothing draws the map, `bundle` has no entry, and the autonomous run adds a third grouping, the batch. Three ways of grouping intents is one too many for a vocabulary a product thinker is meant to hold; the evidence that phases have thinned is that no spec carries a phase anchor and phase membership is recorded editorially in the phase document. + +## Mechanism + +> _Prompted (the claim-recording gradient): why the authors expect this to work, as a falsifiable "we expect X because Y" — not the outcome restated. Replace this line with the claim, or with the exact token `None stated.` alone on its line to record the claim as considered and declined._ + +## Scope Conditions + +> _Required (the claim-recording gradient): the population, platform, scale, or assumptions this claim holds under, one per top-level bullet — `abcd intent plan` stamps each with a persistent identity. Replace this line with those bullets, or with the exact token `None stated.` alone on its line._ + +## What's In Scope + +- **One page**, `glossary/core/record-families.md` or its equivalent, with one row per family: the noun, one definition, what it groups, what groups it, its lifecycle folders, and the verb that moves it; the page is the single source the other entries point at. +- **Two new entries**: `bundle` (intents sharing one spec because they ship as one change; itd-34) and `batch` (the order an autonomous run takes lanes in; the run file). +- **The phase question, answered on the page**: whether a phase remains the sequencing layer, is folded into bundles, or is replaced by the run's batches; the answer is a decision recorded before the page ships, and the retrospective intent (itd-24) and the bundle rule in itd-34 follow it. +- **The lint**: every glossary entry's `not_to_be_confused_with` names a family on the page; a family named in a record's frontmatter that the page does not define is a finding. + +## What's Out of Scope + +- Renaming any family or any folder. +- The dependency graph between intents (itd-78). + +## Acceptance Criteria + +> _Required (the itd-1 discipline): add at least one Given-When-Then bullet describing the verifiable bar for "shipped" before this draft can be planned._ + +## Open Questions + +- **Do phases stay?** Kept as the sequencing layer ending in a milestone; folded into bundles (a bundle is the only grouping, and the roadmap orders bundles); or replaced by the run's batches (the run file's routing is the roadmap). Each answer changes itd-24 and itd-34. +- **Where the page lives**: one glossary page, or the brief's mental-model chapter with the glossary pointing at it. + +## Audit Notes + +_Empty. Populated by intent-auditor when intent moves to shipped/._ diff --git a/.abcd/development/intents/drafts/itd-82-drain-ledger-triage.md b/.abcd/development/intents/drafts/itd-82-drain-ledger-triage.md deleted file mode 100644 index 0c9c5eb84..000000000 --- a/.abcd/development/intents/drafts/itd-82-drain-ledger-triage.md +++ /dev/null @@ -1,116 +0,0 @@ ---- -id: itd-82 -slug: drain-ledger-triage -kind: standalone -suggested_kind: standalone -bundle: null -spec_id: null -reclassification_history: [] -builds_on: [itd-4, itd-46, itd-2609201916151817, itd-2609201925079472] -related_intents: [itd-2609211116005482, itd-84] -related_adrs: [adr-25, adr-27] -severity: major -impact: additive ---- - -# `abcd drain` fixes the issues that need no decision, and hands the rest back by kind - -## Press Release - -> **`abcd drain` works the open issue ledger unattended: it fixes what needs no decision, and routes what does to the place a person decides it.** It reads every open issue that nothing blocks, and lets through only the ones the rules say a machine may touch alone: a stated remedy, one area of the code, no change a user would see, no trust or safety boundary, not security, not `major` or `critical`. Each of those goes to the machinery `abcd build ` uses, one lane, one pull request, the record resolved in the same change, the run never approving its own pull request. An issue that needs a decision is handed back by kind: one that would change what a user sees becomes an intent draft seeded from the issue, for a person to plan; one that turns on a rule about trust or safety is flagged as needing a decision record, with the question stated; anything else is flagged with the home the decision belongs in. A lane that discovers a decision inside a "mechanical" fix stops, discards its work, and hands the issue back the same way rather than guessing. Blocked issues are skipped with the blocker named. It runs until nothing eligible is left, paced by the pace rule, because "drain" means all; `--max ` caps a run. In the morning there are three lists: pull requests to review, drafts to plan, decisions to make. -> -> "I used to point the loop at the ledger and then hover, because the moment it hit something design-shaped it would either stall or, worse, decide it," said a product thinker who ran the first ledger drains by hand. "Now the mechanical ones arrive as pull requests, the ones that are mine arrive as drafts or as questions, and the one time a lane found a decision halfway through, it stopped and told me instead of finishing." - -## Why This Matters - -The pilot run of 2026-09-20 fixed a cluster of five captured issues through lanes, reviewers and a merge queue, and the part a person did by hand at every step was the sorting: which capture a lane may take alone, which is a design decision in disguise, which is blocked. That sorting is the load-bearing, human-shaped piece, and it is the one a person should not run every night. Its failure is one-directional: a machine that decides a thing needs no decision, and then makes one. The evidence gathered on 2026-09-21 is plain about it: models detect their own ambiguity badly (the best reaches 89 per cent with a 3 per cent false-positive rate only under strong prompting; weaker ones flag 93 per cent of clear tasks as ambiguous), two teams that let a triage pass auto-dispatch a coding agent reverted it after unwanted pull requests, and the largest dataset of agent pull requests (878 over ten months) shows clean-ups and tests merging at 85 and 76 per cent while features and performance work merge at 65 and 55. Industrial fix pipelines land 15 to 25 per cent of what they attempt; most lanes end without a merge, and that has to be a normal ending, not a failure. - -The record already has the sorting rule. The four-piece routing every filing goes through (capability to an intent, trust rule to a decision record, stance to a principle, plumbing to the brief) is what a decision-shaped issue needs applied to it, and a capture is not an intent: an intent is a user moment, and an issue whose decision is a rule about trust is not one. So the hand-back routes by kind rather than promoting everything into drafts a person then has to sort again. - -## Mechanism - -We expect a rule-first eligibility filter, a lane that hands back on discovering a decision, and a hand-back routed by kind to drain the mechanical part of a ledger without a person and without a decision being made by a machine, because the rules that predict a safe unattended fix are fields the capture already carries (a stated remedy, category, severity, the blocked-by edges) and the hand-back is a write the verbs already make (`capture promote` for a user moment, a flag for the rest); shown wrong if a merged drain pull request is later found to have decided something a person should have, or if the hand-back queue is mostly issues the filter should have let through. - -## Scope Conditions - -- Holds for a repository abcd manages with a merge queue or branch protection on the default branch, so a pull request the run opens has a path to the default branch that the run does not control. -- Holds for issues captured through `abcd capture` with a category, a severity and a stated remedy; a record without a remedy is not eligible and is listed as such. -- Holds at the ledger sizes seen so far (hundreds of open records), where every eligible issue is classified on every run; a ledger of thousands would want the classification recorded once and reused, which is an open question below. -- The classifier is rules first; where a host-delegated judgement is used for the residual call, its errors run in one direction by construction (an uncertain issue is handed back, never fixed), and that judgement is recorded with the disposition. - -## What's In Scope - -- **Eligibility, rules first**: open, nothing unshipped in `blocked_by`, a stated remedy, category not `security` and not one of the decision categories (`architectural-insight`, `future-work-seed`), severity `minor` or `nitpick`, the remedy naming one package or one surface, no user-visible change and no trust-boundary change on the run's reading of the remedy; `major` and `critical` go to the hand-back by default. A host-delegated judgement (adr-25) is allowed only for the residual "is this genuinely self-contained" call and only to hand back, never to let through. -- **The lane**: the implement machinery (`itd-2609201916151817`), one lane per issue, the record resolved with its commit in the same change, one pull request per issue with the repository's merge rule applied, the run's identity never an approver, nothing pushed to a pull request after its merge is armed. -- **Reproduce, then fix**: the lane arms a detector that fails before the fix and passes after; the adversarial reviewer stays the oracle and the detector is evidence for it, not the gate. -- **The hand-back, routed by kind**: a user-moment issue is promoted to an intent draft (`capture promote`, seeded from the issue, the back-link written); a trust or safety rule is flagged as needing a decision record with the question stated, and nothing is minted; anything else is flagged with the proposed home named. The issue is never edited; the lane's partial work is discarded. -- **The mid-lane stop**: a lane that meets a decision (a reviewer raising a design question, a remedy that turns out to change what a user sees, a second package the remedy did not name) stops, discards, and hands back with the reason, as a first-class outcome. -- **Order**: category and footprint first, smallest first within a category, oldest among equals; the run's summary states the rule. -- **Ceilings**: the pace rule bounds the run; `--max ` caps issues attempted; the pull requests opened, the promotions made and the spend are counted in the summary, and a cap hit is named. -- **The summary**: every issue's disposition (pull request number, draft id, flag with home, skipped with blocker, stopped with reason), machine-readable and re-runnable; "no pull request, reason logged" is a normal terminal state. - -## What's Out of Scope - -- The lane, the validators and the landing: `itd-2609201916151817`; the pace: `itd-2609201925079472`. -- Picking among intents: `abcd build next` (`itd-2609211116005482`), the sibling run. -- Merging pull requests, planning promoted drafts, writing decision records: human gates by design. -- Any change to the ledger schema beyond consuming `list`, `resolve` and `promote`. -- Any estimate the agent makes of its own confidence as a reason to let an issue through. - -## Decisions - -Ruled by the product thinker on 2026-09-21, in the interview that revived this draft: - -1. **This record is the drain run**, updated for the implement machinery; no new record, and the verb keeps its name, `abcd drain`, which is what a person types. It hands over to `abcd implement`, the machinery an agent calls. -2. **The hand-back routes by kind**: user moments become intent drafts; trust rules are flagged for a decision record; the rest are flagged with their home. Not everything promoted, not everything merely listed. -3. **All by default, capped by flag**: "drain" reads all and "next" reads one, so the two runs' defaults differ on purpose and their flags are the same. -4. **The rule for "needs no decision" is a recorded rule.** It decides what a machine may touch alone, so it is written as a decision record and a brief invariant before this path ships, the way `--auto-plan` owes its record; until then the rules above are the draft's statement of it. - -## Prior Art - -- **`itd-2609201916151817`** (the implement machinery): the lane, the validators and the landing this run dispatches onto; it supersedes `itd-29`, which this draft first named. -- **`itd-46`** / **spc-30**: owns `capture promote `, the issue-to-intent elevator the user-moment hand-back uses; shipped, so the earlier cut-A stub is gone. -- **`itd-4`**: the ledger substrate (`list`, `resolve`) the run reads and mutates. -- **`itd-84`**: the four-piece routing the hand-back applies to a decision-shaped issue. -- **adr-25** (host-delegated judgement): the residual classifier call rides the host and only hands back. -- **adr-27** (receipt-gated autonomous runs): the lane's review discipline. - -## SOTA - -Surveyed 2026-09-21 by an independent research pass, primary sources where they exist: - -- **Triage that dispatches is the failure mode.** GitHub's own reference triage workflows label and comment only, never assign or close; two independent repositories that let a triage label auto-dispatch a coding agent reverted it after unwanted pull requests. Adopted: the run's eligibility rules are the only thing that lets an issue through, and the hand-back's only write is a promotion for a user moment. -- **The classifier is rules first.** Copilot's published guidance keeps ambiguous, open-ended, security, PII, authentication and production-critical work for people and assigns bugs, tests, docs and debt to the agent; repositories converged independently on `agent-ready` / `needs-design` labels. Ambiguity self-detection by models is unreliable (Ambig-SWE, ICLR 2026), so the residual judgement may only hand back. -- **Order by category and footprint, not severity.** The dotnet/runtime dataset (878 agent pull requests, ten months): clean-up 84.7, testing 75.6, bug fix 69.4, feature 64.5, performance 54.5 per cent merged; one platform's own fleet report had refactoring and performance queues at zero closure. Adopted for this run because its scope is already bounded to no-decision issues. -- **Detector-first is contested.** Impact-aware test selection cut regressions from 6.08 to 1.82 per cent while a bare "write a failing test first" instruction worsened them to 9.94 (TDAD, 2026); 29.6 per cent of plausible patches on SWE-bench behave differently from the reference. Adopted as reproduce-then-fix with the reviewer as the oracle. -- **The safety envelope is GitHub's two designs combined**: one branch, one pull request per task, the requester cannot approve, a hard session cap (Copilot's cloud agent); read-only agent job with writes through a scoped step, `max: 1` pull request, daily credit and per-user rate limits, `stop-after` (gh-aw). Translated: per-run caps, the run's identity is not an approver, the merge queue is the path to `main`. -- **Expect most lanes not to merge.** Google's sanitizer-fix pipeline landed 15 per cent of candidates; Meta's test-repair agent 25.5 per cent over three months with a judge before human review. "No pull request, reason logged" is a normal ending. -- **Not adopted**: WSJF and cost-of-delay (no outcome evidence, needs human estimates); the agent's self-reported confidence as a gate (vendor feature, no published threshold); fully automatic readiness promotion by a triage model. - -## Open Questions - -- **Whether the classification is recorded on the issue.** Re-derived every run (simple, and the rules may change) or written once as a disposition field the next run reuses and a person can override (durable, and `capture` could set it at file time). Leaning durable at ledger sizes past a few hundred. -- **The decision-record flag's home.** A line in the summary only, or a marker on the issue the person finds without the summary; the marker is a write to a record the run otherwise never edits. -- **Whether `major` may ever be let through.** The default sends it to the hand-back; a `--allow-major` flag is the obvious opt-in, and whether a person wants it is not yet known. - -## Acceptance Criteria - -- **Given** an open ledger holding a `minor` issue with a stated one-package remedy, a `major` issue, an issue whose remedy would change a command's output, an issue in category `security`, and an issue blocked by an open record, **when** `abcd drain` runs, **then** only the first goes to a lane; the `major`, the user-visible and the `security` issues are handed back with the kind named; the blocked one is skipped naming its blocker; and every open issue receives exactly one recorded disposition. -- **Given** an eligible issue, **when** its lane runs, **then** a detector is armed and watched to fail before the fix and pass after, the record is resolved with the fix commit in the same change, one pull request is opened with the repository's merge rule applied, and nothing is pushed to it afterwards. -- **Given** a decision-shaped issue that describes a user moment, **when** it is handed back, **then** an intent draft exists seeded from the issue with the back-link written, the issue is unchanged, and no design or adoption verdict is recorded. -- **Given** a decision-shaped issue that turns on a trust or safety rule, **when** it is handed back, **then** it is flagged as needing a decision record with the question stated, and nothing is minted. -- **Given** a lane whose reviewer raises a design question, or whose fix reaches a second package or a user-visible change the remedy did not name, **when** the lane sees it, **then** the lane stops, its work is discarded, the issue is handed back with the reason, and the summary records the stop as its own outcome. -- **Given** the default run, **when** eligible issues remain at a window's end, **then** the run pauses under the pace rule and resumes; **given** `--max `, **then** it stops at the cap and names it. -- **Given** a completed or paused run, **when** its summary is read, **then** every issue's disposition is there (pull request number, draft id, flag with home, skip with blocker, stop with reason) with the ordering rule and every cap, in text and in `--json`, and a run whose lanes all ended without a merge exits 0 with that stated. -- **Given** a crafted issue id, **when** the run resolves it, **then** it is validated by shape and refused otherwise, and no file outside the ledger directories is read, written or moved. - -## Typed Links - -- **builds_on `itd-2609201916151817`** (the implement machinery) and **`itd-2609201925079472`** (the pace): the lane and its clock. -- **builds_on `itd-46`** (`capture promote`) and **`itd-4`** (the ledger): the writes the hand-back and the resolve make. -- **refines `itd-84`** (decompose before filing): the routing applied to a decision-shaped issue at hand-back. -- **refines `itd-2609211116005482`** (`abcd build next`): the sibling run over intents; the two share the flags and the hand-over and differ in their default count by the meaning of their names. - -## Audit Notes - -_Empty. Populated by intent-auditor when intent moves to shipped/._ diff --git a/.abcd/development/intents/planned/itd-24-reflect-command.md b/.abcd/development/intents/planned/itd-24-reflect-command.md index 0e5dcbc5f..56b410500 100644 --- a/.abcd/development/intents/planned/itd-24-reflect-command.md +++ b/.abcd/development/intents/planned/itd-24-reflect-command.md @@ -1,7 +1,7 @@ --- id: itd-24 slug: reflect-command -spec_id: null +spec_id: spc-2609211751376504 kind: bundle-member bundle: spc-83-operator-surfaces suggested_kind: null @@ -14,6 +14,7 @@ prd_path: null prd_grandfathered: true severity: minor builds_on: [itd-27] +impact: additive --- # Completed Phases Get A Retrospective @@ -34,6 +35,10 @@ The legacy `~/.claude/templates/retrospective.md.template` had the right prompt A phase audit and a phase retrospective are **distinct activities**, the same split `intent-fidelity-reviewer` Role 1 draws at the intent grain — one grain up. The audit asks *did the phase's `## Phase Acceptance` pass* (a per-bullet verdict). The retrospective asks *what did we learn* (transferable insight, interview-driven). `/abcd:reflect` does not replace the audit — it **consumes** it: the audit's verdicts are the seed material the retrospective interview opens from. +## Mechanism + +We expect the value of a retrospective to be in the lifeboat: a new project that starts from an old one's lessons avoids a repeat, because the lessons arrive with the work rather than being re-derived; shown wrong if projects embarked from a lifeboat carrying retrospectives never open a surfaced lesson. + ## What's In Scope - **`/abcd:reflect `** — retrospective for a completed phase, and the command's *only* argument form. Examples: `/abcd:reflect phase-1-substrate`, `/abcd:reflect phase-2-ahoy`. Per-intent reflection is out of scope (see below) — `/abcd:reflect` operates at the phase grain only. @@ -46,7 +51,7 @@ A phase audit and a phase retrospective are **distinct activities**, the same sp - Lessons learned (transferable insights, framed for future-you) - Decisions made (architectural / design choices crystallised during the phase) - Metrics (intents shipped, audit-note severity distribution, time-to-ship if measurable) -- **Output**: `.abcd/retrospectives//README.md` — a peer of `.abcd/intents/` and `.abcd/logbook/`, committed as part of the phase's permanent record. +- **Output**: `.abcd/development/retrospectives//README.md` — a peer of `.abcd/development/intents/`, committed as part of the phase's permanent record. - **Lifeboat integration**: `/abcd:disembark to ` packs *all* of the voyage's phase retrospectives into the lifeboat — the full reflection arc travels. `/abcd:embark from ` surfaces predecessor retrospectives during the press-release interview ("here's what the previous voyage learned about X — does that apply here?"). - **Reference back to intents and the audit**: the retrospective links to the phase doc, to the intents the phase bundled, and to the phase audit; per-intent reviewer notes are referenced (not duplicated). - **`reflection-composer` agent** — runs the interview, drafts the structured output, asks clarifying questions when answers feel thin. @@ -69,14 +74,22 @@ None stated. - **Given** a completed phase with no phase audit yet recorded, **when** the persona runs `/abcd:reflect `, **then** the command reports the missing audit and offers to run the phase-fidelity-reviewer inline before continuing into the retrospective. - **Given** a phase doc that exists but has no spec carrying its `phase:` anchor, **when** the persona runs `/abcd:reflect `, **then** the command refuses with "no specs anchored to `` — nothing shipped to reflect on" and writes no output. - **Given** a draft retrospective with thin answers (e.g. "what went well: it worked"), **when** the agent drafts the output, **then** the agent surfaces the thinness as a clarifying question rather than committing the thin answer. -- **Given** the same repo's lifeboat is then packed via `/abcd:disembark to `, **when** the lifeboat is inspected, **then** every `.abcd/retrospectives//README.md` the voyage produced is included in the lifeboat artefact. -- **Given** a target repo embarked from a lifeboat that includes retrospectives, **when** `/abcd:embark from ` runs the press-release interview, **then** the persona is shown predecessor retrospective lessons and asked which apply to the new voyage. +- **Given** the same repo's lifeboat is then packed via `/abcd:disembark to `, **when** the lifeboat is inspected, **then** every `.abcd/development/retrospectives//README.md` the voyage produced is included in the lifeboat artefact. +- **Given** a target repo embarked from a lifeboat that includes retrospectives, **when** `/abcd:embark from ` runs the press-release interview, **then** the persona is shown the few predecessor lessons ranked most like the new voyage's brief, with the rest as a list, and asked which apply. - **Given** an attempt to reflect on a phase whose specs are not all closed, **when** `/abcd:reflect ` runs, **then** the command warns the persona, lists the open specs anchored to that phase, and asks for confirmation to proceed anyway. +- **Given** the last piece of work anchored to a phase closes, **when** that close completes, **then** abcd says once that a retrospective for the phase is owed and names the command, and says nothing further about it. + +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: + +1. **Nudge once.** When a phase's last piece of work closes, abcd says once that a retrospective is owed; it is not repeated and it is not a gate. +2. **A ranked few on embark.** Predecessor lessons most like the new voyage's brief are shown; the rest are a list opened on request. +3. **Layout.** The retrospective lives under the durable record tier, `.abcd/development/retrospectives//README.md`; the paths this record was written against predate the three-tier layout and are read as that. ## Open Questions -- **Reflection cadence** — is there a soft prompt to encourage running it (e.g. when the last spec of a phase closes), or fully on-demand? -- **Lifeboat surfacing on embark** — how intrusive? The lifeboat carries all phase retrospectives; how should embark present them — show every one, or rank by relevance to the new voyage? Risk of "previous-voyage lessons" feeling like noise on a brand-new voyage. +_None open; decisions 1 and 2 settle the two this record carried (the reflection cadence and the lifeboat surfacing on embark)._ ## Blocking Dependency @@ -125,3 +138,7 @@ spec" as a bundle (`kind: bundle-member` + shared `bundle: spc-83-operator-surfa require. Bundle member by delivery relationship, not a scope change. This intent keeps its real grill linkage (`grill_session_id`); GR002 is handled via `prd_grandfathered`. Full record in the spec's process-exception note. + +## Grounds + +- pursued: nine phases have closed with their lessons living only in session handovers; we expect the first retrospectives to hold what the handovers do not, and a new voyage to open a carried lesson; shown wrong if the first retrospectives say nothing the handovers did not, or if no embarked project opens one diff --git a/.abcd/development/intents/planned/itd-2609201916151817-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md b/.abcd/development/intents/planned/itd-2609201916151817-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md index 42835475e..cec8ff3b1 100644 --- a/.abcd/development/intents/planned/itd-2609201916151817-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md +++ b/.abcd/development/intents/planned/itd-2609201916151817-one-verb-takes-a-single-intent-from-ready-to-delivered-witho.md @@ -6,7 +6,7 @@ kind: standalone suggested_kind: null reclassification_history: [] builds_on: [itd-2609201916056194, itd-2609091014076309, itd-2609091416295622] -supersedes: [itd-29] +supersedes: [itd-29, itd-58] related_intents: [itd-29, itd-50, itd-2, itd-2609170822093401] severity: major impact: additive @@ -71,10 +71,13 @@ Asked and answered on 2026-09-20: Ruled by the product thinker on 2026-09-21, in the interview that filed `itd-2609211116005482` (`abcd build next`) and revived `itd-82` (`abcd drain`): 8. **`build` for people, `implement` for the machinery.** `abcd build ` is what a person types and is the verb this record's press release names; `abcd build next` and `abcd drain` are the two customised runs that hand over to the same loop. `abcd implement` is that loop, with the step-level words a driving host calls renamed from `next` to `step` (`implement step` returns the brief, `implement receipt ` advances) so that `next` is free to mean the pick. Which verbs the command list shows a person and which it shows an agent is its own record (`iss-2609211119023345`). +9. **Only the loop writes a verdict** (itd-58 folded in). A validator's verdict is recorded by the loop from the validator's own return, into the state file, before the advance is decided; a lane has no write to it, and a receipt carrying a verdict the loop did not record is refused at the advance, naming the receipt. + +10. **The loop takes an issue as well as an intent** (ruled 2026-09-21 for `abcd drain`, itd-82). `abcd implement` is keyed by `itd-N` or `iss-N`; for an issue the brief is the record and its `remedy:` field, the checks before it starts are drain's eligibility rule, the validators are the same, and the landing is `capture resolve --commit` in the lane's change; the lane report carries a `handback:` field the loop reads as a first-class outcome. `abcd build ` is the person's form. ## Open Questions -_None open; decisions 5 to 7 settle the interview's questions._ +_None open; decisions 5 to 7 settle the interview's questions, and decisions 8 to 10 record the rulings of 2026-09-21 (the verb's name, the unforgeable verdict, the issue key)._ ## Acceptance Criteria @@ -88,6 +91,8 @@ _None open; decisions 5 to 7 settle the interview's questions._ - **Given** a host with no configured runner, **when** `abcd implement step` returns a brief, **then** the host is told which agent to start and where the receipt goes, and the loop advances only on `abcd implement receipt`. - **Given** a configured runner and the process driver opted in, **when** the loop reaches a lane, **then** it starts the lane through the runner itself, and the run record names the runner. - **Given** a completed run, **when** the run record is read, **then** it names every lane, receipt, reviewer verdict, the model each runner reported, and the transcripts captured into the history store. +- **Given** an eligible issue, **when** `abcd build ` runs, **then** the lane's brief is the record and its remedy, its landing resolves the record with the fix commit, and a `handback:` in the lane report stops the lane and is reported as its own outcome. +- **Given** a validator has returned, **when** the loop records its verdict, **then** the verdict in the state file is the one the loop parsed from the validator's return, and a lane report carrying a verdict the loop did not record is refused at the advance, naming the report; a real SHIP recorded by the loop lets the advance proceed. - **Given** any refusal, **when** it is rendered, **then** it names the step, the reason and the remedy, in text and in `--json`. ## Typed Links diff --git a/.abcd/development/intents/planned/itd-2609211116005482-abcd-build-next-picks-the-readiest-planned-intent-itself-wri.md b/.abcd/development/intents/planned/itd-2609211116005482-abcd-build-next-picks-the-readiest-planned-intent-itself-wri.md new file mode 100644 index 000000000..7203c4800 --- /dev/null +++ b/.abcd/development/intents/planned/itd-2609211116005482-abcd-build-next-picks-the-readiest-planned-intent-itself-wri.md @@ -0,0 +1,111 @@ +--- +id: itd-2609211116005482 +slug: abcd-build-next-picks-the-readiest-planned-intent-itself-wri +spec_id: spc-2609212015048113 +kind: standalone +suggested_kind: null +reclassification_history: [] +builds_on: [itd-2609201916151817, itd-2609201925079472, itd-50] +related_intents: [itd-78, itd-82] +severity: major +impact: additive +origin: researcher-authored +production_mode: hand-written +--- + +# `abcd build next` picks the readiest planned intent itself, and writes down why + +## Press Release + +> **`abcd build next` chooses the next intent to build, says why on the record, and builds it.** It reads every planned intent that passes every refusal check `abcd build ` would run (READY, no open question, no unanswered claim, not held, not blocked, no peer holding it), scores each on facts its own record holds — how clear its acceptance criteria are, whether there is an obvious way to test it, how small the change it asks for looks — takes the readiest, oldest among equals, and writes onto that intent a grounds entry marked as the run's: who else was considered, why each lost, and what would show the pick wrong. The entry is the lane's first commit, so the reason reaches `main` with the work and stays on the branch as history if the lane is discarded. Then it hands the intent to the same machinery `abcd build ` uses and takes it to delivered. One intent per run, because "next" means one; `--until-empty` or `--max ` keeps it going under the pace rule, which is how `abcd drain` reads by default, because "drain" means all. +> +> "I used to open the run file every morning and pick by feel, then forget by Friday why Tuesday's pick was that one," said a product thinker who had planned forty-eight intents and could build one a day. "Now the run picks, the intent says why in its own record, and when a pick turns out wrong the reason is there to argue with." + +## Why This Matters + +Forty-eight planned intents were READY in this repository on 2026-09-20, and the file that routes them into batches was written by hand, ordered by a person's sense of what comes first, with no line saying why. Every autonomous coding platform surveyed on 2026-09-21 (Copilot's coding agent, OpenHands, Jules, Devin, Cursor's cloud agents, Claude Code and Codex routines) is triggered by a human assigning a task or a label; none publishes a policy for choosing among a backlog, so a run that is meant to be unattended still begins with a human's choice. Two things the same survey found do have evidence: readiness is predictable from the record alone (static features of the task predict an agent's success at 0.76 to 0.84 AUC before any code is written), and small, tidy work merges far more often than large features (84.7 per cent for clean-ups against 64.5 per cent for features across 878 agent pull requests in one large repository). A pick made from those facts, written where a person can later check it, is what turns the run file's hand-ordered batches into something the binary does. + +The reason has to be checkable. Agent-written rationales, when audited, were factually wrong more than half the time in one repository's decision log; the survey found no evidence that recording a reason improves later review, only that an unrecorded one cannot be reviewed at all. So the record the run writes is computed, not composed: the candidate set, each candidate's score and the rule that placed it, and a falsifier the lane's own outcome can meet. + +## Mechanism + +We expect a pick made from the record's own facts, with its reason written as computed comparison and a falsifier, to choose intents that deliver and to be correctable when one does not, because the facts that predict an agent's success are in the acceptance criteria, the test path and the footprint before a lane starts, and a written falsifier is what a later reader checks the outcome against; shown wrong if the falsifier on a failed pick did not name what actually failed, or if picks scored readiest end unachievable (itd-50's verdict) more often than not over ten or more picks. + +## Scope Conditions + +- Holds for a repository abcd manages whose planned intents carry acceptance criteria in Given-When-Then form and a linked spec, so the readiness score has fields to read; an intent with a placeholder criteria section is not a candidate. +- Holds where the intent store's grounds section accepts an appended entry the run writes beside a human's; the human's entry is never edited or displaced. +- Holds where `blocked_by` and `builds_on` edges are declared on the records; an intent with an unshipped blocker is not a candidate, and the edges are not otherwise ranked on. + +## What's In Scope + +- **The candidate set**: every intent in `planned/` that passes every refusal check `abcd build ` runs before it starts (READY, no open question, no unanswered claim section, not held, no peer holding it, nothing unshipped in `blocked_by`); the pick never writes onto an intent the build would then refuse. +- **The score**, from the record alone, each component named in the reason: criteria clarity (count and shape of the Given-When-Then bullets), a test path (the spec's `## Footprint` section naming the tests, absent today on every spec), and the expected footprint (that section's package list); three parts at equal weight, the bundled default, revisable after ten picks; age is not a part of the score and breaks ties only. Until a spec carries the section, its test-path and footprint parts read zero and the reason says so. +- **The reason**, written onto the chosen intent's `## Grounds` section as one entry marked as the run's (`pursued:` prefixed with the run's identity and date), carrying the candidates considered with their scores, the rule that placed the winner, the runner-up and why it lost, and the falsifier: a lane outcome the state file can meet (fix rounds past the pace rule's count, or the lane handing back as unachievable under itd-50's stage). The marker ("picked by run on ") goes inside the entry's text after the `pursued:` token, since the token vocabulary is closed. The entry is committed as the lane's first commit on the lane branch. +- **The hand-over**: the chosen intent goes to the implement machinery as `abcd build ` would send it, with one difference: after the loop's branch step and before the implementer starts, the pick step commits the grounds entry as a record-only commit at the branch base, and the loop's receipt verifier counts the implementer's commits from after it. +- **The gate**: `abcd intent ready`'s grounds row skips a run-marked entry when it names the most recent conjecture, so the human's entry stays the one it reports. +- **The run record**: names each pick, its score table, and whether its falsifier was met. +- **The run count**: one by default; `--max ` and `--until-empty` continue under the pace rule, each further pick written the same way. +- **Refusals**: no candidate (says so, names each excluded intent and the check that excluded it, writes nothing); a tie the score cannot break beyond age (takes the oldest and says the tie was broken by age). +- **The spec template**: gains a `## Footprint` section (packages, tests) the score reads; `intent plan` seeds it empty. +- **The command page**: `commands/intent.md` (or the build page once it exists) says what `next` does, its two flags and its refusals. + +## What's Out of Scope + +- The build itself, the lane, the validators, the landing: `itd-2609201916151817`. +- The pace, the ceiling and the pause: `itd-2609201925079472`. +- A derived priority over the dependency graph: `itd-78`; this intent filters on edges and does not rank on them. +- The issue ledger: `abcd drain` (`itd-82`) is the run over issues, and it hands over to the same machinery. +- Any estimate the agent makes of its own confidence as an input to the pick. + +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that filed this draft: + +1. **The verb is `build`, for people; `implement` is the machinery, for agents.** `abcd build ` and `abcd build next` are what a person types; both hand over to `abcd implement`, whose step-level words (`step`, `receipt`) a driving host calls. The planned implement intent is amended to say so; which verbs the command list shows a person and which it shows an agent is a separate record (`iss-2609211119023345`). +2. **Readiest first, oldest among equals.** Not oldest-first, not most-important-first; the evidence for readiness as the predictor is the reason, and the order is declared a heuristic. +3. **The reason lives on the intent itself**, as a grounds entry marked as the run's, so opening the intent shows why it was picked and what would prove the pick wrong. +4. **One by default, many by flag**, and the defaults differ from `drain` on purpose: "next" reads one and "drain" reads all; `--max ` and `--until-empty` are the two flags both share. +5. **The reason is computed, never composed.** Candidates, scores, placing rule, falsifier; no prose the agent writes about its own choice. +6. **The reason is the lane's first commit** on the lane branch (ruled 2026-09-21 on the design review's finding): a record-only commit the pick step makes at the branch base before the implementer starts, so it reaches `main` with the work and stays on the branch if the lane is discarded; the receipt verifier counts the implementer's commits from after it. +8. **The ordering is declared a heuristic** with the evidence in Why This Matters; it is revised on the record, never silently. +7. **Equal weights at first**, declared as the bundled default and revisable after ten picks; age breaks ties and is not scored. + +## Open Questions + +- **Whether a run-made grounds entry needs its own token.** `pursued:` with the run's identity in the text is the shape assumed here; if the vocabulary gains a token for a machine's conjecture, this intent takes it. + +## Acceptance Criteria + +- **Given** planned intents of which some pass every refusal check of `abcd build ` and some fail one (not READY, an open question, an unanswered claim, held, blocked, a peer holding it), **when** `abcd build next` runs, **then** only those that pass are candidates, and the refusal for an empty candidate set names each excluded intent and the check that excluded it, writing nothing. +- **Given** two or more candidates, **when** the pick is made, **then** the chosen intent's `## Grounds` section gains exactly one `pursued:` entry whose text opens with the run's marker, naming every candidate with its score, the rule that placed the winner, the runner-up and why it lost, and a falsifier the state file can meet; no existing entry is changed; the entry is the lane branch's first commit; and `abcd intent ready` still reports the human's entry as the most recent conjecture. +- **Given** candidates whose scores tie, **when** the pick is made, **then** the oldest is taken and the entry says the tie was broken by age. +- **Given** a pick, **when** the lane starts, **then** it is the lane `abcd build ` would start for that intent, with the same brief, worktree, validators and landing, and one record-only commit at the branch base carrying the grounds entry, which the receipt verifier does not count as the implementer's work. +- **Given** the default run count, **when** the picked intent is delivered or its lane stops, **then** the run reports and exits without picking again; **given** `--max ` or `--until-empty`, **then** it picks again under the pace rule and writes each further reason the same way. +- **Given** a pick, **when** the lane's state file shows fix rounds past the pace rule's count or an unachievable hand-back, **then** the run record names the pick as falsified and the intent's entry is not edited. +- **Given** `abcd intent plan` mints a spec, **when** the spec stub is read, **then** it carries an empty `## Footprint` section (packages, tests) for the author to fill. +- **Given** a spec without a `## Footprint` section, **when** the score is computed, **then** its test-path and footprint parts read zero and the reason says the spec carries no footprint. +- **Given** the command page, **when** it is read, **then** it states what `next` does, `--max` and `--until-empty`, and every refusal. +- **Given** `--json`, **when** any of the above renders, **then** the candidate set, scores, the reason and every refusal are in the payload. + +## Prior Art + +- The run file of 2026-09-20 (`abcd-autonomous-run.md`), whose hand-ordered batches this pick replaces. +- `itd-78` (the dependency graph): the derived priority this intent filters on and does not build. +- The state-of-the-art pass of 2026-09-21 (in the record-families of the platforms surveyed and the psychometrics and dotnet/runtime evidence cited above). + +## Typed Links + +- **builds_on `itd-2609201916151817`** (the implement machinery, `abcd build `): the pick hands over to it and adds nothing to the lane. +- **builds_on `itd-2609201925079472`** (the pace): `--until-empty` and `--max` run under it. +- **refines `itd-78`** (the dependency graph): this intent filters on `blocked_by` and does not derive a priority; that draft is where a derived one would come from. +- **related `itd-82`** (`abcd drain`): the sibling run over the issue ledger; the two share `--max` and `--until-empty` and the hand-over to `abcd implement`, and differ in their default count by the meaning of their names. +- **builds_on `itd-50`** (the audit-driven fix round): the unachievable verdict is one of the pick's falsifiers. + +## Audit Notes + +_Empty. Populated by intent-auditor when intent moves to shipped/._ + +## Grounds + +- pursued: the autonomous run's routing table is a person's hand-written file today; we expect the pick to be what replaces it after the run; shown wrong if the next run still starts from a hand-written table diff --git a/.abcd/development/intents/planned/itd-28-rp-reviews-into-flow.md b/.abcd/development/intents/planned/itd-28-rp-reviews-into-flow.md index e809bbb48..2dd9e9469 100644 --- a/.abcd/development/intents/planned/itd-28-rp-reviews-into-flow.md +++ b/.abcd/development/intents/planned/itd-28-rp-reviews-into-flow.md @@ -1,7 +1,7 @@ --- id: itd-28 slug: rp-reviews-into-flow -spec_id: null +spec_id: spc-2609211854150455 kind: standalone suggested_kind: null reclassification_history: [] @@ -20,10 +20,14 @@ glossary_terms_used: - core/transport builds_on: [itd-1] severity: major +impact: additive --- # Spec-Tied Reviews Live Next To The Spec They Reviewed +> **Re-scoped on 2026-09-21** by the product thinker: the record keeps the pin and the staleness view and drops its own review store, its two-stage redaction and its pre-commit verifier. The record already holds dated review folders under `.abcd/work/reviews/` and gate receipts keyed by the commit they gate; what is missing is that every review names the commit it read and the status board says how stale each has become. The scrub rides the scanner that already exists. The press release and scope below are read through this paragraph and the Decisions section. + + ## Press Release > **abcd lands every spec-tied review next to the spec it reviewed, in the native review store, pinned to the commit it reviewed, with a sanitisation pass before commit.** When a plan-review or impl-review finishes — whichever oracle adapter produced it — an abcd-side post-processor captures the review receipt via the native receipt contract and lands a canonical JSON sidecar (`review.json`) and rendered Markdown view (`review.md`) in a per-review directory at `.abcd/reviews//--/`. The sidecar carries the full sanitised review body, structured findings, and a `review_of_commit` SHA pin so future agents can detect when a review's findings have gone stale. Raw transcripts land in a per-review `raw/` subdirectory (gitignored). A two-stage redaction scheme applies: Stage 1 is a write-time sanitiser (AWS, GCP, Azure, Cloudflare, GitHub, Anthropic, OpenAI, Stripe, Slack, JWT, PEM) before any file is written; Stage 2 is a detect-and-block commit gate run through the scanner seam — native patterns by default, `gitleaks protect --staged --redact=100` as the stronger opt-in adapter when the binary is present, the engine always reported — that rejects commits if secrets survive Stage 1. An on-demand `reviews-index --spec ` regenerates `INDEX.md` + `INDEX.json`; CI runs `--check` mode to catch drift without ever writing to the working tree. @@ -85,30 +89,32 @@ The unscoped-transport sweep is adapter-scoped and runs only when an oracle adap - **Unscoped oracle transport storage** — covered by the adapter-scoped sweep into `.abcd/work/reviews/`. - **`Reviewed-by:` git trailer auto-injection on implementation commits** — out of scope (nice-to-have bidirectional linkage). +## Mechanism + +We expect a visible staleness count to make a stale review get re-run before a release rather than trusted, because the count turns "is this still valid" from a question nobody asks into a row on the board everyone sees; shown wrong if a release cut still cites reviews past the threshold with nobody re-running them. + ## Scope Conditions None stated. ## Acceptance Criteria -- **Given** a persona runs a plan-review for `spc-X` end-to-end via any oracle adapter, **when** the review completes, **then** a per-review directory lands at `.abcd/reviews/spc-X/--/` containing `review.json` (all required fields populated, `verdict` ∈ `{SHIP, NEEDS_WORK, MAJOR_RETHINK}`, non-empty `body_markdown` and `reviewed_files`) and `review.md` (mechanically rendered from `review.json`). -- **Given** the post-processor runs twice on the same receipt, **when** both invocations complete, **then** there is exactly one per-review directory in `.abcd/reviews/spc-X/` (idempotent; the second invocation is a no-op). -- **Given** the post-processor is killed mid-write (`kill -9`), **when** the persona inspects the working tree, **then** no `.tmp` or partial files are visible to git. -- **Given** 5 concurrent post-processor invocations on the same spec, **when** they complete, **then** 5 distinct sequence numbers exist (no collisions). -- **Given** the persona sets `ABCD_REVIEW_POSTPROCESS=0` and runs a plan-review, **when** the review completes, **then** the post-processor exits 0 with no side effects and the review remains only in the producing oracle adapter's raw output. -- **Given** a staged `.abcd/reviews/**` file containing a multi-cloud secret (AWS access key, fine-grained GitHub PAT, Anthropic key, JWT, or PEM private key) that was NOT caught by Stage 1, **when** the pre-commit hook runs the Stage-2 scan (native engine, or gitleaks when present), **then** the commit is blocked with the finding path/line/rule reported (never the raw secret value). -- **Given** the gitleaks binary is absent, **when** the pre-commit hook runs the Stage-2 scan, **then** the native engine runs and the hook output names the engine that ran — a downgrade is never silent. -- **Given** a committed `.abcd/reviews/**/*.md` file containing `AKIAIOSFODNN7EXAMPLE`, **when** the pre-commit hook runs, **then** the EXAMPLE-allowlisted value is NOT redacted. -- **Given** a review file with a `review_of_commit` SHA that fails `git rev-parse --verify`, **when** the pre-commit hook runs, **then** the commit is rejected with a clear error message. -- **Given** a `review.json` with `body_markdown` exceeding `body_max_bytes`, **when** the pre-commit verifier runs, **then** the commit is rejected with guidance (the post-processor truncates automatically; this acceptance captures the case where someone manually edits the sidecar to violate the cap). -- **Given** a CI run on a PR touching `.abcd/reviews/**`, **when** `reviews-index --check --all` runs, **then** the workflow fails on drift with the exact remediation command and never writes back to the branch. -- **Given** the brief is updated, **when** a contributor reads `05-internals/02-adapters.md` and `05-internals/03-configuration.md`, **then** they find explicit text describing the two-store carve-out (spec-tied via the native review pipeline; unscoped via a configured oracle adapter). -- **Given** the README is updated, **when** a contributor reads the Acknowledgements section, **then** they find explicit citations of `gitleaks` (Apache-2.0) and `REPPL/abcdZero` F-075 / F-037 prior art. +- **Given** a review folder under `.abcd/work/reviews/` is filed by any of abcd's own review paths, **when** it is written, **then** its summary carries `review_of_commit: ` written by the tool, and the record lint refuses a review folder without one (a folder that predates the rule is named as legacy, not refused). +- **Given** the bare `abcd` status board renders, **when** review folders exist, **then** each is listed with the spec or scope it reviewed and the number of commits the default branch has moved since its `review_of_commit`, and one past twenty is flagged for a re-run. +- **Given** a review of a spec, **when** it is filed, **then** its folder name carries the spec's id, so a reader finds it from the spec. +- **Given** a review body carries a secret, **when** it is committed, **then** the scanner the repository already runs refuses it; this record builds no second scrubber. +- **Given** `--json`, **when** the board renders, **then** the staleness rows and the threshold are in the payload. + +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: + +1. **Pin and staleness only.** The review store, the two-stage redaction and the pre-commit verifier this record first described are dropped: the dated review folders and the gate receipts are the store, and the scanner is the scrub. +2. **The staleness view is on the status board**, flagged past twenty commits. ## Open Questions -- **Hook coverage for hosts without a Stop-hook equivalent**: standalone invocation of the post-processor with `--from-receipt --spec ` is the documented fallback. Should a per-host wrapper also ship? -- **Stale review threshold for `staleness` column**: currently `_commits` since `review_of_commit`. Should we add a "danger" threshold (e.g., `>20_commits` shown red)? Deferred polish. +_None open; the hook-coverage and the danger-threshold questions this record carried fall away with the store, and the threshold is decision 2._ ## Audit Notes @@ -124,3 +130,7 @@ _Empty. Populated by intent-fidelity-reviewer when intent moves to shipped/._ [bias]: https://arxiv.org/html/2603.18740v1 "Confirmation Bias in `LLM`-Assisted Security Code Review" [liip]: https://www.liip.ch/en/blog/preventing-context-pollution-for-%61i-agents "Liip — Preventing Context Pollution for `AI` Agents" + +## Grounds + +- pursued: the release gate now demands review receipts, and a receipt with no commit named cannot be judged fresh; we expect the next cut to read the staleness rows before it trusts a receipt; shown wrong if the next cut never asks how old a receipt is diff --git a/.abcd/development/intents/planned/itd-34-three-intent-kinds.md b/.abcd/development/intents/planned/itd-34-three-intent-kinds.md index dd77bafe6..70ed32464 100644 --- a/.abcd/development/intents/planned/itd-34-three-intent-kinds.md +++ b/.abcd/development/intents/planned/itd-34-three-intent-kinds.md @@ -4,14 +4,18 @@ slug: three-intent-kinds kind: standalone suggested_kind: standalone bundle: null -spec_id: null +spec_id: spc-2609211859391533 reclassification_history: [] builds_on: [itd-1] severity: major +impact: additive --- # Intents Promote In Three Flavours — Standalone, Bundle-Member, Discipline +> **Re-scoped on 2026-09-21** by the product thinker: the three kinds, the `disciplines/` shelf and the two-way supersession link this record described have shipped through other records; what remains, and what this record now plans, is the bundle command and the `reclassify` verb. The press release and scope below are read through this paragraph and the Decisions section. + + ## Press Release > **abcd ships three intent kinds, each with its own lifecycle path through `/abcd:intent plan`.** A standalone intent maps 1:1 to a flow-next `spec` — the default for ~60% of intents, and the only path the prior brief recognised. A bundle-member intent shares a `spec` with its bundle-mates: the persona runs `/abcd:intent plan itd-A itd-B`, abcd creates one shared `spec` with both intents listed, and both files move from `drafts/` to `planned/` together. A discipline intent has no persona moment of its own — it's a cross-cutting rule (every `spec` must carry acceptance criteria, every agent `prompt` must declare a `version`) that lives in `disciplines/` instead of `drafts/`, never gets its own `spec`, but applies to every other `spec` as an inherited acceptance gate. Capture stays format-neutral; the binding `kind` field is set at plan time, with an `oracle` classifier writing an advisory hint at capture time. The discipline shape is `voyage`-agnostic — an application `voyage` (e.g., a macOS app under abcd) produces its own disciplines around privacy-impact review or accessibility passes with the same lifecycle as abcd's own discipline intents. @@ -62,7 +66,6 @@ For late kind changes (a standalone intent realised to be a bundle-member; a dra All intent files gain four new fields: ```yaml -kind: null # set by /abcd:intent plan: standalone | bundle-member | discipline suggested_kind: null # advisory, written by capture-time LLM classifier; can be ignored bundle: null # for kind: bundle-member, the bundle ID reclassification_history: [] # appended to by /abcd:intent reclassify @@ -129,27 +132,21 @@ This mirrors the lint-code namespace pattern in [`05-internals/06-lint.md`](../. - **A fourth kind** for some hypothetical case the three kinds don't cover. If a fourth kind is needed, it lands as a future intent, not this one. **Lineage update (itd-44, spc-56):** the standing-infrastructure-choice case the three kinds don't fit landed as [itd-44](../drafts/itd-44-fourth-intent-kind-decision.md) — but *not* as a fourth persisted `kind`. itd-44 adds a fourth capture-time *verdict*, `decision`, that routes a confirmed standing choice into the existing ADR store (`adr-N`); the persisted `kind` enum here stays **three-valued**, and there is deliberately no `intents/decisions/` lifecycle. - **Migrating disciplines into the brief itself.** Disciplines live in `intents/disciplines/` so they have the same lifecycle, lint, and frontmatter discipline as other intents. Folding them into the brief would lose the lifecycle handling. The brief *describes* the discipline kind in its mental-model and intent-surface sections; the discipline files themselves stay under `intents/`. +## Mechanism + +We expect a `reclassify` verb to end the hand moves that leave one-way supersession links and mismatched kinds, because the two-way write and the shelf move become one command whose refusals name what is missing; shown wrong if superseded records still arrive with a missing back-link after it ships. + ## Scope Conditions None stated. ## Acceptance Criteria -> _BDD format, per the [itd-1 discipline](../disciplines/itd-1-acceptance-gates.md). These gates are checked by `intent-fidelity-reviewer`'s single-document role when this intent moves to `shipped/`._ - -- **Given** a draft intent at `drafts/itd-N-foo.md` with `kind: null`, **when** the user runs `/abcd:intent plan itd-N`, **then** abcd proposes a kind based on the intent body and cross-references, the user confirms or overrides via interactive prompt, and the chosen kind is written to the intent's frontmatter as binding. -- **Given** two draft intents at `drafts/itd-A-foo.md` and `drafts/itd-B-bar.md` whose press releases reference each other in their References sections, **when** the user runs `/abcd:intent plan itd-A itd-B`, **then** `/flow-next:plan` is called once with both intents as joint input; the resulting spec (`spc-N`) in the native spec store has `intent: [itd-A, itd-B]`; each intent's frontmatter has `kind: bundle-member`, `bundle: `, and `epic_id: spc-N`; both files move from `drafts/` to `planned/`. -- **Given** two draft intents `itd-A` and `itd-B` scoped to different phases, **when** the user runs `/abcd:intent plan itd-A itd-B` (multi-arg, kind=bundle-member), **then** abcd hard-blocks promotion with lint code `IL011`, cites the bundle invariant in [`brief/04-surfaces/05-intent.md § 1`](../../brief/04-surfaces/05-intent.md#1-intent-ids-kinds-and-lifecycle), and offers the user two resolutions: (a) re-scope all proposed members into the same phase, or (b) downgrade one or more members to `kind: standalone` and re-run plan. Worked example in the wild: the `intent-capture-discipline` bundle (itd-27 + itd-30) was retired on 2026-05-07 for exactly this reason — both intents now ship as `kind: standalone`. -- **Given** a bundle's shared spec closes in the native spec store, **when** the lifecycle hook fires, **then** all bundle-member intents (those with the matching `bundle:` field) move from `planned/` to `shipped/` together AND `intent-fidelity-reviewer`'s single-document role runs once per member against the same delivered reality, producing per-member Audit Notes. -- **Given** a draft intent that the user wants to promote as a discipline, **when** the user runs `/abcd:intent plan itd-N` and selects `kind: discipline` at the prompt, **then** NO `/flow-next:plan` call is made; the intent's acceptance gates are registered in `.abcd/disciplines/.json`; the file moves from `drafts/` to `disciplines/`. The intent's frontmatter has NO `status` field — the directory IS the state. -- **Given** a discipline-kind intent in `disciplines/`, **when** any subsequent flow-next spec is plan-reviewed, **then** the plan-review verifies the spec's acceptance criteria are compatible with all active disciplines' rules — drift between spec acceptance and discipline gate is flagged. -- **Given** a shipped intent that the user later realises is fully superseded by a newer intent, **when** the user runs `/abcd:intent reclassify --kind superseded --by --reason "fully covered by itd-M"`, **then** the file moves from `shipped/` to `superseded/`; frontmatter gains both `superseded_by: itd-M` AND `kind_at_supersession: standalone` (preserving what the intent was when retired); a `reclassification_history` entry records the date + reason. -- **Given** an active discipline that is being replaced by a stricter successor, **when** the user runs `/abcd:intent reclassify --kind superseded --by `, **then** the file moves from `disciplines/` to `superseded/`; frontmatter gains both `superseded_by: itd-M` AND `kind_at_supersession: discipline` so future readers can tell the retired intent was a rule (not a capability). -- **Given** an intent in `superseded/` lacks either `superseded_by` or `kind_at_supersession`, **when** `internal/core/lint` runs, **then** the lint hard-blocks with a clear error — both fields are required for every superseded intent regardless of original kind. -- **Given** the corpus contains two intents whose press releases reference each other and target the same release, **when** `intent-fidelity-reviewer`'s shape-classification role runs (pre-commit or on `/abcd:intent` invocation), **then** a "bundle candidate" suggestion appears in `/abcd:intent` status output AND in `.abcd/logbook/audit/shape-/report.{json,md}`. -- **Given** the user declines a shape-classification suggestion, **when** the reviewer runs again on the same corpus state, **then** the declined suggestion is NOT re-surfaced (logged in the reviewer's "declined-suggestions" cache) — the reviewer doesn't nag. -- **Given** a project shipping a non-abcd application (e.g., idelphiDev) under the abcd intent framework, **when** the user captures a project-specific discipline (e.g., "every recording feature includes a privacy-impact review"), **then** the same lifecycle applies — the discipline lands in `disciplines/`, never gets its own spec, and is enforced against every other spec via plan-review. The framework treats it identically to abcd's own disciplines. -- **Given** the brief and intent corpus are updated, **when** `internal/core/lint` runs, **then** it verifies (a) every intent in `planned/` has `kind` set and matching the directory; (b) `kind: bundle-member` intents have `bundle:` set bidirectionally with their bundle-mates; (c) `kind: discipline` intents live only in `disciplines/` or `superseded/`; (d) `kind: discipline` intents have `epic_id: null`; (e) discipline-kind intents have NO `status` field (the directory is the state); (f) every intent in `superseded/` has both `superseded_by` and `kind_at_supersession`. Violations are hard-block lint failures. +- **Given** two draft intents, **when** `abcd intent plan itd-A itd-B` runs, **then** it asks for a bundle name, mints one shared spec naming both, writes `kind: bundle-member` and `bundle: ` on each, and moves both to `planned/` together; two drafts scoped to different phases are refused naming both phases, and nothing moves. +- **Given** a bundle's shared spec is closed, **when** the close-hook runs, **then** every member with that `bundle:` ships together. +- **Given** `abcd intent reclassify --kind ` or `--kind superseded --by --reason "…"`, **when** it runs, **then** the record's kind, shelf and links change in one write, with `superseded_by` on the record and `supersedes` on the successor written together; a shipped intent asked to become a discipline is refused and told to file a discipline that supersedes it. +- **Given** one member of a bundle is superseded, **when** the reclassify completes, **then** the surviving member stays `bundle-member` and its record states that the bundle now has one member. +- **Given** the record lint runs, **when** a planned intent's `kind` does not match its shelf, or a `bundle-member` names a bundle whose other members do not exist, **then** the lint refuses naming the record. ## Dependencies @@ -158,13 +155,18 @@ None stated. - **Coordinated with:** [itd-48](itd-48-intent-fidelity-reviewer-roles-2-3.md) — itd-48 owns the cross-document role (Role 2) and the shape-classification role (Role 3) on `intent-fidelity-reviewer`, superseding [itd-31](../superseded/itd-31-cross-document-fidelity-reviewer.md) which originally introduced the cross-document concept. The agent's three-role architecture is documented uniformly across the brief and intent surfaces. - **Coordinated with:** [itd-48](itd-48-intent-fidelity-reviewer-roles-2-3.md) (cross-document role) and [itd-34's own shape-classification role] — these are the second and third roles on `intent-fidelity-reviewer` (the first being single-document fidelity per itd-1). Each role has its own user-facing verb under `/abcd:intent` (consistency, shape, review respectively). The earlier `tier-0-audit-substrate` bundle ([itd-31](../superseded/itd-31-cross-document-fidelity-reviewer.md) + itd-32) was dissolved on 2026-05-07 when the unified-`/abcd:audit`-surface premise no longer held; itd-31 promoted to standalone (later superseded by itd-48 on 2026-05-27), itd-32 superseded. An even earlier attempted bundle (`intent-capture-discipline`, itd-27 + itd-30) was retired on the same day because itd-27 and itd-30 are scoped to different phases — bundles cannot span phases (one shared spec shipped together is the invariant). Both intent-capture intents reclassified to standalone. +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: + +1. **Both remainders are built**: the bundle command and the `reclassify` verb. +2. **You name the bundle**; the command asks for a short name. +3. **A survivor stays.** When one member is superseded the other stays a bundle-member of a bundle of one, and says so. +4. **A shipped intent never changes kind.** A rule discovered after the fact is filed as a discipline that supersedes it. + ## Open Questions -- **Bundle ID assignment.** Multi-arg `/abcd:intent plan itd-A itd-B` needs to generate or accept a bundle ID. Options: (a) auto-generate a slug from the joint title (terse), (b) prompt for an explicit ID at plan time (deliberate), (c) require pre-declaration in one of the member intents' frontmatter before plan runs (strict). Recommend (b) with default suggestion derived from joint title. -- ~~**What happens if a bundle's intents are scoped to different phases?**~~ **Resolved 2026-05-07** — bundle invariant codified: all members belong to the same phase; multi-arg plan hard-blocks via `IL011` if they disagree. See [`brief/04-surfaces/05-intent.md § 1`](../../brief/04-surfaces/05-intent.md#1-intent-ids-kinds-and-lifecycle) "Bundle invariant" and AC bullet above. Worked example: `intent-capture-discipline` retirement (itd-27 + itd-30 scoped to different phases). -- **Discipline reclassification of a *shipped* intent.** Hypothetical: a standalone intent ships, then later we realise it was actually a discipline all along (rule applied to every subsequent spec anyway). Can `/abcd:intent reclassify --kind discipline` work post-ship? Probably yes, with the historical fidelity audit preserved as a `pre-reclassification-audit` field. Edge case; defer until first occurrence. -- **Bundle dissolution.** If two bundle-members ship as a bundle, then later one of them is superseded, what happens to the bundle? Recommend: the surviving member stays at `kind: bundle-member` with `bundle: ` but a new `bundle_dissolved: true` field; the superseded member moves to `superseded/` as usual. -- **Multi-bundle membership.** Can an intent belong to two bundles? Recommend no — bundle membership is exclusive. If two bundles overlap on an intent, that's a sign one of the bundles is mis-scoped. +_None open; decisions 2 to 4 settle the three this record carried._ ## Audit Notes @@ -178,3 +180,7 @@ _Empty. Populated by intent-fidelity-reviewer's single-document role when this i - [`README.md`](../../../../README.md) — top-level "What is intent-driven development?" section gains a paragraph on the three kinds; `/abcd:intent` planned-commands table gains `reclassify` and multi-arg `plan` rows. - 2026-05-07 audit conversation that surfaced the structural finding: ~40% of intents have non-1:1 relationships with at least one other intent. - The discipline-kind framing is consistent across project types (framework projects produce more disciplines than application projects, but both produce some). Cross-project taxonomy revisit: see "Discipline subtypes are deferred" above. + +## Grounds + +- pursued: the autonomous run supersedes and re-plans records unattended and has no verb to do it with; we expect the reclassify verb and the bundle command to be what it calls, and the hand moves with their one-way links to stop; shown wrong if a record superseded after this ships still carries a one-way link, or if the run never calls either diff --git a/.abcd/development/intents/planned/itd-42-coherence-aware-grill.md b/.abcd/development/intents/planned/itd-42-coherence-aware-grill.md index 94bd6d154..0c58b94bb 100644 --- a/.abcd/development/intents/planned/itd-42-coherence-aware-grill.md +++ b/.abcd/development/intents/planned/itd-42-coherence-aware-grill.md @@ -1,7 +1,7 @@ --- id: itd-42 slug: coherence-aware-grill -spec_id: null +spec_id: spc-2609211918551301 kind: standalone suggested_kind: standalone reclassification_history: [] @@ -21,10 +21,14 @@ warrants_assumed: blocked_by: [itd-27] builds_on: [itd-41] severity: major +impact: additive --- # Grill Reads an Intent Against the Brief and Its Siblings, Not Just the Glossary +> **Re-scoped on 2026-09-21** by the product thinker: this record is the automated pre-pass the decomposition discipline (itd-84) names as its next rung. Before the planning interview, abcd reads the brief's invariants, the principles and a one-line index of every intent, and writes the coherence questions into the planning brief the interview starts from. The grill it first named is superseded (itd-27); the press release and scope below are read through this paragraph and the Decisions section. + + ## Press Release > **abcd's grill stops checking an intent in isolation: a full grill now reads it against the brief's invariants, its scope boundary, and every other intent — and asks the coherence questions a solo interrogation cannot.** Capturing an idea stays one line and zero friction. But when a product thinker promotes a draft, the grill loads more than the terminology glossary: it loads the brief's invariants and scope sections, and a one-line index of every other intent — drafted, planned, and shipped. Now it can ask the question that actually catches wrong code: "Invariant 3 says config is never written outside `~/.abcd/` — your intent implies it is; which gives?" and "itd-19 already covers stage-aware behaviour — how is this different?" It still grills vague terms and hidden assumptions; it now also grills *incoherence* — the intent that is locally clear and globally wrong. Grilling against the corpus is grounded the way the phase negotiator is grounded: a coherence concern that cannot be tied to a named invariant, scope clause, or sibling intent is asked as a Socratic question, never asserted as a conflict. @@ -33,7 +37,7 @@ severity: major ## Why This Matters -The intent layer is abcd's highest-leverage moment for product clarity ([itd-27](itd-27-grill-skill-and-glossary.md) built `/abcd:intent grill` on exactly that premise). But the grill itd-27 shipped has a blind spot its `--with-docs` flag name actively hides: "glossary-aware" mode loads **only** the terminology database. It checks that an intent *speaks* the brief's words. It never reads the brief, never reads another intent, never reads a shipped spec. It enforces vocabulary; it does not enforce coherence. This intent also corrects the misnomer: `--with-docs` becomes `--glossary` (the terminology tier it always was), the new coherence tier is `--coherence`, and `--full` runs both. +The intent layer is abcd's highest-leverage moment for product clarity ([itd-27](../superseded/itd-27-grill-skill-and-glossary.md) built `/abcd:intent grill` on exactly that premise). But the grill itd-27 shipped has a blind spot its `--with-docs` flag name actively hides: "glossary-aware" mode loads **only** the terminology database. It checks that an intent *speaks* the brief's words. It never reads the brief, never reads another intent, never reads a shipped spec. It enforces vocabulary; it does not enforce coherence. This intent also corrects the misnomer: `--with-docs` becomes `--glossary` (the terminology tier it always was), the new coherence tier is `--coherence`, and `--full` runs both. So an intent can pass a full grill — crisp terms, EARS-clean acceptance, warrants surfaced — and still: @@ -66,30 +70,33 @@ The brief is already structured for selective loading — numbered sections, inv - **Grilling against shipped *code*** — Tier 3 reads shipped *intents*, not the implementation. Delivered-reality comparison remains `intent-fidelity-reviewer`'s job at the shipped transition. - **A new sub-verb or command** — this is a capability of the existing `/abcd:intent grill`, not a sibling verb. +## Mechanism + +We expect a pre-pass that reads the invariants and the sibling index to catch the contradiction or the duplicate before a spec exists, because both are visible from the record alone and the interview today finds them only when the person happens to remember; shown wrong if planned intents still turn out to duplicate or contradict one another after it ships. + ## Scope Conditions None stated. ## Acceptance Criteria -> _BDD format, per the itd-1 discipline._ +- **Given** a draft intent, **when** the pre-pass runs, **then** the planning brief it writes names every brief invariant the draft's text implies a conflict with, quoting the invariant's line and the draft's. +- **Given** a draft that overlaps an existing intent on any shelf, **when** the pre-pass runs, **then** the brief names the sibling from a one-line index of every intent, writes the question with the four answers (keep both, bundle, supersede, refine), and, where it has one, its recommendation with the reason in the prose beside the question and never as a marked option. +- **Given** the pre-pass has run, **when** the tree is inspected, **then** the draft, the brief and the sibling intents are unchanged; the pre-pass read the invariants, the principles, the index and the draft, and wrote only the planning brief under the local tier. +- **Given** a planning brief with questions, **when** the interview runs, **then** each question is asked, and the answer lands on the record as a decision or a typed link. +- **Given** a concern the pre-pass cannot anchor to a named invariant or record, **when** it writes the brief, **then** the concern is a question marked unanchored, not a finding. -- **Given** an intent in `drafts/`, **when** the product thinker runs a light grill on it, **then** the grill loads at most the glossary tier and does not load brief or sibling context — capture-stage grilling stays cheap. -- **Given** an intent being promoted out of `drafts/`, **when** the full grill runs, **then** it loads the glossary tier, the named brief invariant/scope/surface slices, the `principles/` set, and the one-line sibling-intent index. -- **Given** an intent whose body implies behaviour that a `02-constraints/03-invariants.md` invariant forbids, **when** the full grill runs, **then** it surfaces the conflict and cites the specific invariant; a conflict surfaced without such a citation is a defect. -- **Given** an intent that overlaps a sibling intent, **when** the full grill runs, **then** it names the sibling intent ID and asks the product thinker to state the difference as a Socratic question — sibling overlap is never asserted as fact, because a draft sibling is a mutable anchor. -- **Given** the product thinker answers a surfaced overlap question, **when** the session ends, **then** the answer — the stated difference, or a kill/merge decision — is captured in the `grill-report` against that question, so the claimed distinction is on record. -- **Given** a coherence concern the grill cannot anchor to a named invariant or scope clause, **when** it surfaces that concern, **then** it is phrased as a Socratic question tagged with a named move — never as an asserted conflict. -- **Given** a full grill has run, **when** the `grill-report.json` is written, **then** coherence questions and any grounded conflicts appear in it alongside the existing question stream, each grounded conflict carrying its invariant or scope-clause anchor reference. -- **Given** the grill has surfaced a coherence conflict, **when** the session ends, **then** the intent, the brief, and the sibling intents are unchanged on disk — the grill's coherence output is advisory. +## Decisions -## Open Questions +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: -> _Tier-selection surface and the Tier 3 index mechanism were resolved during the grill (2026-05-16): tier set is lifecycle-derived with `--light`/`--full` override; the sibling index is built fresh each grill, not a maintained file. Brief-slice degradation is resolved in scope (skip-and-warn). The questions below remain genuine plan-time decisions._ +1. **The record is the pre-pass**, run before the interview and writing into the planning brief; the interview stays the human's. +2. **Overlaps are asked with the four standard answers**, and the pre-pass may recommend one with its reason in the prose beside the question, never as a marked option (the GRILL rule). +3. **The loader is the interview's own** (the planning-brief writer the intent page describes), not a module shared with the phase negotiator or the fidelity reviewer; the context is the invariants, the principles, the index and the draft, and nothing else, which is the budget. -- Does coherence grilling share a context-loading module with itd-41's phase negotiator and itd-31's cross-document fidelity reviewer, or keep its own loader until the three demonstrably converge? -- Should a surfaced sibling-overlap question, once the thinker answers it, be allowed to *recommend* reclassification (bundle-member) or supersession, or strictly surface-and-record? itd-27's grill already touches reclassification-adjacent territory. -- Token budget — at 40+ intents the one-line index is small, but the brief slices plus glossary plus intent body must still fit. Is there a point where Tier 2/3 must itself become selective (the itd-39 boundary)? +## Open Questions + +_None open; decisions 2 and 3 settle the three this record carried._ ## Audit Notes @@ -97,7 +104,11 @@ _Empty. Populated by intent-fidelity-reviewer when intent moves to shipped/._ ## References -- Extends: [itd-27](itd-27-grill-skill-and-glossary.md) (grill skill & glossary) — adds a coherence tier to the grill itd-27 built; the glossary tier's behaviour is unchanged. **Also renames itd-27's `--with-docs` flag to `--glossary` and adds `--coherence` / `--full`** — itd-27's surface table and the grill `SKILL.md` flag list must be updated when this intent is planned. +- Extends: [itd-27](../superseded/itd-27-grill-skill-and-glossary.md) (grill skill & glossary) — adds a coherence tier to the grill itd-27 built; the glossary tier's behaviour is unchanged. **Also renames itd-27's `--with-docs` flag to `--glossary` and adds `--coherence` / `--full`** — itd-27's surface table and the grill `SKILL.md` flag list must be updated when this intent is planned. - Shares the grounded-adversary pattern with: [itd-41](../drafts/itd-41-phase-negotiator.md) (phase negotiator) — Socratic where it questions, grounded where it asserts. - Defers to: [itd-39](../drafts/itd-39-scope-aware-memory-retrieval.md) (scope-aware memory retrieval) — full-body cross-intent comparison at scale is itd-39's problem, not this intent's. - Coordinates with: [itd-48](itd-48-intent-fidelity-reviewer-roles-2-3.md) (cross-document fidelity reviewer — supersedes [itd-31](../superseded/itd-31-cross-document-fidelity-reviewer.md)) — different register: itd-48's Role 2 reviews delivered documents for drift; this grills an intent for coherence before it is planned. + +## Grounds + +- pursued: the autonomous run may prepare interviews but not perform them, and the planning brief it writes is only worth reading if this pass fed it; we expect the first briefs to carry a conflict or an overlap the person had not seen; shown wrong if the briefs raise nothing the person did not already know diff --git a/.abcd/development/intents/planned/itd-48-intent-fidelity-reviewer-roles-2-3.md b/.abcd/development/intents/planned/itd-48-intent-fidelity-reviewer-roles-2-3.md index 5660b7630..ee0f38b18 100644 --- a/.abcd/development/intents/planned/itd-48-intent-fidelity-reviewer-roles-2-3.md +++ b/.abcd/development/intents/planned/itd-48-intent-fidelity-reviewer-roles-2-3.md @@ -1,7 +1,7 @@ --- id: itd-48 slug: intent-fidelity-reviewer-roles-2-3 -spec_id: null +spec_id: spc-2609211921272106 kind: standalone suggested_kind: bundle-member reclassification_history: @@ -11,10 +11,14 @@ related_adrs: [] routed_from: ["spc-33:A1", "spc-33:A2", "spc-33:A3", "spc-33:A4", "spc-33:G1"] builds_on: [itd-34, itd-5] severity: major +impact: additive --- # `intent-fidelity-reviewer` Gains Its Cross-Doc And Kind-Classification Roles +> **Re-scoped on 2026-09-21** by the product thinker: the consistency pass only. The shape role goes to the kinds lint (itd-34) and the overlap question to the pre-pass (itd-42); the headless oracle leg this record named no longer exists. The press release and scope below are read through this paragraph and the Decisions section. + + ## Press Release > **abcd's `intent-fidelity-reviewer` agent grows from one `role` to three.** `Role 1` (per-intent fidelity) grades a shipped intent's acceptance criteria against delivered reality and writes `MET`/`MET_WITH_CONCERNS`/`NOT_MET`/`INCONCLUSIVE` verdicts back into the intent's `## Audit Notes`. This intent ships `Role 2` (cross-document consistency — surfaces terminology drift, premise contradictions, scope leakage, sequencing impossibilities, naming conflicts across the brief + intents corpus) and `Role 3` (kind classification — examines whether intents' declared `kind` still fits the corpus and surfaces suggested reclassifications). With all three roles live, `/abcd:intent consistency` and `/abcd:intent shape` move from documented command surface to working command surface. The reviewer becomes the corpus's continuous fidelity auditor. @@ -125,32 +129,27 @@ The intent is project-agnostic: every abcd project that uses the intent corpus b None stated. +## Mechanism + +We expect a whole-corpus read to find contradictions no single-record review can, because terminology drift, a premise that two records state differently and a scope that leaked are visible only with both ends in view; shown wrong if a pass over the current corpus finds nothing the ledger did not already hold. + ## Acceptance Criteria -- *Given* the agent file with Roles 2 and 3 sections added, *when* `lint_prompts.py` runs, *then* `prompt_version` is bumped, the CHANGELOG carries the bump entry, and at least one injection canary per role exists in `agents/intent-fidelity-reviewer/fixtures/`. -- *Given* a facilitator runs `/abcd:intent consistency` (bare), *when* the corpus contains a known seeded drift (terminology, premise, scope, sequencing, or naming), *then* the command writes a structured report to `.abcd/logbook/audit/consistency-/report.{json,md}` whose findings name the judgement category, the conflicting documents, and the drift kind. -- *Given* a facilitator runs `/abcd:intent consistency itd-N`, *when* the named intent contradicts another corpus document, *then* the persisted report identifies both ends of the contradiction. -- *Given* a facilitator runs `/abcd:intent shape` (bare), *when* an intent in the corpus has drifted from its declared `kind` (e.g., a `standalone` that has become bundle-shaped), *then* the command writes `.abcd/logbook/audit/shape-/report.{json,md}` with a suggestion naming the reclassification target. -- *Given* a facilitator runs `/abcd:intent shape itd-N`, *when* the named intent's kind still fits, *then* a `KIND_OK` scoped verdict is emitted in the persisted report. -- *Given* a Ralph session runs the consistency or shape verb in headless mode, *when* the call reaches `_build_cli_oracle()`, *then* the Codex leg is used (per itd-47) and the command completes with a real verdict. -- *Given* itd-48 is the standalone intent owning Roles 2 and 3, *when* it is planned, *then* the planned spec ships only the on-demand verbs (no pre-commit hook installation) and records pre-commit scheduling for both roles as deferred follow-ups. +- **Given** `abcd intent consistency` runs bare, **when** it completes, **then** a dated report exists under the reviews shelf listing each contradiction found across the brief and every intent (terminology, premise, scope, sequencing, naming), with both ends quoted and located, and the commit it read named. +- **Given** the report's findings, **when** the pass files them, **then** each is an issue naming both ends with the report as its evidence, and a finding the ledger already holds is linked to the existing record rather than filed twice. +- **Given** `abcd intent consistency `, **when** it runs, **then** the pass is scoped to that intent against the corpus and the report says so. +- **Given** the pass has run, **when** the tree is inspected, **then** the brief and the intents are unchanged; the binary assembled the input and validated the return, and the judgement rode the host. + +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: + +1. **The consistency pass only.** The shape role is the kinds lint in itd-34; the overlap question is the pre-pass in itd-42. +2. **A report, and a capture per finding**, deduplicated against the ledger. ## Open Questions -- **Standalone vs. two standalones.** Plan time resolved this in favour of - one standalone intent whose scope covers both roles: the shared agent - file, shared `prompt_version` family, shared injection-canary discipline, - and shared oracle infrastructure all argue against splitting. A future - spec could split the prompt sections without splitting the intent if the - roles diverge. -- **Pre-commit follow-up shape.** A follow-up intent will define when and - how `/abcd:intent consistency` and `/abcd:intent shape` run in - pre-commit (every intent-touching commit, every kind-frontmatter-touching - commit, or only at state transitions). Out of scope for this intent. -- **Mechanical Role 2 categories.** Schema/state contradictions, reference - rot, and acknowledgement gaps are deferred to a separate intent that - owns the mechanical (lint-driven) half of cross-doc fidelity. spc-29 - ships only the judgement half. +_None open; the standalone-versus-two question this record carried is moot with one role left._ ## Routed Deferrals (spc-33) @@ -199,3 +198,7 @@ scope captured here — NOT active spc-33 work: rows this intent makes real. - **a dated working-log entry (2026-05-16)** — the gap entry that motivated this intent. + +## Grounds + +- pursued: the autonomous run builds forty-eight intents against this corpus, and a contradiction between two of them is a stop condition it cannot resolve; we expect the first whole-corpus pass to find contradictions the ledger does not hold; shown wrong if it finds none diff --git a/.abcd/development/intents/planned/itd-50-loop-toward-acceptance.md b/.abcd/development/intents/planned/itd-50-loop-toward-acceptance.md index 2eb69732a..485bc2476 100644 --- a/.abcd/development/intents/planned/itd-50-loop-toward-acceptance.md +++ b/.abcd/development/intents/planned/itd-50-loop-toward-acceptance.md @@ -1,7 +1,7 @@ --- id: itd-50 slug: loop-toward-acceptance -spec_id: null +spec_id: spc-2609211924346308 kind: standalone suggested_kind: standalone reclassification_history: [] @@ -12,10 +12,14 @@ grilled_at: 2026-06-02 blocked_by: [itd-53] builds_on: [itd-44, itd-43] severity: major +impact: additive --- # The Audit Loop Drives An Intent To Acceptance — Or Calls For A Replan +> **Re-scoped on 2026-09-21** by the product thinker: the loop is the last stage of `abcd build` (itd-2609201916151817), not a per-intent mode. After the fidelity audit, a not-met verdict starts a fix round with a fresh implementer, bounded by the pace rule's fix-round count; exhaustion hands the intent back as unachievable, reopened to drafts with the reason. The press release and scope below are read through this paragraph and the Decisions section. + + ## Press Release > **abcd's fidelity audit stops being a report card and becomes a loop the facilitator can drive toward acceptance.** Today, when a shipped intent's delivered reality fails a criterion, the `intent-fidelity-reviewer` records a `NOT_MET` verdict and the voyage moves on — the divergence is logged, but nothing closes it. With this change, the facilitator elects, per intent, whether that intent should *loop toward acceptance*: a `NOT_MET` verdict re-opens the work and iterates against the same acceptance criteria until they read `MET`, bounded by a budget so the loop can never grind forever. When an intent genuinely *cannot* be met as written, the loop doesn't thrash — it terminates with an explicit `UNACHIEVABLE` verdict that summons the product thinker and facilitator to sit together and replan the intent. And only once the machine-checkable criteria all read `MET` is the product thinker invited to manually verify the intention — so a human is never asked to hand-test something the audit already knows is broken. @@ -65,17 +69,17 @@ This intent is **project-agnostic**: every abcd project ships intents whose deli None stated. +## Mechanism + +We expect a bounded fix round after the audit to turn most not-met verdicts into met without a person, because a not-met criterion names exactly what to change and a fresh implementer with that criterion is the shape the pilot's fix rounds already proved; shown wrong if the loop's fix rounds mostly end unachievable. + ## Acceptance Criteria -- *Given* an intent carries `audit_mode: loop-to-acceptance` in its frontmatter (set at plan time; portable with the intent), *when* a fidelity review returns `NOT_MET` for a criterion, *then* the linked work is re-opened and re-reviewed against the same criteria, and the cycle repeats until all criteria read `MET` or the iteration budget is exhausted. -- *Given* an intent in `loop-to-acceptance` whose iteration budget is exhausted (or whose criteria are judged unmeetable as written), *when* the loop terminates, *then* the intent-level Family-2 rollup becomes `UNACHIEVABLE`, the loop **stops** (never auto-continues), the intent is **marked with a written explanation of why it is unachievable**, and a replan invitation is recorded naming both the product thinker and facilitator — with no automatic rollback of delivered reality and no machine-authored replan. -- *Given* an intent whose machine-checkable criteria all read `MET`, *when* the product thinker is invited to verify, *then* a manual-verification step is offered and its sign-off is recorded as a verification receipt distinct from the machine verdict of record. -- *Given* an intent whose machine-checkable criteria all read `MET` **but** the product thinker judges the criteria themselves were wrong (the why is not delivered despite every criterion passing), *when* manual verification is rejected, *then* the intent routes to the **replan** path (revise the intent's criteria), **not** a synthetic `NOT_MET` that would re-loop the implementation against criteria that already pass. -- *Given* an intent whose machine-checkable criteria do **not** all read `MET`, *when* the workflow reaches the manual-verification point, *then* the product thinker is **not** asked to verify — the loop (or the replan invitation) runs first. -- *Given* a fidelity review returns `INCONCLUSIVE` (the fail-closed result of a malformed or unreachable reviewer), *when* the loop processes it, *then* it is recorded as today — `INCONCLUSIVE` does **not** summon the product thinker and does **not** itself trigger replan (it is a "could not run the audit" signal, not a "the intent is impossible" signal). -- *Given* an intent left at the default `audit_mode: record-only`, *when* a fidelity review returns `NOT_MET`, *then* behaviour is unchanged from today — the verdict is recorded to `## Audit Notes` and no re-work is triggered. -- *Given* an intent reaches `UNACHIEVABLE` (loop exit) or its manual verification is rejected (wrong-criteria replan), *when* the product thinker takes it up, *then* they use the `/abcd:intent grill` skill to think the replan through, and the recorded `why-unachievable` / rejection justification seeds that grill session. -- *Given* the on-close lifecycle hook (`intent_lifecycle`), *when* any of these modes is active, *then* the hook remains a pure data function (no subprocess, no oracle dispatch) — the mode logic lives in the drainer/policy layer. +- **Given** a `build` lane whose fidelity audit returns not-met on any criterion, **when** the verdict is ingested, **then** a fix round starts with a fresh implementer briefed on those criteria, and the audit re-runs after it. +- **Given** the pace rule's fix-round count is exhausted with a criterion still not met, or the auditor judges a criterion unmeetable as written, **when** the loop reaches that point, **then** the lane stops with the verdict unachievable and starts nothing further. +- **Given** an unachievable verdict, **when** the loop hands back, **then** the intent is moved to `drafts/` carrying `replan_reason` and its audit notes, the spec stays open, and the run's summary lists it for a replan. +- **Given** every machine-checkable criterion reads met, **when** the lane reaches its landing, **then** the product thinker is offered a hand verification and the answer is recorded as a grounds entry on the intent in their words; a rejection of the criteria themselves reopens the intent as above. +- **Given** an inconclusive audit (a malformed or unreachable reviewer), **when** the loop processes it, **then** no fix round starts, nothing counts against the budget, and the run names the inconclusive audit in its summary. ## Resolved (grill 2026-06-02) @@ -87,12 +91,18 @@ The 2026-06-02 grill (5 questions across Dialectic / Definition / Counterfactual - **`UNACHIEVABLE` always stops and summons the product thinker**, with a written `why-unachievable` explanation; no machine auto-replan (the product thinker owns the why). The product thinker uses `/abcd:intent grill` to think the replan through, seeded by that explanation. - **`INCONCLUSIVE` stays fail-closed only — no summons, no replan.** It is the result of a malformed/unreachable reviewer (a backend signal), not a "the evidence is contradictory" signal, so it cannot be trusted as a human-summons trigger. +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: + +1. **The loop is `build`'s last stage**, not a per-intent mode; every intent the run ships goes through it. +2. **An iteration is a fix round** (fresh implementer plus re-audit), bounded by the pace rule's count; exhaustion, or a criterion judged unmeetable, is unachievable. +3. **Unachievable reopens the intent to drafts** with the reason and its audit notes. +4. **Hand verification is a grounds entry** in the product thinker's words, not a separate receipt. + ## Open Questions -- **The loop budget for `loop-to-acceptance`.** What consumes an iteration (a full re-review? a re-open + re-implement + re-review cycle?), and what the default budget is. Mirror `MAX_REVIEW_ITERATIONS` or set an intent-grain equivalent. -- **How the replan invitation surfaces.** Re-open the intent to `drafts/` with a `replan_reason`? A dedicated replan queue/surface the facilitator drains? What state the original delivered reality is left in. (Both replan entry points — `UNACHIEVABLE` and wrong-criteria rejection — share this surface.) -- **How manual-verification sign-off is recorded.** A verification receipt schema, distinct from the machine verdict of record, with an explicit `rejected` state that carries the wrong-criteria justification into the seeded grill. -- **Iteration autonomy bound.** `loop-to-acceptance` iterates unattended up to budget; the precise rule for when a budget-exhausted loop flips to `UNACHIEVABLE` vs. is judged unmeetable earlier. +_None open; decisions 2 to 4 settle the four this record carried._ ## Related @@ -124,3 +134,7 @@ In the predecessor implementation each acceptance criterion above is satisfied a | on-close hook stays a pure data function (no subprocess / oracle) | Satisfied | The mode logic rides the spc-43 drainer / policy layer; `intent_lifecycle` is untouched by the loop | **Open questions (predecessor answers, to re-adjudicate at spec time):** loop budget = one re-open+re-review cycle per iteration, default `3` (spc-52.1 § Decision context); replan surface = no `drafts/` move, a `why-unachievable` + replan block in `## Audit Notes` with the intent kept in `shipped/` (spc-52.2 R4); manual-verification sign-off = the receipt schema `{intent_id, machine_rollup, state, justification?, recorded_by_role, ts}` with the `rejected_wrong_criteria` state carrying the justification to the shared replan surface (spc-52.3 R5). + +## Grounds + +- pursued: the run audits every intent it ships and today a not-met verdict is a note nobody acts on; we expect the bounded fix round to close most of them without a person; shown wrong if most fix rounds end unachievable diff --git a/.abcd/development/intents/planned/itd-53-review-queue-auto-drain-fidelity-gate.md b/.abcd/development/intents/planned/itd-53-review-queue-auto-drain-fidelity-gate.md index b43133355..97307b1dd 100644 --- a/.abcd/development/intents/planned/itd-53-review-queue-auto-drain-fidelity-gate.md +++ b/.abcd/development/intents/planned/itd-53-review-queue-auto-drain-fidelity-gate.md @@ -1,7 +1,7 @@ --- id: itd-53 slug: review-queue-auto-drain-fidelity-gate -spec_id: null +spec_id: spc-2609211930059886 kind: standalone suggested_kind: standalone reclassification_history: [] @@ -9,10 +9,14 @@ related_adrs: [adr-16] routed_from: ["spc-33:I-D2"] prd_path: null severity: major +impact: additive --- # A Shipped Intent No Longer Drifts Out Of Audit Just Because Nobody Ran The Review +> **Re-scoped on 2026-09-21** by the product thinker: one bounded command that pays the backlog, `abcd intent audit --owed`, run once by the autonomous run's first batch. The standing list is itd-2609150819445595 and the inline audit of every new lane is `abcd build`'s (itd-2609201916151817, itd-50); the boundary-hooked autodrain and the blocking gate this record first described are dropped. The press release and scope below are read through this paragraph and the Decisions section. + + ## Press Release > **abcd closes the audit back-edge: when a specced block of work ships, its owed fidelity review actually gets run at a safe moment, and a standing gate surfaces any shipped intent whose review is missing or unmet — without ever blocking the autonomous loop.** Today abcd does the honest half: closing a spec moves its intent to shipped and enqueues a fidelity-review entry. But the review itself only runs when someone manually invokes it, so the queue can quietly accumulate owed reviews that nobody drains, and a shipped intent can sit with its acceptance never machine-checked. This intent adds an opt-in drainer that runs queued reviews at a safe boundary (after a loop, at a session edge, in a pre-commit or CI step — never inside the pure close hook), leaving entries deferred rather than failed when no review backend is reachable, plus a consistency gate that lists shipped intents whose latest review is absent or not-met. Enforcement, not just bookkeeping — and loop purity preserved. @@ -43,22 +47,28 @@ The fix is not to make the close hook run the review — that would break loop p None stated. +## Mechanism + +We expect the backlog of owed audits to clear once paying it is one bounded command, because the debt accrued only while each audit was a hand-run request and ingest; shown wrong if the owed count is unchanged a month after it ships. + ## Acceptance Criteria -> _Given-When-Then per the itd-1 discipline._ +- **Given** `abcd intent audit --owed [--max ]`, **when** it runs, **then** it lists the owed audits oldest first and runs each as the host pass a single-intent audit uses, ingesting each verdict; the cap stops it and the summary names how many remain. +- **Given** no reviewer is reachable, **when** it runs, **then** every entry is left owed, not failed, and the summary says why nothing ran. +- **Given** a verdict, **when** it is ingested, **then** it lands exactly as a single audit's does (audit notes, receipt, scope-condition dispositions), and a not-met verdict on a long-shipped intent is captured, not fixed. +- **Given** a spec closes, **when** the close hook runs, **then** it still only enqueues; nothing starts a reviewer from it. -- **Given** `review.autodrain` is enabled, **when** a safe boundary is reached and a review backend is reachable, **then** pending fidelity-review queue entries are run and their verdicts recorded. -- **Given** no review backend is reachable, **when** the drainer runs, **then** entries are left `deferred` (not failed) and nothing blocks. -- **Given** the close hook, **when** a spec closes, **then** it still only enqueues — no subprocess or oracle dispatch is added to it (loop purity preserved). -- **Given** the consistency gate, **when** it runs, **then** it lists every shipped intent whose latest fidelity review is absent or not-met. -- **Given** `review.autodrain` is off (default), **when** specs close, **then** behavior is unchanged from today (enqueue-only, manual review). +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: + +1. **One bounded command for the backlog**, run once by the autonomous run's batch 0; no boundary-hooked autodrain. +2. **The gate reports and never blocks**; the report is itd-2609150819445595's listing. +3. **Cost is bounded by `--max`** and by the host pass running one audit at a time. ## Open Questions -- Which safe boundary is the primary drain point — a post-turn hook, a session-edge step, a pre-commit/CI step, or several, configurably? -- Does the gate merely report, or can it be wired to block a commit / a phase transition when a shipped intent is unaudited or not-met? (Report first; blocking is a policy decision.) -- How does the drainer bound its own cost (number of reviews per drain, token budget) so a large backlog does not stall the boundary it runs at? -- Interaction with itd-50 (loop-toward-acceptance): does the drainer just run reviews, with itd-50's policy deciding what a not-met verdict triggers, or does the drainer need hooks for that policy from the outset? +_None open; decisions 1 to 3 settle the three this record carried._ ## Audit Notes @@ -75,3 +85,7 @@ _Empty. Populated by intent-fidelity-reviewer when intent moves to shipped/._ NO; add a drainer instead). - Touches: the pure on-close lifecycle hook (spc-28) and the review-queue drain/claim machinery; the fidelity reviewer (spc-12) is the run target. + +## Grounds + +- pursued: the run's first batch audits every shipped-but-open intent and has no verb to run the audits with; we expect the owed count to fall to zero in that batch and stay near it once build audits inline; shown wrong if the owed count is unchanged a month after it ships diff --git a/.abcd/development/intents/planned/itd-6-rp-mcp-only-integration.md b/.abcd/development/intents/planned/itd-6-rp-mcp-only-integration.md index 186043d92..b4b112a90 100644 --- a/.abcd/development/intents/planned/itd-6-rp-mcp-only-integration.md +++ b/.abcd/development/intents/planned/itd-6-rp-mcp-only-integration.md @@ -1,17 +1,21 @@ --- id: itd-6 slug: rp-mcp-only-integration -spec_id: null +spec_id: spc-2609211950427074 kind: standalone suggested_kind: null reclassification_history: [] builds_on: [itd-2] severity: minor +impact: additive --- > **⚠️ Framing superseded by [ADR-25](../../decisions/adrs/0025-host-delegated-llm-default.md)** (host-delegated LLM is the default; RepoPrompt is one optional oracle adapter among many, not abcd's single integration — see also [ADR-22](../../decisions/adrs/0022-bundled-deps-as-pluggable-adapters.md)). The intent itself stays live and is scheduled in Phase 0's `## Scope` as the oracle adapter seam: read "abcd's only RP integration is MCP" below as the contract of the *RP adapter*, not of abcd — the RP-specific mechanics (MCP bridge, cascade position, `chat_id` semantics) are adapted to the adapter seam at spec time. -# RP-Only Integration: abcd Talks to RepoPrompt via MCP, Period +# RepoPrompt over MCP is one opt-in reviewer route, beside the host's own agents + +> **Re-filed on 2026-09-21** by the product thinker: not abcd's one integration but one opt-in reviewer adapter. abcd is host-delegated by default; with `oracle.review = rp` configured, a lane's reviews go to RepoPrompt over MCP, and when it is unreachable the host's own agent runs the review and the receipt says so. The three-step cascade, the setup discovery and the non-Mac flow this record first described are dropped; the press release and scope below are read through this paragraph and the Decisions section. itd-7 waits on this record. + ## Press Release @@ -54,23 +58,29 @@ This intent re-frames the brief's RP integration: drop "select RP backend with p None stated. +## Mechanism + +We expect a reviewer route a person has already configured and paid for to be used over one abcd would have to own, because the review is the run's scarcest step and the person's own tool is where their model choices already live; shown wrong if nobody opts in within a release of it shipping. + ## Acceptance Criteria -> _BDD format, per `itd-1-acceptance-gates`. These gates are checked by `intent-fidelity-reviewer` when this intent moves to `shipped/`._ +- **Given** `oracle.review = rp` in the repository's or the machine's abcd configuration, **when** a `build` lane reaches its reviews, **then** the ruthless and security review requests are sent to RepoPrompt over MCP and each returned verdict is recorded by the loop as any validator's is. +- **Given** the adapter is configured and RepoPrompt is unreachable, **when** a review is due, **then** the host's own agent runs it and the receipt states that the review fell back and why; no review is silently skipped. +- **Given** the adapter is not configured, **when** a review runs, **then** nothing differs from today, and abcd never spawns, installs or configures RepoPrompt. +- **Given** the adapter ships, **when** the brief's adapters chapter and the command page are read, **then** the adapter is one entry beside the command-line runner, with the opt-in named. -- **Given** a macOS user with RepoPrompt installed and `oracle.backend = "rp"` (or `"auto"` with RP available), **when** any abcd command invokes the oracle (lifeboat-oracle, press-release-composer, intent-fidelity-reviewer, plan-review, impl-review, prompt SOTA audit), **then** the call goes through `mcp__RepoPrompt__*` tools exclusively — no `claude -p` subprocess spawn, no direct OpenAI / Anthropic / Google API call, and no preset-selection prompt to the user. -- **Given** abcd issues an RP MCP audit call, **when** the call returns, **then** `McpResult.chat_id` is populated AND the audit-fix loop in `oracle.py` threads that `chat_id` back as the `chat_id` argument on the next call — never `--new-chat`, never a fresh `rp builder`. This holds within a single `abcd-cli` invocation across plan-review→fix→re-review, impl-review→fix→re-review, and lifeboat-oracle→fix→re-audit cycles. (Per ADR-02 § 3: "same chat" is scoped to within one `abcd-cli` command invocation; cross-invocation chat continuation requires fresh GUI approval and is not supported for autonomous operation.) -- **Given** an audit-fix iteration produces a verdict change in either direction (NEEDS_WORK→SHIP after fixes; SHIP→NEEDS_WORK after a regression), **when** abcd's `re_audit` runs, **then** both directions are accepted as valid signal — no rejection of downgrades, no special-casing of upgrades. -- **Given** RP MCP is unreachable (RP not running, MCP server config missing, network failure), **when** abcd attempts an oracle call with `oracle.backend = "auto"`, **then** the resolution chain falls through to Codex CLI (if `codex` is on PATH) and then to in-session subagent (per itd-2) — three-step cascade, surfaced in the run log so the user can see which backend served the call. -- **Given** the user runs `/abcd:ahoy` for the first time on a macOS machine with RP installed, **when** ahoy's setup discovery runs, **then** the RP MCP config path is detected (in `~/Library/Application Support/RepoPrompt/MCP/` or the project's `.mcp.json`), recorded in `.abcd/config.json` → `oracle.rp.mcp_config_path`, reachability is tested, and `oracle.backend` is locked to `"rp"` on success or to the next available backend on failure (with a one-time hint about how to enable RP later). -- **Given** the user has configured Claude, Codex, and Gemini as separate model presets inside RP, **when** abcd issues different oracle call types (review, audit, question), **then** abcd makes no preset-selection decision — RP routes to whatever the user has configured for that call type. abcd's logs record only the MCP call shape, never the resolved model. -- **Given** a non-Mac user with Codex CLI but no RP, **when** abcd's resolution chain runs, **then** Codex CLI is selected and the user is never prompted about RP setup; RP-related friction is invisible to non-Mac users. +## Decisions + +Ruled by the product thinker on 2026-09-21, in the interview that gave this intent its spec: + +1. **One opt-in reviewer route**, not the integration; the host stays the default. +2. **Reviews only**, not audits. +3. **Unreachable falls back to the host**, with the receipt saying so. +4. **itd-7 waits** on this record and is not in the run. ## Open Questions -- What does the RP MCP API actually expose for "review this prompt" vs "ask this question" vs "audit this content"? Need to verify which `mcp__RepoPrompt__*` tools cover the abcd oracle use cases (lifeboat-oracle, press-release-composer, intent-fidelity-reviewer, prompt SOTA audit, plan-review, impl-review). _(Open: this is an RP-API-shape sub-question; spc-5 delivered the `MCPBridge` transport but not the per-oracle-call tool mapping.)_ -- How does this interact with itd-22 (OpenCode portability)? OpenCode probably has its own equivalent integration pattern (its own MCP setup, or a different surface entirely). The harness's `mcp_call(server, tool, args)` shim should treat "RP" as one server name; OpenCode's equivalent picks up via its own server config. -- **Standardised review-chat naming — naming affordance sub-question (still Open).** When abcd opens an RP chat for an oracle/review call, the chat should carry a deterministic, identifiable label — e.g. `spc-6: Phase 1 reconciliation` (or `: ` for intent-stage work) — not RP's auto-generated `untitled-chat-`. Field evidence (2026-05-16, spc-6 plan-review): the RepoPrompt MCP `oracle_send` tool exposes **no chat-name parameter** — it always auto-names — and the MCP toolset has no rename op. Naming only worked via `flowctl rp chat-send --chat-name`, a path broken against rp-cli 2.x. itd-6 must still settle: does abcd's RP integration name chats at creation (needs an MCP affordance RP may not currently provide — verify against the RP MCP API), name the *tab* instead (the `builder` path does title tabs from its summary), or rename post-creation? A consistent `: