fix: refutation pass should re-derive evidence of any category #122

Merged
jercik merged 3 commits from fix/verify-doc-drift-verification-gate into main 2026-10-08 18:07:55 +00:00
Owner

The refutation pass in verify-doc-drift is told to re-derive the evidence "from the code", but duplicate and obvious findings are cited from documentation. The sentence now says "re-derive that evidence", which defers to the category's citation requirement that main already names.

Follow-up to a review finding on #110, which did not introduce the line. #124 fixed the first half of the sentence; this PR fixes the rest. I did not run an audit to measure how often a model resolves the conflict on its own.

🤖 Generated with Claude Code

The refutation pass in `verify-doc-drift` is told to re-derive the evidence "from the code", but `duplicate` and `obvious` findings are cited from documentation. The sentence now says "re-derive that evidence", which defers to the category's citation requirement that main already names. Follow-up to a review finding on [#110](https://code.j4k.dev/j4k-oss/agent-skills/pulls/110#issuecomment-111947), which did not introduce the line. #124 fixed the first half of the sentence; this PR fixes the rest. I did not run an audit to measure how often a model resolves the conflict on its own. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix: verification gate should accept documentary evidence
All checks were successful
Review / Review (pull_request_target) Successful in 3s
commit-msg / commitlint (pull_request) Successful in 24s
Node tests / node:test (pull_request) Successful in 3m14s
9e60e4c50f
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>

Review 01M4DR47JSGD8XFTJTFAMH7S7J — head 3dd961929db6c66329103c0cf20afafc3e372780

Review — j4k-oss/agent-skills @ 8bece749d7

Scope: diff against base tree 8e004ccf90c2
Status: dispatched — coverage complete (3/3 slots terminal)
Facts: current review-wide projection

Computed under:

{
  "abandonment": "abandonment-v1",
  "anchor_recipe": 1,
  "batch_policy": "batch-v1",
  "coverage": "coverage-v3",
  "dispatch_policy": "dispatch-v2",
  "grounder_version": 1,
  "grounding_read_rule": "grounding-read-v1",
  "promotion_policy": "promotion-v1",
  "report": "report-v4",
  "tally": "tally-v1",
  "triage_settle": "triage-settle-v2"
}

Findings (1)

low — The catalog description repeats the audit intent and includes workflow detail

  • claim: 01M4DR8EQ4172TVEDX6JA2F51H
  • anchor: skills/verify-doc-drift/SKILL.md (snippet)
description: Audit a repository's documentation against its source code and fix the factual drift. Use when the user wants to verify documentation matches the code, audit docs for accuracy, hunt doc-vs-code drift, or check that READMEs / CONTEXT.md / ADRs / standards are still true. Triggers on "verify docs", "audit documentation", "doc drift", "do the docs match the code", "check the docs against the code", "are the docs still accurate".
  • lens: writing-quality · arm: default
  • verdicts: 1 valid / 0 invalid / 0 uncertain
    • pass 01M4DR9H2MR6TCC5S4H4XD4W95 · valid: Grounded exact snippet shows the description opening with a task summary ('Audit a repository's documentation against its source code and fix the factual drift') before 'Use when', then restating the same intent under 'Use when' and again as a 'Triggers on' phrase list. The installed skill-packaging reference states a capability description is a routing rule, begins with 'Use when …', and does not describe the skill's method or contents; it also notes the description consumes catalog context on every load. The reviewer reports no user-only invocation policy, so the capability contract applies. The proposed replacement keeps the routing intent (verify docs match source, audit for factual drift); 'repository documentation' covers READMEs/CONTEXT.md/ADRs/standards, and each dropped literal phrase ('verify docs', 'doc drift', 'do the docs match the code', etc.) is a paraphrase a model would route to the same intent, so no match is lost. The fix-vs-report direction lives in the body, which the reviewer says handles the report-only override. The packaging guidance permits but does not require literal phrases, so dropping them breaks no constraint. Low severity is appropriate.
  • disposition: none

The description spends catalog context on a task summary and repeated paraphrases of the same selection intent. Every catalog reader pays that cost before deciding whether to load the skill. The opening sentence also promises fixing, although the body explicitly supports report-only requests.

The description opens with “Audit a repository's documentation against its source code and fix the factual drift,” then describes the audit intent again under “Use when” and again under “Triggers on.” The installed writing-for-agents packaging guidance says a capability description is a routing rule, begins with “Use when,” and describes distinguishing requests rather than the method or contents. Its One Idea, One Place guidance also calls for removing repeated instructions.

Replace the description with: “Use when the user asks to verify that repository documentation matches its source code or audit documentation for factual drift.” This preserves the selection intent, while the body remains the home for the default correction direction and the user's report-only override.

I read skills/verify-doc-drift/SKILL.md in full, its declared project-docs dependency, and the repository README's skill-packaging context. This is a static writing assessment; no routing experiment was run. The subject declares no user-only invocation policy for this skill. Evidence that a supported client requires these exact literal trigger phrases would justify retaining those phrases; no such requirement appears in the examined material.

Other claims

  • grounding-pending (0)
  • ungrounded (0)
  • rejected (0)
  • duplicate-of (0)
  • unadjudicated (0)

Coverage

Coverage pass: 01M4DR4M5EFWFQ23XN1HSN0N60
Accounting: complete
Slot health: healthy

lens part arm unit status runs loss
general-bug whole default no-claims 1 no
writing-quality whole default claims-emitted 1 no
project-docs whole default no-claims 1 no
  • test-trimming — skipped-by-dispatch: no tests changed
  • restated-sets — skipped-by-dispatch: single-sentence wording edit introduces no enumerations or restated sets
<!-- review:summary --> **Review** `01M4DR47JSGD8XFTJTFAMH7S7J` — head `3dd961929db6c66329103c0cf20afafc3e372780` # Review — j4k-oss/agent-skills @ 8bece749d72a Scope: diff against base tree `8e004ccf90c2` Status: dispatched — coverage complete (3/3 slots terminal) Facts: current review-wide projection Computed under: ```json { "abandonment": "abandonment-v1", "anchor_recipe": 1, "batch_policy": "batch-v1", "coverage": "coverage-v3", "dispatch_policy": "dispatch-v2", "grounder_version": 1, "grounding_read_rule": "grounding-read-v1", "promotion_policy": "promotion-v1", "report": "report-v4", "tally": "tally-v1", "triage_settle": "triage-settle-v2" } ``` ## Findings (1) ### low — The catalog description repeats the audit intent and includes workflow detail - claim: `01M4DR8EQ4172TVEDX6JA2F51H` - anchor: `skills/verify-doc-drift/SKILL.md` (snippet) ``` description: Audit a repository's documentation against its source code and fix the factual drift. Use when the user wants to verify documentation matches the code, audit docs for accuracy, hunt doc-vs-code drift, or check that READMEs / CONTEXT.md / ADRs / standards are still true. Triggers on "verify docs", "audit documentation", "doc drift", "do the docs match the code", "check the docs against the code", "are the docs still accurate". ``` - lens: writing-quality · arm: default - verdicts: 1 valid / 0 invalid / 0 uncertain - pass `01M4DR9H2MR6TCC5S4H4XD4W95` · valid: Grounded exact snippet shows the description opening with a task summary ('Audit a repository's documentation against its source code and fix the factual drift') before 'Use when', then restating the same intent under 'Use when' and again as a 'Triggers on' phrase list. The installed skill-packaging reference states a capability description is a routing rule, begins with 'Use when …', and does not describe the skill's method or contents; it also notes the description consumes catalog context on every load. The reviewer reports no user-only invocation policy, so the capability contract applies. The proposed replacement keeps the routing intent (verify docs match source, audit for factual drift); 'repository documentation' covers READMEs/CONTEXT.md/ADRs/standards, and each dropped literal phrase ('verify docs', 'doc drift', 'do the docs match the code', etc.) is a paraphrase a model would route to the same intent, so no match is lost. The fix-vs-report direction lives in the body, which the reviewer says handles the report-only override. The packaging guidance permits but does not require literal phrases, so dropping them breaks no constraint. Low severity is appropriate. - disposition: none > The description spends catalog context on a task summary and repeated paraphrases of the same selection intent. Every catalog reader pays that cost before deciding whether to load the skill. The opening sentence also promises fixing, although the body explicitly supports report-only requests. > > The description opens with “Audit a repository's documentation against its source code and fix the factual drift,” then describes the audit intent again under “Use when” and again under “Triggers on.” The installed writing-for-agents packaging guidance says a capability description is a routing rule, begins with “Use when,” and describes distinguishing requests rather than the method or contents. Its One Idea, One Place guidance also calls for removing repeated instructions. > > Replace the description with: “Use when the user asks to verify that repository documentation matches its source code or audit documentation for factual drift.” This preserves the selection intent, while the body remains the home for the default correction direction and the user's report-only override. > > I read skills/verify-doc-drift/SKILL.md in full, its declared project-docs dependency, and the repository README's skill-packaging context. This is a static writing assessment; no routing experiment was run. The subject declares no user-only invocation policy for this skill. Evidence that a supported client requires these exact literal trigger phrases would justify retaining those phrases; no such requirement appears in the examined material. ## Other claims - grounding-pending (0) - ungrounded (0) - rejected (0) - duplicate-of (0) - unadjudicated (0) ## Coverage Coverage pass: 01M4DR4M5EFWFQ23XN1HSN0N60 Accounting: complete Slot health: healthy | lens | part | arm | unit status | runs | loss | | --- | --- | --- | --- | --- | --- | | general-bug | whole | default | no-claims | 1 | no | | writing-quality | whole | default | claims-emitted | 1 | no | | project-docs | whole | default | no-claims | 1 | no | - test-trimming — skipped-by-dispatch: no tests changed - restated-sets — skipped-by-dispatch: single-sentence wording edit introduces no enumerations or restated sets
Author
Owner

Closing as obsolete. Main already contains the equivalent change: #124 (7669945, "fix: doc drift verification should accept each category's citation") rewrote the same sentence in skills/verify-doc-drift/SKILL.md's "Adversarial verification" section so the refutation pass defaults to REJECTED unless it reproduces the citation the finding's category requires. This PR's rewording of that line now conflicts with main and would add nothing.

Closing as obsolete. Main already contains the equivalent change: #124 (7669945, "fix: doc drift verification should accept each category's citation") rewrote the same sentence in `skills/verify-doc-drift/SKILL.md`'s "Adversarial verification" section so the refutation pass defaults to REJECTED unless it reproduces the citation the finding's category requires. This PR's rewording of that line now conflicts with main and would add nothing.
jercik closed this pull request 2026-10-07 21:59:46 +00:00
jercik reopened this pull request 2026-10-08 07:17:14 +00:00
Author
Owner

Reopening: the close was a mistake. #124 rewrote the sentence in the "Adversarial verification" section, but main still lacks this PR's other change: skills/verify-doc-drift/SKILL.md line 37 still tells the refutation pass to "re-derive the evidence from the code rather than trusting the first pass", which this PR replaces with fresh evidence that satisfies the citation requirement in "The per-unit audit loop". The next sweep pass will trim this PR to the hunks main still needs.

Reopening: the close was a mistake. #124 rewrote the sentence in the "Adversarial verification" section, but main still lacks this PR's other change: `skills/verify-doc-drift/SKILL.md` line 37 still tells the refutation pass to "re-derive the evidence from the code rather than trusting the first pass", which this PR replaces with fresh evidence that satisfies the citation requirement in "The per-unit audit loop". The next sweep pass will trim this PR to the hunks main still needs.
jercik changed title from fix: verification gate should accept documentary evidence to fix: refutation pass should re-derive evidence of any category 2026-10-08 08:44:24 +00:00
Author
Owner

Pushed d9f1187: merged main and trimmed this PR to the one phrase main still lacks. #124 already made the refutation pass cite the category's required evidence; the remaining "re-derive the evidence from the code" is now "re-derive that evidence", since duplicate and obvious findings are cited from docs. The earlier review of 9e60e4c reported no findings, so there was nothing to settle.

Pushed d9f1187: merged main and trimmed this PR to the one phrase main still lacks. #124 already made the refutation pass cite the category's required evidence; the remaining "re-derive the evidence from the code" is now "re-derive that evidence", since duplicate and obvious findings are cited from docs. The earlier review of 9e60e4c reported no findings, so there was nothing to settle.
fix: merge main and trim the verification gate change
All checks were successful
commit-msg / commitlint (pull_request) Successful in 21s
Node tests / node:test (pull_request) Successful in 2m8s
Review / Review (pull_request_target) Successful in 9m16s
d9f1187ac8
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Author
Owner

Replying to the review summary (review 01M4DAYW84N056TKJ0SG5ZDJS3, head d9f1187ac86b414683f60032b7a7b00532d2109e)

No push for this pass: the change is still needed (main line 37 of skills/verify-doc-drift/SKILL.md still says "re-derive the evidence from the code") and the trial merge with main is clean.

Finding 01M4DBC94GDDGY33EQ7HRGWSKN (description says every finding is "verified against the code"): valid, but not introduced here. #124 already made obvious and duplicate findings cite documentation, so the description was already inconsistent on main. Tracked in #128, which removes the method summary from the description.

Unadjudicated claim 01M4DBCNZT1W72QA03FBWWCB7T (the adversarial pass "describes survivors as real drift"): not a defect. The sentence before it already makes the refutation pass default to REJECTED unless the category's citation reproduces, so the "real drift" phrasing is a gloss and does not change what a pass does. No change.

> Replying to the review summary (review `01M4DAYW84N056TKJ0SG5ZDJS3`, head `d9f1187ac86b414683f60032b7a7b00532d2109e`) No push for this pass: the change is still needed (main line 37 of `skills/verify-doc-drift/SKILL.md` still says "re-derive the evidence from the code") and the trial merge with `main` is clean. Finding `01M4DBC94GDDGY33EQ7HRGWSKN` (description says every finding is "verified against the code"): valid, but not introduced here. #124 already made `obvious` and `duplicate` findings cite documentation, so the description was already inconsistent on main. Tracked in #128, which removes the method summary from the description. Unadjudicated claim `01M4DBCNZT1W72QA03FBWWCB7T` (the adversarial pass "describes survivors as real drift"): not a defect. The sentence before it already makes the refutation pass default to REJECTED unless the category's citation reproduces, so the "real drift" phrasing is a gloss and does not change what a pass does. No change.
Author
Owner

Replying to the review summary (review 01M4DAYW84N056TKJ0SG5ZDJS3, head d9f1187ac86b414683f60032b7a7b00532d2109e)

Pushed 3dd9619: merged main (which moved the sentence again in #126) and re-applied the one phrase main still lacks. Line 37 of skills/verify-doc-drift/SKILL.md now says "re-derive that evidence" instead of "re-derive the evidence from the code", since obvious and duplicate findings are cited from docs. The earlier trial merge was clean; #126 made it conflict.

Finding 01M4DBC94GDDGY33EQ7HRGWSKN (description says findings are "verified against the code"): already fixed on main by #128, which removed the method summary from the description. No change in this PR.

Unadjudicated claim 01M4DBCNZT1W72QA03FBWWCB7T: not a defect, as recorded earlier; the refutation pass defaults to REJECTED unless the category's citation reproduces.

> Replying to the review summary (review `01M4DAYW84N056TKJ0SG5ZDJS3`, head `d9f1187ac86b414683f60032b7a7b00532d2109e`) Pushed 3dd9619: merged main (which moved the sentence again in #126) and re-applied the one phrase main still lacks. Line 37 of `skills/verify-doc-drift/SKILL.md` now says "re-derive that evidence" instead of "re-derive the evidence from the code", since `obvious` and `duplicate` findings are cited from docs. The earlier trial merge was clean; #126 made it conflict. Finding `01M4DBC94GDDGY33EQ7HRGWSKN` (description says findings are "verified against the code"): already fixed on main by #128, which removed the method summary from the description. No change in this PR. Unadjudicated claim `01M4DBCNZT1W72QA03FBWWCB7T`: not a defect, as recorded earlier; the refutation pass defaults to REJECTED unless the category's citation reproduces.
fix: merge main into the verification gate change
All checks were successful
commit-msg / commitlint (pull_request) Successful in 15s
Node tests / node:test (pull_request) Successful in 2m22s
Review / Review (pull_request_target) Successful in 4m35s
3dd961929d
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Author
Owner

Replying to the review summary (review 01M4DR47JSGD8XFTJTFAMH7S7J, head 3dd961929db6c66329103c0cf20afafc3e372780)

No push for this pass: the review of 3dd9619 is complete with no unadjudicated claims, and the trial merge with main is clean.

Finding 01M4DR8EQ4172TVEDX6JA2F51H (the verify-doc-drift description repeats the audit intent and opens with a task summary): valid, but this PR does not touch the description. Tracked in #129, which replaces it with a single "Use when" sentence.

> Replying to the review summary (review `01M4DR47JSGD8XFTJTFAMH7S7J`, head `3dd961929db6c66329103c0cf20afafc3e372780`) No push for this pass: the review of `3dd9619` is complete with no unadjudicated claims, and the trial merge with `main` is clean. Finding `01M4DR8EQ4172TVEDX6JA2F51H` (the `verify-doc-drift` description repeats the audit intent and opens with a task summary): valid, but this PR does not touch the description. Tracked in https://code.j4k.dev/j4k-oss/agent-skills/pulls/129, which replaces it with a single "Use when" sentence.
jercik merged commit 81e603aad9 into main 2026-10-08 18:07:55 +00:00
jercik deleted branch fix/verify-doc-drift-verification-gate 2026-10-08 18:07:55 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
j4k-oss/agent-skills!122
No description provided.