fix(skills): report unavailable release claims as unverified #113

Merged
jercik merged 4 commits from fix/report-unverified-release-claims into main 2026-10-08 18:07:47 +00:00
Owner

Claims whose documented-release source can't be read are now reported as unverified, instead of being called drift or silently left out of the coverage statement.

Addresses the report-location finding in #109.

Claims whose documented-release source can't be read are now reported as unverified, instead of being called drift or silently left out of the coverage statement. Addresses the report-location finding in [#109](https://code.j4k.dev/j4k-oss/agent-skills/pulls/109).
fix: report claims whose release source is unavailable
All checks were successful
commit-msg / commitlint (pull_request) Successful in 27s
Node tests / node:test (pull_request) Successful in 4m25s
Review / Review (pull_request_target) Successful in 5m37s
56a85abed5

Review 01M4DR66EDKBAK9PV7MYR2EB6Z — head b604c074194ba93536540d8f4a17bd3f6f300f1b

Review — j4k-oss/agent-skills @ 2e28781874

Scope: diff against base tree 8e004ccf90c2
Status: dispatched — coverage complete (3/3 slots terminal)
Facts: current review-wide projection

Computed under:

{
  "abandonment": "abandonment-v1",
  "anchor_recipe": 1,
  "batch_policy": "batch-v1",
  "coverage": "coverage-v3",
  "dispatch_policy": "dispatch-v2",
  "grounder_version": 1,
  "grounding_read_rule": "grounding-read-v1",
  "promotion_policy": "promotion-v1",
  "report": "report-v4",
  "tally": "tally-v1",
  "triage_settle": "triage-settle-v2"
}

Findings (1)

low — The catalog description repeats the audit outcome and routing synonyms

  • claim: 01M4DRA8H67T1XCY5YBNNNTKNX
  • anchor: skills/verify-doc-drift/SKILL.md (snippet)
description: Audit a repository's documentation against its source code and fix the factual drift. Use when the user wants to verify documentation matches the code, audit docs for accuracy, hunt doc-vs-code drift, or check that READMEs / CONTEXT.md / ADRs / standards are still true. Triggers on "verify docs", "audit documentation", "doc drift", "do the docs match the code", "check the docs against the code", "are the docs still accurate".
  • lens: writing-quality · arm: default
  • verdicts: 1 valid / 0 invalid / 0 uncertain
    • pass 01M4DRBB777ZKD15T0HRM62C1Q · valid: The exact-grounded snippet shows the description opening with a task summary ('Audit a repository's documentation against its source code and fix the factual drift') before 'Use when', then restating the same intent in a 'Triggers on' list of paraphrases. The installed skill-packaging reference says a capability description is a routing rule that begins with 'Use when …', does not describe the skill's method or contents, and costs catalog context on every load; it also says each literal phrase acts as an independent trigger and belongs only when that phrase alone identifies a match, which generic phrases such as 'audit documentation' do not. The fix direction lives in the body, so cutting it from the description loses no routing boundary. 'Repository documentation' still covers READMEs, CONTEXT.md, ADRs, and standards, so the proposed replacement keeps the selection intent. One caveat: the replacement adds 'including source doc-comments and examples', which the original description does not contain. Only the reviewer's report that they read the body supports that scope, so keep the clause only if the body really covers doc-comments. That does not affect whether the defect is real. The prior adjudication 01M4DR8EQ4172TVEDX6JA2F51H on the same snippet reached valid for the same reasons. Low severity is appropriate.
  • disposition: none

Each catalog reader must process an audit summary followed by two overlapping descriptions of when to select the skill. The extra wording consumes persistent catalog context without supplying another selection boundary.

The description begins “Audit a repository's documentation against its source code and fix the factual drift,” then gives a “Use when” sentence and another sentence of synonymous trigger phrases. The body already defines the audit outcome and fix direction. The installed writing standard's skill-packaging guidance requires a capability description to begin with “Use when” and describe routing rather than its method or contents; “One Idea, One Place” also calls for one home for repeated instructions.

Replace the description with: “Use when the user asks to audit repository documentation against source code, identify factual drift, or correct documentation accuracy, including source doc-comments and examples.” This retains the accuracy-audit boundary and selection intent while leaving fix direction and audit mechanics in the body.

I read the full verify-doc-drift skill, its project-docs dependency, and the README's skill-catalog contract. The README says a capability description supports autonomous selection and its body loads only after selection. This is a wording finding, not an observed selection failure; there is no model-routing reproduction. A distinct routing meaning carried by one of the removed phrases would refute its redundancy and should be retained as a boundary in the concise description.

Other claims

  • grounding-pending (0)
  • ungrounded (0)
  • rejected (0)
  • duplicate-of (0)
  • unadjudicated (0)

Coverage

Coverage pass: 01M4DR6M13QAJ88Q2K9HPDS3AH
Accounting: complete
Slot health: healthy

lens part arm unit status runs loss
general-bug whole default no-claims 1 no
writing-quality whole default claims-emitted 1 no
project-docs whole default no-claims 1 no
  • test-trimming — skipped-by-dispatch: no tests changed; only a skill markdown file edited
  • restated-sets — skipped-by-dispatch: no enumerated sets or lists restated in the change
<!-- review:summary --> **Review** `01M4DR66EDKBAK9PV7MYR2EB6Z` — head `b604c074194ba93536540d8f4a17bd3f6f300f1b` # Review — j4k-oss/agent-skills @ 2e28781874ab Scope: diff against base tree `8e004ccf90c2` Status: dispatched — coverage complete (3/3 slots terminal) Facts: current review-wide projection Computed under: ```json { "abandonment": "abandonment-v1", "anchor_recipe": 1, "batch_policy": "batch-v1", "coverage": "coverage-v3", "dispatch_policy": "dispatch-v2", "grounder_version": 1, "grounding_read_rule": "grounding-read-v1", "promotion_policy": "promotion-v1", "report": "report-v4", "tally": "tally-v1", "triage_settle": "triage-settle-v2" } ``` ## Findings (1) ### low — The catalog description repeats the audit outcome and routing synonyms - claim: `01M4DRA8H67T1XCY5YBNNNTKNX` - anchor: `skills/verify-doc-drift/SKILL.md` (snippet) ``` description: Audit a repository's documentation against its source code and fix the factual drift. Use when the user wants to verify documentation matches the code, audit docs for accuracy, hunt doc-vs-code drift, or check that READMEs / CONTEXT.md / ADRs / standards are still true. Triggers on "verify docs", "audit documentation", "doc drift", "do the docs match the code", "check the docs against the code", "are the docs still accurate". ``` - lens: writing-quality · arm: default - verdicts: 1 valid / 0 invalid / 0 uncertain - pass `01M4DRBB777ZKD15T0HRM62C1Q` · valid: The exact-grounded snippet shows the description opening with a task summary ('Audit a repository's documentation against its source code and fix the factual drift') before 'Use when', then restating the same intent in a 'Triggers on' list of paraphrases. The installed skill-packaging reference says a capability description is a routing rule that begins with 'Use when …', does not describe the skill's method or contents, and costs catalog context on every load; it also says each literal phrase acts as an independent trigger and belongs only when that phrase alone identifies a match, which generic phrases such as 'audit documentation' do not. The fix direction lives in the body, so cutting it from the description loses no routing boundary. 'Repository documentation' still covers READMEs, CONTEXT.md, ADRs, and standards, so the proposed replacement keeps the selection intent. One caveat: the replacement adds 'including source doc-comments and examples', which the original description does not contain. Only the reviewer's report that they read the body supports that scope, so keep the clause only if the body really covers doc-comments. That does not affect whether the defect is real. The prior adjudication 01M4DR8EQ4172TVEDX6JA2F51H on the same snippet reached valid for the same reasons. Low severity is appropriate. - disposition: none > Each catalog reader must process an audit summary followed by two overlapping descriptions of when to select the skill. The extra wording consumes persistent catalog context without supplying another selection boundary. > > The description begins “Audit a repository's documentation against its source code and fix the factual drift,” then gives a “Use when” sentence and another sentence of synonymous trigger phrases. The body already defines the audit outcome and fix direction. The installed writing standard's skill-packaging guidance requires a capability description to begin with “Use when” and describe routing rather than its method or contents; “One Idea, One Place” also calls for one home for repeated instructions. > > Replace the description with: “Use when the user asks to audit repository documentation against source code, identify factual drift, or correct documentation accuracy, including source doc-comments and examples.” This retains the accuracy-audit boundary and selection intent while leaving fix direction and audit mechanics in the body. > > I read the full verify-doc-drift skill, its project-docs dependency, and the README's skill-catalog contract. The README says a capability description supports autonomous selection and its body loads only after selection. This is a wording finding, not an observed selection failure; there is no model-routing reproduction. A distinct routing meaning carried by one of the removed phrases would refute its redundancy and should be retained as a boundary in the concise description. ## Other claims - grounding-pending (0) - ungrounded (0) - rejected (0) - duplicate-of (0) - unadjudicated (0) ## Coverage Coverage pass: 01M4DR6M13QAJ88Q2K9HPDS3AH Accounting: complete Slot health: healthy | lens | part | arm | unit status | runs | loss | | --- | --- | --- | --- | --- | --- | | general-bug | whole | default | no-claims | 1 | no | | writing-quality | whole | default | claims-emitted | 1 | no | | project-docs | whole | default | no-claims | 1 | no | - test-trimming — skipped-by-dispatch: no tests changed; only a skill markdown file edited - restated-sets — skipped-by-dispatch: no enumerated sets or lists restated in the change
@ -27,3 +27,3 @@
## The per-unit audit loop
Split the docs into units — root docs (README, standards, glossary, ADRs) and one unit per package/app (its README + CONTEXT + ADRs + source doc-comments). For each unit: read the docs in full, then read the corresponding source, **following imports to the definition**, and check each claim. For versioned documentation, read the source for its documented release; if that source is unavailable, record a proof gap rather than inferring drift from current code. Require a concrete citation — `file:line` plus a short quoted snippet — for every finding: the contradicting source line for `incorrect`/`code-drift`, the canonical doc location for a `duplicate`, the doc line itself for an `obvious`. No citation, no finding. Returning zero findings for accurate docs is the correct outcome; do not pad.
Split the docs into units — root docs (README, standards, glossary, ADRs) and one unit per package/app (its README + CONTEXT + ADRs + source doc-comments). For each unit: read the docs in full, then read the corresponding source, **following imports to the definition**, and check each claim. For versioned documentation, read the source for its documented release; if that source is unavailable, list the claim as unverified in the report rather than inferring drift from current code. Require a concrete citation — `file:line` plus a short quoted snippet — for every finding: the contradicting source line for `incorrect`/`code-drift`, the canonical doc location for a `duplicate`, the doc line itself for an `obvious`. No citation, no finding. Returning zero findings for accurate docs is the correct outcome; do not pad.

high — Audit-loop citation rule lists every finding category instead of attaching the citation to each category's definition

What I examined: skills/verify-doc-drift/SKILL.md, the paragraph under ## The per-unit audit loop (rewritten by this diff), and the section that defines the finding categories, ## Finding categories, which opens "Classify every discrepancy as exactly one of:" and defines the set as four bullets: - **incorrect** — the code contradicts the doc ..., - **code-drift** — the doc states the intended/correct behavior but the code no longer matches ..., - **obvious** — a self-evident / generic-engineering fact ..., - **duplicate** — the same fact is restated (near-)verbatim in another doc .... No other file in the subject defines these categories.

What the subject says: the changed paragraph requires a citation "for every finding: the contradicting source line for incorrect/code-drift, the canonical doc location for a duplicate, the doc line itself for an obvious." That clause names all four members of the category set, so it is a complete copy of the set defined a few lines above, carrying per-member detail (which citation each category needs).

What goes wrong: today the copy agrees with the source, but the per-category citation rule lives apart from the category definitions. Adding, renaming, or splitting a category under ## Finding categories leaves this clause silently incomplete: a new category would have no citation rule here, and the agent following the skill would have "No citation, no finding" applied to a category the clause never mentions. The same mapping is repeated a second time in the # Output paragraph, so every category change needs two extra edits.

Why no exception applies: the reader (an agent loading this SKILL.md) can read the ## Finding categories section in the same file; the clause is not a table of contents, a dated record, a generated file, or a partial list of examples — it names every member.

Correction: move each category's citation requirement into that category's bullet under ## Finding categories (e.g. append "Cite the contradicting source line." to the incorrect and code-drift bullets, "Cite the canonical doc location." to duplicate, "Cite the doc line itself." to obvious), and reword this sentence to: "Require a concrete citation — file:line plus a short quoted snippet — for every finding, of the kind its category's definition under Finding categories names. No citation, no finding."

Decisive evidence: the four bullets under ## Finding categories versus the four category names in the quoted clause; this is static reading, nothing was run.

lens restated-sets · arm default · tally 1 valid / 0 invalid / 0 uncertain
claim 01M42JE53MH769BAEQ2BTCHQFR of review 01M42JAMB37R0G59RRZTRT3Y60

<!-- review:claim:01M42JE53MH769BAEQ2BTCHQFR --> **high** — Audit-loop citation rule lists every finding category instead of attaching the citation to each category's definition > What I examined: `skills/verify-doc-drift/SKILL.md`, the paragraph under `## The per-unit audit loop` (rewritten by this diff), and the section that defines the finding categories, `## Finding categories`, which opens "Classify every discrepancy as exactly one of:" and defines the set as four bullets: `- **incorrect** — the code contradicts the doc ...`, `- **code-drift** — the doc states the intended/correct behavior but the code no longer matches ...`, `- **obvious** — a self-evident / generic-engineering fact ...`, `- **duplicate** — the same fact is restated (near-)verbatim in another doc ...`. No other file in the subject defines these categories. > > What the subject says: the changed paragraph requires a citation "for every finding: the contradicting source line for `incorrect`/`code-drift`, the canonical doc location for a `duplicate`, the doc line itself for an `obvious`." That clause names all four members of the category set, so it is a complete copy of the set defined a few lines above, carrying per-member detail (which citation each category needs). > > What goes wrong: today the copy agrees with the source, but the per-category citation rule lives apart from the category definitions. Adding, renaming, or splitting a category under `## Finding categories` leaves this clause silently incomplete: a new category would have no citation rule here, and the agent following the skill would have "No citation, no finding" applied to a category the clause never mentions. The same mapping is repeated a second time in the `# Output` paragraph, so every category change needs two extra edits. > > Why no exception applies: the reader (an agent loading this SKILL.md) can read the `## Finding categories` section in the same file; the clause is not a table of contents, a dated record, a generated file, or a partial list of examples — it names every member. > > Correction: move each category's citation requirement into that category's bullet under `## Finding categories` (e.g. append "Cite the contradicting source line." to the **incorrect** and **code-drift** bullets, "Cite the canonical doc location." to **duplicate**, "Cite the doc line itself." to **obvious**), and reword this sentence to: "Require a concrete citation — `file:line` plus a short quoted snippet — for every finding, of the kind its category's definition under Finding categories names. No citation, no finding." > > Decisive evidence: the four bullets under `## Finding categories` versus the four category names in the quoted clause; this is static reading, nothing was run. lens `restated-sets` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain claim `01M42JE53MH769BAEQ2BTCHQFR` of review `01M42JAMB37R0G59RRZTRT3Y60`
jercik marked this conversation as resolved
@ -74,3 +74,3 @@
# Output
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State the clean bill of health too — which claims verified — so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State which claims verified and which could not be checked, so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.

low — Changed text calls the same uncheckable claims "unverified" in the audit loop but "could not be checked" in Output

Examined: the two sentences this diff changes in skills/verify-doc-drift/SKILL.md, and how they connect.

The audit loop now says: "For versioned documentation, read the source for its documented release; if that source is unavailable, list the claim as unverified in the report rather than inferring drift from current code." The Output section, which defines the report, now says: "State which claims verified and which could not be checked, so the user knows the coverage." The Output paragraph opens with "A severity-ranked report: for each confirmed finding give …".

So one concept, a claim the agent could not check against the right source, has two names, "unverified" and "could not be checked". The writing skill's "Use Precise Language" section says to "use one term for one concept". Here is the plausible misreading: the loop says to list the claim "in the report", and the Output's word "could not be checked" doesn't visibly match "unverified". An agent can then put the unverified claim into the severity-ranked findings list, giving it a category and severity it has no evidence for, instead of into the coverage statement. That's exactly the drift-inferring the loop forbids. An agent could also keep two separate lists for what is one category.

Correction (the smallest one): use the loop's term in Output: "State which claims verified and which are unverified, so the user knows the coverage." Both changed sentences then plainly refer to the same report section, and nothing is added or lost.

lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain
claim 01M42JFG15QBTYVBQ61AWK7SVJ of review 01M42JAMB37R0G59RRZTRT3Y60

<!-- review:claim:01M42JFG15QBTYVBQ61AWK7SVJ --> **low** — Changed text calls the same uncheckable claims "unverified" in the audit loop but "could not be checked" in Output > Examined: the two sentences this diff changes in skills/verify-doc-drift/SKILL.md, and how they connect. > > The audit loop now says: "For versioned documentation, read the source for its documented release; if that source is unavailable, list the claim as unverified in the report rather than inferring drift from current code." The Output section, which defines the report, now says: "State which claims verified and which could not be checked, so the user knows the coverage." The Output paragraph opens with "A severity-ranked report: for each confirmed finding give …". > > So one concept, a claim the agent could not check against the right source, has two names, "unverified" and "could not be checked". The writing skill's "Use Precise Language" section says to "use one term for one concept". Here is the plausible misreading: the loop says to list the claim "in the report", and the Output's word "could not be checked" doesn't visibly match "unverified". An agent can then put the unverified claim into the severity-ranked findings list, giving it a category and severity it has no evidence for, instead of into the coverage statement. That's exactly the drift-inferring the loop forbids. An agent could also keep two separate lists for what is one category. > > Correction (the smallest one): use the loop's term in Output: "State which claims verified and which are unverified, so the user knows the coverage." Both changed sentences then plainly refer to the same report section, and nothing is added or lost. lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain claim `01M42JFG15QBTYVBQ61AWK7SVJ` of review `01M42JAMB37R0G59RRZTRT3Y60`
Author
Owner

Fixed in 4e4199e72d: audit loop and coverage output both call unavailable-source claims unverified.

<!-- gh-feedback:reply-to:112183 --> Fixed in 4e4199e72d6b5f8a31f82661684ddaaa37d06b7f: audit loop and coverage output both call unavailable-source claims unverified.
jercik marked this conversation as resolved
fix: use one term for unverified release claims
All checks were successful
commit-msg / commitlint (pull_request) Successful in 23s
Node tests / node:test (pull_request) Successful in 3m10s
Review / Review (pull_request_target) Successful in 3m35s
4e4199e72d
Author
Owner

Replying to review comment #112182

Citation-map duplication (claim 01M42JE53MH769BAEQ2BTCHQFR, including duplicate claims 01M42JEE9DD8J64X5QPYR7HQ8Y and 01M42JF93ZG8KQANNP3ZSMVVYP) is pre-existing and verified as repeated category/citation wording. It is scope-deferred outside this bounded unavailable-release-source report repair. No category-definition restructuring is adopted, and no fix or disproof is claimed.

Replying to review comment #112181

Current review 01M42JAMB37R0G59RRZTRT3Y60 report-only medium 01M42JFS1SCEG9B66VV07DQFS8: tracked in #114. Actual main already assigns ADR changes to project-docs; the child only declares delivery and explicitly loads that existing contract after the user changes a decision. It targets main independently and preserves accepted/status-less/proposed ADR authority and the user's gate. It is not merged; #113 is not claimed fixed by the child.

Report-only high 01M42JG3HEXSGXYNDJWAVZXN73: the existing frontmatter repeats categories and describes method. This is the same packaging cleanup already retained as scope-deferred on #110, outside minimal composition. It remains unfixed, not disproved.

Report-only low 01M42JGBB4X7JJFBV97H3Q6XM2: opening fix-direction wording repeats its detailed section. It is pre-existing wording cleanup and scope-deferred; no authority rule is removed and no fix is claimed.

The three report-only claims have no corresponding inline anchors in the stable surface. No per-claim native transitions are invented. Scoped low 01M42JFG15QBTYVBQ61AWK7SVJ is fixed in 4e4199e72d6b5f8a31f82661684ddaaa37d06b7f: the coverage statement and audit loop both use "unverified". Review of that new head is still required.

> Replying to review comment #112182 Citation-map duplication (claim `01M42JE53MH769BAEQ2BTCHQFR`, including duplicate claims `01M42JEE9DD8J64X5QPYR7HQ8Y` and `01M42JF93ZG8KQANNP3ZSMVVYP`) is pre-existing and verified as repeated category/citation wording. It is scope-deferred outside this bounded unavailable-release-source report repair. No category-definition restructuring is adopted, and no fix or disproof is claimed. > Replying to review comment #112181 Current review `01M42JAMB37R0G59RRZTRT3Y60` report-only medium `01M42JFS1SCEG9B66VV07DQFS8`: tracked in [#114](https://code.j4k.dev/j4k-oss/agent-skills/pulls/114). Actual main already assigns ADR changes to `project-docs`; the child only declares delivery and explicitly loads that existing contract after the user changes a decision. It targets main independently and preserves accepted/status-less/proposed ADR authority and the user's gate. It is not merged; #113 is not claimed fixed by the child. Report-only high `01M42JG3HEXSGXYNDJWAVZXN73`: the existing frontmatter repeats categories and describes method. This is the same packaging cleanup already retained as scope-deferred on #110, outside minimal composition. It remains unfixed, not disproved. Report-only low `01M42JGBB4X7JJFBV97H3Q6XM2`: opening fix-direction wording repeats its detailed section. It is pre-existing wording cleanup and scope-deferred; no authority rule is removed and no fix is claimed. The three report-only claims have no corresponding inline anchors in the stable surface. No per-claim native transitions are invented. Scoped low `01M42JFG15QBTYVBQ61AWK7SVJ` is fixed in `4e4199e72d6b5f8a31f82661684ddaaa37d06b7f`: the coverage statement and audit loop both use "unverified". Review of that new head is still required.
@ -74,3 +74,3 @@
# Output
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State the clean bill of health too — which claims verified — so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State which claims verified and which are unverified, so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.

high — Output and the audit loop each restate the per-category citation mapping across all four finding categories

What I examined: skills/verify-doc-drift/SKILL.md in full. The diff modifies the audit-loop paragraph ("## The per-unit audit loop") and the Output paragraph. Both modified paragraphs carry the same mapping from every finding category to the citation it requires. The categories are defined in one place, the "## Finding categories" list, which has bullets for incorrect, code-drift, obvious, and duplicate.

What the subject says: The audit loop says "Require a concrete citation — file:line plus a short quoted snippet — for every finding: the contradicting source line for incorrect/code-drift, the canonical doc location for a duplicate, the doc line itself for an obvious." Output says "its citation (contradicting code for incorrect/code-drift, the canonical doc for a duplicate, the doc line for an obvious)". Both passages name every member of the category set that the Finding categories list defines, and they already word it differently ("contradicting source line" vs "contradicting code", "canonical doc location" vs "canonical doc").

What goes wrong: When a category is added or its evidence rule changes, three places must change together: the Finding categories bullet, the audit loop, and Output. If one is missed, an agent running this skill gets conflicting rules about which evidence makes a finding valid. The skill itself teaches this rule in "Restated sets: point instead of recopying". The writing standard's "One Idea, One Place" says to give each instruction one home and keep a term's definition, rule, and caveat together.

Correction: Put each category's citation requirement into that category's own bullet in Finding categories. For example, add "Cite the contradicting source line." to incorrect and code-drift, "Cite the canonical doc location." to duplicate, and "Cite the doc line itself." to obvious. In the audit loop, shorten the sentence to "Require a concrete citation — file:line plus a short quoted snippet — of the kind the finding's category names." In Output, replace the parenthetical with plain "its citation". Each requirement keeps its meaning but has one home, sitting next to the definition the agent reads when classifying a finding.

Evidence: static reading of the file. The three passages quoted above are the whole of the duplication.

lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain
claim 01M42JW3A9KEM17AVHRYYPQDSE of review 01M42JSA8391GTPPAJHXEX3YJV

<!-- review:claim:01M42JW3A9KEM17AVHRYYPQDSE --> **high** — Output and the audit loop each restate the per-category citation mapping across all four finding categories > What I examined: skills/verify-doc-drift/SKILL.md in full. The diff modifies the audit-loop paragraph ("## The per-unit audit loop") and the Output paragraph. Both modified paragraphs carry the same mapping from every finding category to the citation it requires. The categories are defined in one place, the "## Finding categories" list, which has bullets for **incorrect**, **code-drift**, **obvious**, and **duplicate**. > > What the subject says: The audit loop says "Require a concrete citation — `file:line` plus a short quoted snippet — for every finding: the contradicting source line for `incorrect`/`code-drift`, the canonical doc location for a `duplicate`, the doc line itself for an `obvious`." Output says "its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`)". Both passages name every member of the category set that the Finding categories list defines, and they already word it differently ("contradicting source line" vs "contradicting code", "canonical doc location" vs "canonical doc"). > > What goes wrong: When a category is added or its evidence rule changes, three places must change together: the Finding categories bullet, the audit loop, and Output. If one is missed, an agent running this skill gets conflicting rules about which evidence makes a finding valid. The skill itself teaches this rule in "Restated sets: point instead of recopying". The writing standard's "One Idea, One Place" says to give each instruction one home and keep a term's definition, rule, and caveat together. > > Correction: Put each category's citation requirement into that category's own bullet in Finding categories. For example, add "Cite the contradicting source line." to incorrect and code-drift, "Cite the canonical doc location." to duplicate, and "Cite the doc line itself." to obvious. In the audit loop, shorten the sentence to "Require a concrete citation — `file:line` plus a short quoted snippet — of the kind the finding's category names." In Output, replace the parenthetical with plain "its citation". Each requirement keeps its meaning but has one home, sitting next to the definition the agent reads when classifying a finding. > > Evidence: static reading of the file. The three passages quoted above are the whole of the duplication. lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain claim `01M42JW3A9KEM17AVHRYYPQDSE` of review `01M42JSA8391GTPPAJHXEX3YJV`
jercik marked this conversation as resolved
Author
Owner

Replying to review comment #112271

Current claim 01M42JW3A9KEM17AVHRYYPQDSE and its two duplicate-of claims repeat the verified pre-existing citation-mapping cleanup already recorded for #112182. This remains scope-deferred outside the bounded unavailable-release-source coverage repair. No category/citation restructuring is adopted, and the repeated mappings remain unfixed. This disposition records the explicit user scope limitation, not empirical disproof or a correction delivered by this PR.

> Replying to review comment #112271 Current claim `01M42JW3A9KEM17AVHRYYPQDSE` and its two duplicate-of claims repeat the verified pre-existing citation-mapping cleanup already recorded for #112182. This remains scope-deferred outside the bounded unavailable-release-source coverage repair. No category/citation restructuring is adopted, and the repeated mappings remain unfixed. This disposition records the explicit user scope limitation, not empirical disproof or a correction delivered by this PR.
chore: merge fix/version-matched-source-pointers into fix/report-unverified-release-claims
All checks were successful
commit-msg / commitlint (pull_request) Successful in 22s
Node tests / node:test (pull_request) Successful in 3m1s
Review / Review (pull_request_target) Successful in 9m38s
cb70087a88
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@ -80,3 +80,3 @@
# Output
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State the clean bill of health too — which claims verified — so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State which claims verified and which are unverified, so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.

medium — Citation guidance recopies the finding-category set

The report instruction creates another maintenance site for the category taxonomy: changing the categories requires updating this citation list as well as their definitions, and an omitted update silently leaves report guidance stale. The per-unit audit paragraph repeats the same mapping.

The defining source is the “Finding categories” section of skills/verify-doc-drift/SKILL.md. It defines incorrect, code-drift, obvious, and duplicate. The anchored copy covers incorrect, code-drift, duplicate, and obvious; no member appears in only one list, so the copy currently agrees. The parenthetical enumerates the complete category set while attaching the required citation to each member.

Move each category’s citation requirement into its definition in “Finding categories”. Replace the parenthetical with “using the citation required by the Finding categories section”, and make the per-unit audit instruction use that same pointer. This preserves the citation detail without recopied membership.

I read the complete skill, its project-docs dependency, and the served diff. This is a static comparison of the changed report and audit paragraphs with the category definitions. The copy is neither a dated record nor a table of contents, example subset, external standard, or third-party contract. The agent can read the defining section in the same skill, and those definitions can hold the citation detail. Although the definitions supply instructions the agent reads, this separate enumeration is a copy of them.

The decisive evidence is that the defined category membership and the two citation enumerations coincide at this revision. A distribution boundary that withholds the defining section from the report-writing agent would establish an exception; none is present in the examined skill.

lens restated-sets · arm default · tally 1 valid / 0 invalid / 0 uncertain
claim 01M484A3EY918NDMKQ95QRPVSV of review 01M4843PG2WYFJPWKJVXKCPX7V

<!-- review:claim:01M484A3EY918NDMKQ95QRPVSV --> **medium** — Citation guidance recopies the finding-category set > The report instruction creates another maintenance site for the category taxonomy: changing the categories requires updating this citation list as well as their definitions, and an omitted update silently leaves report guidance stale. The per-unit audit paragraph repeats the same mapping. > > The defining source is the “Finding categories” section of skills/verify-doc-drift/SKILL.md. It defines incorrect, code-drift, obvious, and duplicate. The anchored copy covers incorrect, code-drift, duplicate, and obvious; no member appears in only one list, so the copy currently agrees. The parenthetical enumerates the complete category set while attaching the required citation to each member. > > Move each category’s citation requirement into its definition in “Finding categories”. Replace the parenthetical with “using the citation required by the Finding categories section”, and make the per-unit audit instruction use that same pointer. This preserves the citation detail without recopied membership. > > I read the complete skill, its project-docs dependency, and the served diff. This is a static comparison of the changed report and audit paragraphs with the category definitions. The copy is neither a dated record nor a table of contents, example subset, external standard, or third-party contract. The agent can read the defining section in the same skill, and those definitions can hold the citation detail. Although the definitions supply instructions the agent reads, this separate enumeration is a copy of them. > > The decisive evidence is that the defined category membership and the two citation enumerations coincide at this revision. A distribution boundary that withholds the defining section from the report-writing agent would establish an exception; none is present in the examined skill. lens `restated-sets` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain claim `01M484A3EY918NDMKQ95QRPVSV` of review `01M4843PG2WYFJPWKJVXKCPX7V`
jercik marked this conversation as resolved
Author
Owner

Replying to review comment #133709

Claim 01M484A3EY918NDMKQ95QRPVSV (and its duplicate 01M484C5BWJX849DQSREEPZ777) is valid and pre-existing: the citation mapping is restated in the audit loop and in Output. Tracked in #126, which moves each category's citation into its definition and points the audit loop at it. Output is covered by #121. Not fixed in this PR.

Report-only claims of review 01M4843PG2WYFJPWKJVXKCPX7V (head cb70087):

Replying to review #112181

  • 01M484ATDPD0NT22JB7SXTP4B7 (description recopies the categories): pre-existing, tracked in #119.
  • 01M484BBF5KCK3959HK889J1MF (verification gate rejects documentary evidence): pre-existing, tracked in #121, #122 and #124, which overlap.

None is fixed in #113. This PR stays stacked on #109 and is unchanged.

> Replying to review comment #133709 Claim `01M484A3EY918NDMKQ95QRPVSV` (and its duplicate `01M484C5BWJX849DQSREEPZ777`) is valid and pre-existing: the citation mapping is restated in the audit loop and in `Output`. Tracked in [#126](https://code.j4k.dev/j4k-oss/agent-skills/pulls/126), which moves each category's citation into its definition and points the audit loop at it. `Output` is covered by [#121](https://code.j4k.dev/j4k-oss/agent-skills/pulls/121). Not fixed in this PR. Report-only claims of review `01M4843PG2WYFJPWKJVXKCPX7V` (head `cb70087`): > Replying to review #112181 - `01M484ATDPD0NT22JB7SXTP4B7` (description recopies the categories): pre-existing, tracked in [#119](https://code.j4k.dev/j4k-oss/agent-skills/pulls/119). - `01M484BBF5KCK3959HK889J1MF` (verification gate rejects documentary evidence): pre-existing, tracked in [#121](https://code.j4k.dev/j4k-oss/agent-skills/pulls/121), [#122](https://code.j4k.dev/j4k-oss/agent-skills/pulls/122) and [#124](https://code.j4k.dev/j4k-oss/agent-skills/pulls/124), which overlap. None is fixed in #113. This PR stays stacked on #109 and is unchanged.
jercik changed target branch from fix/version-matched-source-pointers to main 2026-10-08 08:20:51 +00:00
jercik force-pushed fix/report-unverified-release-claims from cb70087a88
All checks were successful
commit-msg / commitlint (pull_request) Successful in 22s
Node tests / node:test (pull_request) Successful in 3m1s
Review / Review (pull_request_target) Successful in 9m38s
to 595fbb4569
All checks were successful
commit-msg / commitlint (pull_request) Successful in 23s
Node tests / node:test (pull_request) Successful in 3m6s
Review / Review (pull_request_target) Successful in 6m22s
2026-10-08 08:24:34 +00:00
Compare
@ -82,3 +82,3 @@
# Output
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State the clean bill of health too — which claims verified — so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.
A severity-ranked report: for each confirmed finding give the doc `file:line`, the category, the claim, its citation (contradicting code for `incorrect`/`code-drift`, the canonical doc for a `duplicate`, the doc line for an `obvious`), and the precise fix. Group by severity (a non-compiling example or a runtime-crashing snippet outranks a stale dependency label). State which claims verified and which are unverified, so the user knows the coverage. Then apply the confirmed fixes; open a PR only when the user asked for one rather than local edits or work inside an existing PR.

medium — The output contract recopies the category-specific citation rules

Maintaining the evidence contract requires editing both the audit instructions and the report instructions. A later change to the required evidence can leave the report following an outdated copy even when the audit uses the current rule.

‘The per-unit audit loop’ defines the citation rules: incorrect and code-drift require contradicting source evidence, duplicate requires the canonical documentation location, and obvious requires the documentation passage itself. The Output paragraph repeats the same category-to-evidence mapping. The source's members are incorrect, code-drift, duplicate, and obvious; the copy covers incorrect, code-drift, duplicate, and obvious. Neither list has a member absent from the other, and their evidence requirements agree at this revision.

Replace the anchored phrase with ‘the citation required by the per-unit audit loop’. This keeps the report's obligation to include evidence while leaving the classification-specific requirements in their existing authoritative home. It follows the writing standard's ‘One Idea, One Place’ guidance and the review's requirement to replace a restated set with a source pointer.

I compared the complete per-unit audit loop, adversarial verification section, and Output paragraph in verify-doc-drift/SKILL.md. The adversarial verification section already uses a pointer to the per-unit audit loop for this same evidence contract, demonstrating that the file can express the requirement without copying the mapping. This is a static writing assessment; no audit execution was needed. A requirement that Output must be distributed separately without the audit-loop instructions would refute the proposed pointer, but the examined skill states no such delivery contract.

lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain
claim 01M4D9Z1WTWV4X4G2R6Z8MWRYX of review 01M4D9TJBR3XEQ2BG1FXN2T6WF

<!-- review:claim:01M4D9Z1WTWV4X4G2R6Z8MWRYX --> **medium** — The output contract recopies the category-specific citation rules > Maintaining the evidence contract requires editing both the audit instructions and the report instructions. A later change to the required evidence can leave the report following an outdated copy even when the audit uses the current rule. > > ‘The per-unit audit loop’ defines the citation rules: incorrect and code-drift require contradicting source evidence, duplicate requires the canonical documentation location, and obvious requires the documentation passage itself. The Output paragraph repeats the same category-to-evidence mapping. The source's members are incorrect, code-drift, duplicate, and obvious; the copy covers incorrect, code-drift, duplicate, and obvious. Neither list has a member absent from the other, and their evidence requirements agree at this revision. > > Replace the anchored phrase with ‘the citation required by the per-unit audit loop’. This keeps the report's obligation to include evidence while leaving the classification-specific requirements in their existing authoritative home. It follows the writing standard's ‘One Idea, One Place’ guidance and the review's requirement to replace a restated set with a source pointer. > > I compared the complete per-unit audit loop, adversarial verification section, and Output paragraph in verify-doc-drift/SKILL.md. The adversarial verification section already uses a pointer to the per-unit audit loop for this same evidence contract, demonstrating that the file can express the requirement without copying the mapping. This is a static writing assessment; no audit execution was needed. A requirement that Output must be distributed separately without the audit-loop instructions would refute the proposed pointer, but the examined skill states no such delivery contract. lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain claim `01M4D9Z1WTWV4X4G2R6Z8MWRYX` of review `01M4D9TJBR3XEQ2BG1FXN2T6WF`
Author
Owner

Already fixed on main in de47351 (#127): the Output paragraph now says "its citation (as the per-unit audit loop requires)". Merged into this branch in 0c1b135. The companion catalog-description finding in the same report (claim 01M4D9Y9SVRXEVMVPQYJ008DNH) is likewise already fixed on main in 664392a (#119).

<!-- gh-feedback:reply-to:141255 --> Already fixed on main in de47351 (#127): the Output paragraph now says "its citation (as the per-unit audit loop requires)". Merged into this branch in 0c1b135. The companion catalog-description finding in the same report (claim 01M4D9Y9SVRXEVMVPQYJ008DNH) is likewise already fixed on main in 664392a (#119).
jercik marked this conversation as resolved
fix: merge main into report-unverified-release-claims
All checks were successful
commit-msg / commitlint (pull_request) Successful in 21s
Node tests / node:test (pull_request) Successful in 2m57s
Review / Review (pull_request_target) Successful in 8m53s
0c1b135b5b
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@ -31,3 +31,3 @@
## The per-unit audit loop
Split the docs into units — root docs (README, standards, glossary, ADRs) and one unit per package/app (its README + CONTEXT + ADRs + source doc-comments). For each unit: read the docs in full, then read the corresponding source, **following imports to the definition**, and check each claim. For versioned documentation, read the source for its documented release; if that source is unavailable, record a proof gap rather than inferring drift from current code. Require a concrete citation — `file:line` plus a short quoted snippet — for every finding: the contradicting source line for `incorrect`/`code-drift`, the canonical doc location for a `duplicate`, the doc line itself for an `obvious`. No citation, no finding. Returning zero findings for accurate docs is the correct outcome; do not pad.
Split the docs into units — root docs (README, standards, glossary, ADRs) and one unit per package/app (its README + CONTEXT + ADRs + source doc-comments). For each unit: read the docs in full, then read the corresponding source, **following imports to the definition**, and check each claim. For versioned documentation, read the source for its documented release; if that source is unavailable, list the claim as unverified in the report rather than inferring drift from current code. Require a concrete citation — `file:line` plus a short quoted snippet — for every finding: the contradicting source line for `incorrect`/`code-drift`, the canonical doc location for a `duplicate`, the doc line itself for an `obvious`. No citation, no finding. Returning zero findings for accurate docs is the correct outcome; do not pad.

medium — The citation rule recopies the complete finding-category set

The audit's category vocabulary has a second maintenance home: changing a category requires finding and updating this citation rule as well as its definition. A new category can otherwise lack a stated evidence requirement while the audit still requires evidence for every finding.

The authoritative Finding categories section defines incorrect, code-drift, obvious, and duplicate. This passage covers incorrect, code-drift, duplicate, and obvious. No member appears in only one list; the copy currently agrees. The restated-sets guideline nevertheless treats an agreeing copy as a defect, and the writing standard's One Idea, One Place guidance puts each rule with its definition.

Move each category's specific citation requirement into its existing definition. Replace this passage with: ‘Require a concrete citation — file:line plus a short quoted snippet — using the evidence requirement in Finding categories.’ This retains the citation format and every category's evidence requirement without maintaining another category list. Keep the surrounding no-citation/no-finding and zero-findings rules.

I compared the category definitions, the per-unit audit loop, and the adversarial-verification paragraph in skills/verify-doc-drift/SKILL.md. This is a static comparison of instructions, not an executed audit. The matching definitions and copied list establish the duplicate set; a generated relationship or an audience unable to access the definitions would refute the maintenance premise, but neither is declared here.

lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain
claim 01M4DKV6E92QC95RE571H0G8JY of review 01M4DKQQ14XP5JAJ8GS3PMER2V

<!-- review:claim:01M4DKV6E92QC95RE571H0G8JY --> **medium** — The citation rule recopies the complete finding-category set > The audit's category vocabulary has a second maintenance home: changing a category requires finding and updating this citation rule as well as its definition. A new category can otherwise lack a stated evidence requirement while the audit still requires evidence for every finding. > > The authoritative Finding categories section defines incorrect, code-drift, obvious, and duplicate. This passage covers incorrect, code-drift, duplicate, and obvious. No member appears in only one list; the copy currently agrees. The restated-sets guideline nevertheless treats an agreeing copy as a defect, and the writing standard's One Idea, One Place guidance puts each rule with its definition. > > Move each category's specific citation requirement into its existing definition. Replace this passage with: ‘Require a concrete citation — `file:line` plus a short quoted snippet — using the evidence requirement in Finding categories.’ This retains the citation format and every category's evidence requirement without maintaining another category list. Keep the surrounding no-citation/no-finding and zero-findings rules. > > I compared the category definitions, the per-unit audit loop, and the adversarial-verification paragraph in skills/verify-doc-drift/SKILL.md. This is a static comparison of instructions, not an executed audit. The matching definitions and copied list establish the duplicate set; a generated relationship or an audience unable to access the definitions would refute the maintenance premise, but neither is declared here. lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain claim `01M4DKV6E92QC95RE571H0G8JY` of review `01M4DKQQ14XP5JAJ8GS3PMER2V`
Author
Owner

Already fixed on main in d152f38 (#126), which moved each category's citation into its definition and left this rule pointing at "Finding categories". Merged into this branch in b604c07; the sentence no longer lists the categories.

<!-- gh-feedback:reply-to:142967 --> Already fixed on main in d152f38 (#126), which moved each category's citation into its definition and left this rule pointing at "Finding categories". Merged into this branch in b604c07; the sentence no longer lists the categories.
jercik marked this conversation as resolved
fix: merge main into report-unverified-release-claims
All checks were successful
commit-msg / commitlint (pull_request) Successful in 29s
Review / Review (pull_request_target) Successful in 3m37s
Node tests / node:test (pull_request) Successful in 4m21s
b604c07419
Author
Owner

Replying to review #112181

Report-only low 01M4DRA8H67T1XCY5YBNNNTKNX of review 01M4DR66EDKBAK9PV7MYR2EB6Z (head b604c07): the verify-doc-drift description does open with a task summary and restate the intent in the "Triggers on" list. I verified this against current main, where the text is unchanged. It is wording in lines this PR does not touch, so it is not fixed here. It is tracked in #129, which rewrites the description to state only when to use the skill. No fix is claimed in #113, and #129 is not merged.

> Replying to review #112181 Report-only low `01M4DRA8H67T1XCY5YBNNNTKNX` of review `01M4DR66EDKBAK9PV7MYR2EB6Z` (head `b604c07`): the `verify-doc-drift` description does open with a task summary and restate the intent in the "Triggers on" list. I verified this against current `main`, where the text is unchanged. It is wording in lines this PR does not touch, so it is not fixed here. It is tracked in [#129](https://code.j4k.dev/j4k-oss/agent-skills/pulls/129), which rewrites the description to state only when to use the skill. No fix is claimed in #113, and #129 is not merged.
jercik merged commit 58a6d734ed into main 2026-10-08 18:07:47 +00:00
jercik deleted branch fix/report-unverified-release-claims 2026-10-08 18:07:48 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
j4k-oss/agent-skills!113
No description provided.