fix(skills): dated records should use their release evidence #110

Merged
jercik merged 2 commits from fix/dates-audit-release-evidence into main 2026-10-07 22:00:47 +00:00
Owner

Moves the dated-record caveat in verify-doc-drift next to claim classification, where an auditor reads it before deciding what counts as drift, and drops the blanket "documentation rots" opening sentence it contradicts. Addresses the clarification in #97, which has since merged.

🤖 Generated with Claude Code

Moves the dated-record caveat in `verify-doc-drift` next to claim classification, where an auditor reads it before deciding what counts as drift, and drops the blanket "documentation rots" opening sentence it contradicts. Addresses the clarification in [#97](https://code.j4k.dev/j4k-oss/agent-skills/pulls/97#issuecomment-111906), which has since merged. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
fix(skills): dated records should use their release evidence
All checks were successful
commit-msg / commitlint (pull_request) Successful in 24s
Node tests / node:test (pull_request) Successful in 3m4s
Review / Review (pull_request_target) Successful in 5m20s
4f43d63aa3

Review 01M482S4AGS61A02RY67T8Y9Z2 — head 8ee8e3675ab95c27cb9b0c453199306101ae674c

Review — j4k-oss/agent-skills @ eac2f079b8

Scope: diff against base tree bc96a7769926
Status: dispatched — coverage complete (5/5 slots terminal)
Facts: current review-wide projection

Computed under:

{
  "abandonment": "abandonment-v1",
  "anchor_recipe": 1,
  "batch_policy": "batch-v1",
  "coverage": "coverage-v3",
  "dispatch_policy": "dispatch-v2",
  "grounder_version": 1,
  "grounding_read_rule": "grounding-read-v1",
  "promotion_policy": "promotion-v1",
  "report": "report-v4",
  "tally": "tally-v1",
  "triage_settle": "triage-settle-v2"
}

Findings (2)

medium — The catalog description copies the authoritative finding-category set

  • claim: 01M48322K8DAMB4RX5JHJEERTR
  • anchor: skills/verify-doc-drift/SKILL.md (snippet)
Each finding is categorized (incorrect / code-drift / obvious / duplicate) and adversarially verified against the code before it is reported, then docs are fixed to match the code (the reverse only when the doc is the intended source of truth).
  • lens: writing-quality · arm: default
  • verdicts: 1 valid / 0 invalid / 0 uncertain
    • pass 01M4833PDTKGW4ZP1SW16CFJF5 · valid: The exact-grounded description copies the category members. The current body names the source path and its Finding categories instruction, lists the defining entries, and compares them with the copy: both contain incorrect, code-drift, obvious, and duplicate. That specific reviewer-reported comparison meets the restated-sets standard; the router has no need to use these members and can load the defining body, so no exception applies. Deleting the anchored method-summary sentence preserves routing text and leaves the reported verification and fix-direction requirements in their operative sections, consistent with skill-packaging. Reassessing the historical verdicts, their supported copying rationale still applies, but the earlier high-severity label does not: an agreeing copy is medium under the present guideline. The current correction also avoids the earlier full-description replacement concern about losing duplicate/obsolete-doc routing scope.
  • disposition: none

Changing the audit's classification now requires updating both its operative definitions and the catalog description. The second copy can silently retain obsolete categories even though readers of this skill can open the defining section.

The source is “Finding categories” in this file, whose “Classify every discrepancy as exactly one of” instruction defines incorrect, code-drift, obvious, and duplicate. The frontmatter parenthetical copies incorrect, code-drift, obvious, and duplicate. The lists currently agree; neither contains a member absent from the other. The restated-sets guideline treats an agreeing copy as a defect because it creates a second maintenance obligation. The installed writing standard's “One Idea, One Place” guidance also puts each instruction in one home, and its packaging reference reserves capability descriptions for routing rather than methods.

Delete the anchored method-summary sentence from the description. Keep the request-based routing text; the “Finding categories” section remains the authoritative definition, and the body's verification and fix-direction sections retain the useful procedural meaning.

I read the complete skill, its declared project-docs dependency, and the README's catalog-loading explanation, and compared the parenthetical with the category definitions directly. This is a static comparison, not an execution test. The defining classification instruction and the identical parenthetical establish the copy; it would be refuted if the description were generated from that definition or its intended reader could not access the body, neither of which is indicated in the examined files.

medium — The universal verification gate excludes valid documentary findings

  • claim: 01M4832MYZS7687PYDQGMW7WT1
  • anchor: skills/verify-doc-drift/SKILL.md (snippet)
Every candidate finding gets a second, independent pass that tries to **refute** it and defaults to REJECTED unless it can point at the contradicting line itself — re-derive the evidence from the code rather than trusting the first pass. This is what separates real drift from a plausible misread; skipping it ships false corrections. Only survivors reach the report.
  • lens: writing-quality · arm: default
  • verdicts: 1 valid / 0 invalid / 0 uncertain
    • pass 01M4833PDTKGW4ZP1SW16CFJF5 · valid: The exact-grounded gate applies to Every candidate finding, defaults to rejection without the contradicting line, and requires evidence re-derived from the code. The reviewer supplies a specific surrounding instruction trace: The per-unit audit loop accepts a canonical document location for a duplicate and the passage itself for an obvious finding, while the other verification sections supply no category-specific exception. An accurate fact duplicated in documentation can satisfy the reported duplicate criterion without any contradicting code line, so the universal gate excludes an otherwise eligible finding. This is a concrete evidence-contract conflict under Specify the Discipline and Use Precise Language, not a demand for an unnecessary exception. The proposed anchored rewrite preserves independent refutation, fresh evidence, default rejection and the reporting boundary while pointing to the existing citation requirement without restating its set. Reported model behavior remains unmeasured, but that does not defeat the instruction conflict; medium severity is appropriate.
  • disposition: none

A verifier following this gate literally rejects valid documentary findings that have no contradicting code. For example, the same accurate fact duplicated in two documents satisfies the audit's duplicate criterion but cannot supply a source line contradicting that fact. Those findings disappear before reporting or correction.

“The per-unit audit loop” explicitly accepts the canonical document location as evidence for a duplicate and the document passage itself for an obvious finding. In contrast, the anchored rule applies to “Every candidate finding” and defaults to rejection unless the verifier can point at “the contradicting line itself”, then requires evidence re-derived “from the code”. The documentary evidence allowed by the earlier contract cannot satisfy that universal gate. The installed writing standard's “Specify the Discipline” and “Use Precise Language” guidance requires checkable evidence appropriate to the task and conditions scoped to the actions they govern.

Replace the gate with: “Every candidate finding gets a second, independent pass that tries to refute it and defaults to REJECTED unless fresh evidence satisfies the citation requirement in ‘The per-unit audit loop’. Re-derive that evidence rather than trusting the first pass. Only survivors reach the report.” This retains independent verification, rejection of unsupported findings, and the reporting boundary while referring to the existing evidence contract rather than copying its category set.

I traced “Finding categories”, “The per-unit audit loop”, “Adversarial verification”, the Task's verification step, and Output in the complete skill. None scopes the contradicting-code requirement to factual contradictions or exempts documentary findings. This is a static instruction conflict; I did not run an audit with a served model and cannot establish how often a model would resolve the conflict correctly on its own. A category-scoped gate or an explicit category-appropriate evidence requirement would refute the textual finding; successful documentary verification in a model trial would qualify the practical impact without removing the conflicting instructions.

Other claims

  • grounding-pending (0)
  • ungrounded (0)
  • rejected (0)
  • duplicate-of (0)
  • unadjudicated (0)

Coverage

Coverage pass: 01M482S51F18NVSNJGVWX5Z2PB
Accounting: complete
Slot health: healthy

lens part arm unit status runs loss
general-bug whole default no-claims 1 no
writing-quality whole default claims-emitted 1 no
test-trimming whole default no-claims 1 no
restated-sets whole default no-claims 1 no
project-docs whole default no-claims 1 no
<!-- review:summary --> **Review** `01M482S4AGS61A02RY67T8Y9Z2` — head `8ee8e3675ab95c27cb9b0c453199306101ae674c` # Review — j4k-oss/agent-skills @ eac2f079b8c1 Scope: diff against base tree `bc96a7769926` Status: dispatched — coverage complete (5/5 slots terminal) Facts: current review-wide projection Computed under: ```json { "abandonment": "abandonment-v1", "anchor_recipe": 1, "batch_policy": "batch-v1", "coverage": "coverage-v3", "dispatch_policy": "dispatch-v2", "grounder_version": 1, "grounding_read_rule": "grounding-read-v1", "promotion_policy": "promotion-v1", "report": "report-v4", "tally": "tally-v1", "triage_settle": "triage-settle-v2" } ``` ## Findings (2) ### medium — The catalog description copies the authoritative finding-category set - claim: `01M48322K8DAMB4RX5JHJEERTR` - anchor: `skills/verify-doc-drift/SKILL.md` (snippet) ``` Each finding is categorized (incorrect / code-drift / obvious / duplicate) and adversarially verified against the code before it is reported, then docs are fixed to match the code (the reverse only when the doc is the intended source of truth). ``` - lens: writing-quality · arm: default - verdicts: 1 valid / 0 invalid / 0 uncertain - pass `01M4833PDTKGW4ZP1SW16CFJF5` · valid: The exact-grounded description copies the category members. The current body names the source path and its Finding categories instruction, lists the defining entries, and compares them with the copy: both contain incorrect, code-drift, obvious, and duplicate. That specific reviewer-reported comparison meets the restated-sets standard; the router has no need to use these members and can load the defining body, so no exception applies. Deleting the anchored method-summary sentence preserves routing text and leaves the reported verification and fix-direction requirements in their operative sections, consistent with skill-packaging. Reassessing the historical verdicts, their supported copying rationale still applies, but the earlier high-severity label does not: an agreeing copy is medium under the present guideline. The current correction also avoids the earlier full-description replacement concern about losing duplicate/obsolete-doc routing scope. - disposition: none > Changing the audit's classification now requires updating both its operative definitions and the catalog description. The second copy can silently retain obsolete categories even though readers of this skill can open the defining section. > > The source is “Finding categories” in this file, whose “Classify every discrepancy as exactly one of” instruction defines incorrect, code-drift, obvious, and duplicate. The frontmatter parenthetical copies incorrect, code-drift, obvious, and duplicate. The lists currently agree; neither contains a member absent from the other. The restated-sets guideline treats an agreeing copy as a defect because it creates a second maintenance obligation. The installed writing standard's “One Idea, One Place” guidance also puts each instruction in one home, and its packaging reference reserves capability descriptions for routing rather than methods. > > Delete the anchored method-summary sentence from the description. Keep the request-based routing text; the “Finding categories” section remains the authoritative definition, and the body's verification and fix-direction sections retain the useful procedural meaning. > > I read the complete skill, its declared project-docs dependency, and the README's catalog-loading explanation, and compared the parenthetical with the category definitions directly. This is a static comparison, not an execution test. The defining classification instruction and the identical parenthetical establish the copy; it would be refuted if the description were generated from that definition or its intended reader could not access the body, neither of which is indicated in the examined files. ### medium — The universal verification gate excludes valid documentary findings - claim: `01M4832MYZS7687PYDQGMW7WT1` - anchor: `skills/verify-doc-drift/SKILL.md` (snippet) ``` Every candidate finding gets a second, independent pass that tries to **refute** it and defaults to REJECTED unless it can point at the contradicting line itself — re-derive the evidence from the code rather than trusting the first pass. This is what separates real drift from a plausible misread; skipping it ships false corrections. Only survivors reach the report. ``` - lens: writing-quality · arm: default - verdicts: 1 valid / 0 invalid / 0 uncertain - pass `01M4833PDTKGW4ZP1SW16CFJF5` · valid: The exact-grounded gate applies to Every candidate finding, defaults to rejection without the contradicting line, and requires evidence re-derived from the code. The reviewer supplies a specific surrounding instruction trace: The per-unit audit loop accepts a canonical document location for a duplicate and the passage itself for an obvious finding, while the other verification sections supply no category-specific exception. An accurate fact duplicated in documentation can satisfy the reported duplicate criterion without any contradicting code line, so the universal gate excludes an otherwise eligible finding. This is a concrete evidence-contract conflict under Specify the Discipline and Use Precise Language, not a demand for an unnecessary exception. The proposed anchored rewrite preserves independent refutation, fresh evidence, default rejection and the reporting boundary while pointing to the existing citation requirement without restating its set. Reported model behavior remains unmeasured, but that does not defeat the instruction conflict; medium severity is appropriate. - disposition: none > A verifier following this gate literally rejects valid documentary findings that have no contradicting code. For example, the same accurate fact duplicated in two documents satisfies the audit's duplicate criterion but cannot supply a source line contradicting that fact. Those findings disappear before reporting or correction. > > “The per-unit audit loop” explicitly accepts the canonical document location as evidence for a duplicate and the document passage itself for an obvious finding. In contrast, the anchored rule applies to “Every candidate finding” and defaults to rejection unless the verifier can point at “the contradicting line itself”, then requires evidence re-derived “from the code”. The documentary evidence allowed by the earlier contract cannot satisfy that universal gate. The installed writing standard's “Specify the Discipline” and “Use Precise Language” guidance requires checkable evidence appropriate to the task and conditions scoped to the actions they govern. > > Replace the gate with: “Every candidate finding gets a second, independent pass that tries to refute it and defaults to REJECTED unless fresh evidence satisfies the citation requirement in ‘The per-unit audit loop’. Re-derive that evidence rather than trusting the first pass. Only survivors reach the report.” This retains independent verification, rejection of unsupported findings, and the reporting boundary while referring to the existing evidence contract rather than copying its category set. > > I traced “Finding categories”, “The per-unit audit loop”, “Adversarial verification”, the Task's verification step, and Output in the complete skill. None scopes the contradicting-code requirement to factual contradictions or exempts documentary findings. This is a static instruction conflict; I did not run an audit with a served model and cannot establish how often a model would resolve the conflict correctly on its own. A category-scoped gate or an explicit category-appropriate evidence requirement would refute the textual finding; successful documentary verification in a model trial would qualify the practical impact without removing the conflicting instructions. ## Other claims - grounding-pending (0) - ungrounded (0) - rejected (0) - duplicate-of (0) - unadjudicated (0) ## Coverage Coverage pass: 01M482S51F18NVSNJGVWX5Z2PB Accounting: complete Slot health: healthy | lens | part | arm | unit status | runs | loss | | --- | --- | --- | --- | --- | --- | | general-bug | whole | default | no-claims | 1 | no | | writing-quality | whole | default | claims-emitted | 1 | no | | test-trimming | whole | default | no-claims | 1 | no | | restated-sets | whole | default | no-claims | 1 | no | | project-docs | whole | default | no-claims | 1 | no |
Author
Owner

Replying to review comment #111947

Correction: this report has two findings; delivered coverage and zero native review conversations did not make it a clean verdict. The summary is reopened as in-progress while this work remains incomplete.

Report-only claim 01M422TTQMGMVTX6J7V0TVDR26 (low): removing the blanket opening “Documentation rots the moment code changes” sentence is staged as a scoped contradiction repair. The existing release-specific dated-record rule remains. This is not committed or fixed on the PR; normal signing is blocked pending the user's 1Password unlock.

Report-only claim 01M422S5Z0ZDES9YNFC66HDQ7N (high): the frontmatter repeats the finding categories defined below in the same file. This is pre-existing packaging cleanup beyond #110's bounded dated-caveat relocation and the user's minimal composition/no-new-policy scope. It is scope-deferred with the source evidence retained, not claimed fixed or disproved. No additional PR or category-policy change is adopted.

Both claims occur in the stamped summary only: the stable canonical sweep returned zero reviews/conversations and no corresponding inline anchors. No native disposition transition is claimed.

> Replying to review comment #111947 Correction: this report has two findings; delivered coverage and zero native review conversations did not make it a clean verdict. The summary is reopened as in-progress while this work remains incomplete. Report-only claim `01M422TTQMGMVTX6J7V0TVDR26` (low): removing the blanket opening “Documentation rots the moment code changes” sentence is staged as a scoped contradiction repair. The existing release-specific dated-record rule remains. This is not committed or fixed on the PR; normal signing is blocked pending the user's 1Password unlock. Report-only claim `01M422S5Z0ZDES9YNFC66HDQ7N` (high): the frontmatter repeats the finding categories defined below in the same file. This is pre-existing packaging cleanup beyond #110's bounded dated-caveat relocation and the user's minimal composition/no-new-policy scope. It is scope-deferred with the source evidence retained, not claimed fixed or disproved. No additional PR or category-policy change is adopted. Both claims occur in the stamped summary only: the stable canonical sweep returned zero reviews/conversations and no corresponding inline anchors. No native disposition transition is claimed.
fix: preserve dated documentation authority in audit opening
All checks were successful
commit-msg / commitlint (pull_request) Successful in 27s
Review / Review (pull_request_target) Successful in 1m21s
Node tests / node:test (pull_request) Successful in 4m24s
686096787f
Author
Owner

Replying to review #01M422NEF7X7VN2QYE9QVSAKWR

Summary #111947, low claim 01M422TTQMGMVTX6J7V0TVDR26: fixed in 686096787f7b14fbcd32824bb4728bb1fe6fc9c9. Removed the blanket opening sentence; dated-release comparison and ADR authority remain.

High claim 01M422S5Z0ZDES9YNFC66HDQ7N remains verified and scope-deferred: frontmatter repeats the finding categories. This existing packaging cleanup is outside the authorized dated-record repair and minimal composition. It is not fixed or disproved, and no broader policy was adopted.

> Replying to review #01M422NEF7X7VN2QYE9QVSAKWR Summary #111947, low claim `01M422TTQMGMVTX6J7V0TVDR26`: fixed in `686096787f7b14fbcd32824bb4728bb1fe6fc9c9`. Removed the blanket opening sentence; dated-release comparison and ADR authority remain. High claim `01M422S5Z0ZDES9YNFC66HDQ7N` remains verified and scope-deferred: frontmatter repeats the finding categories. This existing packaging cleanup is outside the authorized dated-record repair and minimal composition. It is not fixed or disproved, and no broader policy was adopted.
jercik changed target branch from fix/skills-point-to-sources to main 2026-10-06 07:45:16 +00:00
jercik force-pushed fix/dates-audit-release-evidence from 686096787f
All checks were successful
commit-msg / commitlint (pull_request) Successful in 27s
Review / Review (pull_request_target) Successful in 1m21s
Node tests / node:test (pull_request) Successful in 4m24s
to 8ee8e3675a
All checks were successful
commit-msg / commitlint (pull_request) Successful in 22s
Node tests / node:test (pull_request) Successful in 3m7s
Review / Review (pull_request_target) Successful in 4s
2026-10-06 07:45:16 +00:00
Compare
Author
Owner

Rebased onto main after #97 was squash-merged and retargeted this PR to main. The diff is now just the dated-record caveat and the opening-sentence removal, head 8ee8e36. The frontmatter-duplication finding stays deferred as pre-existing cleanup outside this PR. Merge after the re-triggered review and checks pass.

Rebased onto main after #97 was squash-merged and retargeted this PR to main. The diff is now just the dated-record caveat and the opening-sentence removal, head `8ee8e36`. The frontmatter-duplication finding stays deferred as pre-existing cleanup outside this PR. Merge after the re-triggered review and checks pass.
Author
Owner

Replying to review comment #111947

Summary findings on head 8ee8e36 (review 01M482S4AGS61A02RY67T8Y9Z2):

  • Frontmatter copies the finding categories: verified, but the text is not touched by this PR. Tracked in #119, which removes the list from the description.
  • Verification gate demands a contradicting code line for every finding: verified against "The per-unit audit loop", which cites documentation for duplicate and obvious. The gate is not touched by this PR either. Tracked in #122.

Neither finding makes this PR harmful to merge, so no change is pushed here. The earlier outcome comments on this PR for the opening-sentence claim and the frontmatter finding still stand.

> Replying to review comment #111947 Summary findings on head `8ee8e36` (review `01M482S4AGS61A02RY67T8Y9Z2`): - Frontmatter copies the finding categories: verified, but the text is not touched by this PR. Tracked in [#119](https://code.j4k.dev/j4k-oss/agent-skills/pulls/119), which removes the list from the description. - Verification gate demands a contradicting code line for every finding: verified against "The per-unit audit loop", which cites documentation for `duplicate` and `obvious`. The gate is not touched by this PR either. Tracked in [#122](https://code.j4k.dev/j4k-oss/agent-skills/pulls/122). Neither finding makes this PR harmful to merge, so no change is pushed here. The earlier outcome comments on this PR for the opening-sentence claim and the frontmatter finding still stand.
jercik merged commit 9fc4b067ff into main 2026-10-07 22:00:47 +00:00
jercik deleted branch fix/dates-audit-release-evidence 2026-10-07 22:00:47 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
j4k-oss/agent-skills!110
No description provided.