feat(rules): share the Claude and Codex sub-agent model rules #91

Merged
jercik merged 9 commits from feat/share-subagent-model-rules into main 2026-10-01 09:07:11 +00:00
Owner

Moves the Claude and Codex sub-agent model rules here from j4k/setup-atlas and reduces each to two tiers:

Work Claude Codex
Regular: exploration, research, commands, CI watching, browsing, implementing scoped changes sonnet at low gpt-6.1-sol at low
Original thought: plans, design, architecture, reviews, opinions, diagnoses (ceiling) opus at xhigh gpt-6-astra at high

Rule IDs are unchanged, so existing selections need no edit.

Skills now leave model and effort to these rules. Three sites in improve-codebase-architecture and review-agent-instructions say only what the sub-agent delivers and drop Claude Code-specific tool names (#92).

j4k/setup-atlas#134 already deleted the private copies. Until this merges, a refreshed setup-atlas checkout selects rules that no source provides, and axskills refuses to run.

Moves the Claude and Codex sub-agent model rules here from `j4k/setup-atlas` and reduces each to two tiers: | Work | Claude | Codex | | --- | --- | --- | | Regular: exploration, research, commands, CI watching, browsing, implementing scoped changes | `sonnet` at `low` | `gpt-6.1-sol` at `low` | | Original thought: plans, design, architecture, reviews, opinions, diagnoses (ceiling) | `opus` at `xhigh` | `gpt-6-astra` at `high` | Rule IDs are unchanged, so existing selections need no edit. Skills now leave model and effort to these rules. Three sites in `improve-codebase-architecture` and `review-agent-instructions` say only what the sub-agent delivers and drop Claude Code-specific tool names (#92). j4k/setup-atlas#134 already deleted the private copies. Until this merges, a refreshed setup-atlas checkout selects rules that no source provides, and `axskills` refuses to run.
feat(rules): share the Claude and Codex sub-agent model rules
All checks were successful
commit-msg / commitlint (pull_request) Successful in 16s
Node tests / node:test (pull_request) Successful in 1m14s
Review / Review (pull_request_target) Successful in 11m11s
57dc5449df

Review 01M3VA1JHFYRT7H3A16SX436B9 — head 17e17788f269a2a307d94a1f33e7636bc17d7b7e

Review — j4k-oss/agent-skills @ 9bc616862d

Scope: diff against base tree 87531f305d3a
Status: dispatched — coverage complete (3/3 slots terminal)
Facts: current review-wide projection

Computed under:

{
  "abandonment": "abandonment-v1",
  "anchor_recipe": 1,
  "batch_policy": "batch-v1",
  "coverage": "coverage-v3",
  "dispatch_policy": "dispatch-v2",
  "grounder_version": 1,
  "grounding_read_rule": "grounding-read-v1",
  "promotion_policy": "promotion-v1",
  "report": "report-v3",
  "tally": "tally-v1",
  "triage_settle": "triage-settle-v2"
}

Findings (3)

medium — review-agent-instructions replaces explicit strongest-model guidance with an unexplained "delivers a conclusion" aside that only works with an agent-specific rule

  • claim: 01M3VA560TT6X25SRN4RYZ1K93
  • anchor: skills/review-agent-instructions/SKILL.md (snippet)
For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test.
  • lens: writing-quality · arm: default
  • verdicts: 1 valid / 0 invalid / 0 uncertain
  • disposition: none

What I examined: the changed sentence in skills/review-agent-instructions/SKILL.md ("## Judge independently"), the new rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, and the README.

What changed: the diff removed "use a fresh reviewer with the strongest available model at its deepest reasoning setting" and put in "use a fresh reviewer, which delivers a conclusion and need not be the model under test". The words "which delivers a conclusion" carry no instruction of their own. They only matter to a reader who has also loaded the new tier rule ("Use the higher tier only when the deliverable is a judgment the caller will rely on"). Those rules exist only under rules/claude/ and rules/codex/. The README says they are "selected by an agent override", and it lists Claude, Codex, Cursor, Gemini, OpenCode, and Grok as delivery targets.

What goes wrong: (1) Portability. The writing guide says "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have." An agent that runs this skill without the override rule (any of the other four harnesses, or Claude/Codex without the override selected) no longer gets any model guidance for the judge. The judge is the evidence-bearing step of this skill, so it may run at a default or cheap tier. (2) Clarity. Read alone, the relative clause looks like an odd description of the reviewer, not an instruction. That fails the no-op test for those readers. (3) Terminology. The rule keys on "judgment" and this skill says "conclusion", so a literal reader has to infer that the two words mean the same thing (see the companion finding on the rule).

Proposed correction: say the decision in portable terms that don't name a model, for example "For a clean-context review, use a fresh reviewer on your strongest available model tier, since its verdict is a judgment the caller relies on; it need not be the model under test." This keeps the change's goal of keeping model names out of the skill. It restores the instruction for readers without the rule and lines up with the rule's "judgment" trigger for readers who have it.

Proof gap: I couldn't check which rules a given installation selects. The claim rests on the README's statement that these rules are selected per agent and on the rule directories covering only two of the six listed harnesses.

low — Claude sub-agent rule says the sonnet alias is "Sonnet 5.5", a model that does not exist

  • claim: 01M3VA3YDKATBHNCMCJV6WH9X3
  • anchor: rules/claude/subagent-model-selection.md (snippet)
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work:
  • lens: general-bug · arm: default
  • verdicts: 1 valid / 0 invalid / 0 uncertain
  • duplicates: 01M3VA5H888VW2937FP658QPEE (writing-quality)
  • disposition: none

What I examined: the new rule rules/claude/subagent-model-selection.md, which is delivered to Claude Code sessions as user-level guidance for choosing sub-agent models. Its regular tier reads "sonnet (Sonnet 5.5) at low", and the higher tier reads "opus (Opus 5.5) at xhigh".

What goes wrong: the current Claude 5 family is Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), and Sonnet 5 (claude-sonnet-5), plus Haiku 4.5. No Sonnet 5.5 exists. The Opus label is correct, but the sonnet alias resolves to Sonnet 5, not 5.5. An agent that follows the rule and pins a full model ID instead of the alias (for example in agent-definition frontmatter, where the rule says to set the model) could derive claude-sonnet-5-5 from this label, and that ID would not resolve. A reader also gets the wrong idea of which model the regular tier uses. The cost is small because the rule's main instruction is to use the sonnet alias, which works.

Evidence basis: I checked the label against the current Claude model list, which I know from my own environment. I did not run anything against the Claude API from this sandbox, which has no network access. To confirm or refute the claim, check whether claude-sonnet-5-5 appears in Anthropic's model list. If it does not, change the parenthetical to "Sonnet 5" or remove it.

low — Tier rules key on a "judgment" deliverable while the skills that feed them signal "conclusion"/"facts", splitting one concept across two terms

  • claim: 01M3VA56GWMZDAV9MR15BBV4Z0
  • anchor: rules/claude/subagent-model-selection.md (snippet)
Use the higher tier only when the deliverable is a judgment the caller will rely on.
  • lens: writing-quality · arm: default
  • verdicts: 1 valid / 0 invalid / 0 uncertain
  • disposition: none

What I examined: the trigger sentence shared by rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, and the three skill passages this change rewrote to feed it:

  • improve-codebase-architecture/SKILL.md: "Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions."
  • improve-codebase-architecture/INTERFACE-DESIGN.md: "Spawn 3+ sub-agents in parallel. Each delivers a design conclusion, not a fact list, ..."
  • review-agent-instructions/SKILL.md: "use a fresh reviewer, which delivers a conclusion ..."

What goes wrong: the rules decide the tier on "the deliverable is a judgment the caller will rely on". The skills describe deliverables as "conclusions" or "facts". The same skill also uses "judgment" for something else: "The judgment calls stay with you" means the main agent's own work, not a sub-agent deliverable. The writing guide says "use one term for one concept". A literal reader has to infer that "design conclusion" means "judgment" before the rule applies. In the architecture skill, the nearby "judgment calls stay with you" points away from that reading.

Proposed correction: pick one term and use it in both places. One option is to change the rules to "Use the higher tier only when the deliverable is a conclusion the caller will rely on — a design, review, diagnosis, or choice — not gathered facts." Then the skills' existing "conclusion"/"facts" wording maps onto the rule directly. The other option is to keep "judgment" and change the skills to say the sub-agent "delivers a judgment". Either way the tier split and its examples stay the same, and the rule-to-skill link becomes explicit.

Other claims

  • grounding-pending (0)
  • ungrounded (0)
  • rejected (0)
  • duplicate-of (1)
    • 01M3VA5H888VW2937FP658QPEE low — Claude tier rule pins the sonnet/opus aliases to version numbers that will go stale, and "Sonnet 5.5" doesn't match the current model lineup → 01M3VA3YDKATBHNCMCJV6WH9X3
  • unadjudicated (1)
    • 01M3VA4KZZ0AJDT0SQEP2FWQPV medium — Claude model rule says every sub-agent must use one of two settings but has no fallback when no agent type sets the needed effort

Coverage

Coverage pass: 01M3VA1JJXFAM3C57BK4Z5PFEF
Accounting: complete
Slot health: healthy

lens part arm unit status runs loss
general-bug whole default claims-emitted 1 no
writing-quality whole default claims-emitted 1 no
test-trimming whole default no-claims 1 no
<!-- review:summary --> **Review** `01M3VA1JHFYRT7H3A16SX436B9` — head `17e17788f269a2a307d94a1f33e7636bc17d7b7e` # Review — j4k-oss/agent-skills @ 9bc616862d3a Scope: diff against base tree `87531f305d3a` Status: dispatched — coverage complete (3/3 slots terminal) Facts: current review-wide projection Computed under: ```json { "abandonment": "abandonment-v1", "anchor_recipe": 1, "batch_policy": "batch-v1", "coverage": "coverage-v3", "dispatch_policy": "dispatch-v2", "grounder_version": 1, "grounding_read_rule": "grounding-read-v1", "promotion_policy": "promotion-v1", "report": "report-v3", "tally": "tally-v1", "triage_settle": "triage-settle-v2" } ``` ## Findings (3) ### medium — review-agent-instructions replaces explicit strongest-model guidance with an unexplained "delivers a conclusion" aside that only works with an agent-specific rule - claim: `01M3VA560TT6X25SRN4RYZ1K93` - anchor: `skills/review-agent-instructions/SKILL.md` (snippet) ``` For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. ``` - lens: writing-quality · arm: default - verdicts: 1 valid / 0 invalid / 0 uncertain - disposition: none > What I examined: the changed sentence in `skills/review-agent-instructions/SKILL.md` ("## Judge independently"), the new `rules/claude/subagent-model-selection.md` and `rules/codex/subagent-model-selection.md`, and the README. > > What changed: the diff removed "use a fresh reviewer with the strongest available model at its deepest reasoning setting" and put in "use a fresh reviewer, which delivers a conclusion and need not be the model under test". The words "which delivers a conclusion" carry no instruction of their own. They only matter to a reader who has also loaded the new tier rule ("Use the higher tier only when the deliverable is a judgment the caller will rely on"). Those rules exist only under `rules/claude/` and `rules/codex/`. The README says they are "selected by an agent override", and it lists Claude, Codex, Cursor, Gemini, OpenCode, and Grok as delivery targets. > > What goes wrong: (1) Portability. The writing guide says "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have." An agent that runs this skill without the override rule (any of the other four harnesses, or Claude/Codex without the override selected) no longer gets any model guidance for the judge. The judge is the evidence-bearing step of this skill, so it may run at a default or cheap tier. (2) Clarity. Read alone, the relative clause looks like an odd description of the reviewer, not an instruction. That fails the no-op test for those readers. (3) Terminology. The rule keys on "judgment" and this skill says "conclusion", so a literal reader has to infer that the two words mean the same thing (see the companion finding on the rule). > > Proposed correction: say the decision in portable terms that don't name a model, for example "For a clean-context review, use a fresh reviewer on your strongest available model tier, since its verdict is a judgment the caller relies on; it need not be the model under test." This keeps the change's goal of keeping model names out of the skill. It restores the instruction for readers without the rule and lines up with the rule's "judgment" trigger for readers who have it. > > Proof gap: I couldn't check which rules a given installation selects. The claim rests on the README's statement that these rules are selected per agent and on the rule directories covering only two of the six listed harnesses. ### low — Claude sub-agent rule says the `sonnet` alias is "Sonnet 5.5", a model that does not exist - claim: `01M3VA3YDKATBHNCMCJV6WH9X3` - anchor: `rules/claude/subagent-model-selection.md` (snippet) ``` - `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: ``` - lens: general-bug · arm: default - verdicts: 1 valid / 0 invalid / 0 uncertain - duplicates: `01M3VA5H888VW2937FP658QPEE` (writing-quality) - disposition: none > What I examined: the new rule `rules/claude/subagent-model-selection.md`, which is delivered to Claude Code sessions as user-level guidance for choosing sub-agent models. Its regular tier reads "`sonnet` (Sonnet 5.5) at `low`", and the higher tier reads "`opus` (Opus 5.5) at `xhigh`". > > What goes wrong: the current Claude 5 family is Fable 5.1 (`claude-fable-5-1`), Opus 5.5 (`claude-opus-5-5`), and Sonnet 5 (`claude-sonnet-5`), plus Haiku 4.5. No Sonnet 5.5 exists. The Opus label is correct, but the `sonnet` alias resolves to Sonnet 5, not 5.5. An agent that follows the rule and pins a full model ID instead of the alias (for example in agent-definition frontmatter, where the rule says to set the model) could derive `claude-sonnet-5-5` from this label, and that ID would not resolve. A reader also gets the wrong idea of which model the regular tier uses. The cost is small because the rule's main instruction is to use the `sonnet` alias, which works. > > Evidence basis: I checked the label against the current Claude model list, which I know from my own environment. I did not run anything against the Claude API from this sandbox, which has no network access. To confirm or refute the claim, check whether `claude-sonnet-5-5` appears in Anthropic's model list. If it does not, change the parenthetical to "Sonnet 5" or remove it. ### low — Tier rules key on a "judgment" deliverable while the skills that feed them signal "conclusion"/"facts", splitting one concept across two terms - claim: `01M3VA56GWMZDAV9MR15BBV4Z0` - anchor: `rules/claude/subagent-model-selection.md` (snippet) ``` Use the higher tier only when the deliverable is a judgment the caller will rely on. ``` - lens: writing-quality · arm: default - verdicts: 1 valid / 0 invalid / 0 uncertain - disposition: none > What I examined: the trigger sentence shared by `rules/claude/subagent-model-selection.md` and `rules/codex/subagent-model-selection.md`, and the three skill passages this change rewrote to feed it: > - improve-codebase-architecture/SKILL.md: "Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions." > - improve-codebase-architecture/INTERFACE-DESIGN.md: "Spawn 3+ sub-agents in parallel. Each delivers a design conclusion, not a fact list, ..." > - review-agent-instructions/SKILL.md: "use a fresh reviewer, which delivers a conclusion ..." > > What goes wrong: the rules decide the tier on "the deliverable is a judgment the caller will rely on". The skills describe deliverables as "conclusions" or "facts". The same skill also uses "judgment" for something else: "The judgment calls stay with you" means the main agent's own work, not a sub-agent deliverable. The writing guide says "use one term for one concept". A literal reader has to infer that "design conclusion" means "judgment" before the rule applies. In the architecture skill, the nearby "judgment calls stay with you" points away from that reading. > > Proposed correction: pick one term and use it in both places. One option is to change the rules to "Use the higher tier only when the deliverable is a conclusion the caller will rely on — a design, review, diagnosis, or choice — not gathered facts." Then the skills' existing "conclusion"/"facts" wording maps onto the rule directly. The other option is to keep "judgment" and change the skills to say the sub-agent "delivers a judgment". Either way the tier split and its examples stay the same, and the rule-to-skill link becomes explicit. ## Other claims - grounding-pending (0) - ungrounded (0) - rejected (0) - duplicate-of (1) - `01M3VA5H888VW2937FP658QPEE` low — Claude tier rule pins the `sonnet`/`opus` aliases to version numbers that will go stale, and "Sonnet 5.5" doesn't match the current model lineup → `01M3VA3YDKATBHNCMCJV6WH9X3` - unadjudicated (1) - `01M3VA4KZZ0AJDT0SQEP2FWQPV` medium — Claude model rule says every sub-agent must use one of two settings but has no fallback when no agent type sets the needed effort ## Coverage Coverage pass: 01M3VA1JJXFAM3C57BK4Z5PFEF Accounting: complete Slot health: healthy | lens | part | arm | unit status | runs | loss | | --- | --- | --- | --- | --- | --- | | general-bug | whole | default | claims-emitted | 1 | no | | writing-quality | whole | default | claims-emitted | 1 | no | | test-trimming | whole | default | no-claims | 1 | no |
@ -0,0 +4,4 @@
- `gpt-6-luna` at `high` for mechanical work with a clear output and little judgment: running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, applying a specified bulk edit. The small model needs `high` even for mechanical work.
- `gpt-6-sol` at `medium` for work that takes trial and error to finish: implementing a scoped change, getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.
- `gpt-6-astra` at `high` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.

medium — The Astra effort rule conflicts with the interface-design skill maximum-effort instruction
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

I read the new Codex model rule and the full skills/improve-codebase-architecture/INTERFACE-DESIGN.md, which the architecture skill invokes when the user wants alternative interfaces. Its designer step directs three or more sub-agents to use the strongest reasoning model at maximum effort because they deliver design conclusions. The new Codex rule assigns exactly that kind of conclusion to gpt-6-astra at high, then directs the agent to move up a model tier rather than raise effort. Astra is already the highest tier named here, so a Codex agent using the skill receives incompatible effort instructions for the same designers. State whether an explicit skill instruction takes precedence over this default, or revise the interface-design instruction to the chosen effort. That preserves the intended model guidance and lets the agent select a single effort for the design task. This is a textual conflict; I did not run the skill or a model comparison.

claim 01M3TYZ7M0FFRN5CGE1P6BX2N7 of review 01M3TYMYHQ04WGZ7WBT0Y2MJHN

<!-- review:claim:01M3TYZ7M0FFRN5CGE1P6BX2N7 --> **medium** — The Astra effort rule conflicts with the interface-design skill maximum-effort instruction lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > I read the new Codex model rule and the full skills/improve-codebase-architecture/INTERFACE-DESIGN.md, which the architecture skill invokes when the user wants alternative interfaces. Its designer step directs three or more sub-agents to use the strongest reasoning model at maximum effort because they deliver design conclusions. The new Codex rule assigns exactly that kind of conclusion to gpt-6-astra at high, then directs the agent to move up a model tier rather than raise effort. Astra is already the highest tier named here, so a Codex agent using the skill receives incompatible effort instructions for the same designers. State whether an explicit skill instruction takes precedence over this default, or revise the interface-design instruction to the chosen effort. That preserves the intended model guidance and lets the agent select a single effort for the design task. This is a textual conflict; I did not run the skill or a model comparison. claim `01M3TYZ7M0FFRN5CGE1P6BX2N7` of review `01M3TYMYHQ04WGZ7WBT0Y2MJHN`
Author
Owner

Fixed in 9ac36d692a. INTERFACE-DESIGN.md now says the designers deliver a conclusion and defers model and effort to the agent's own instructions, falling back to the strongest model at maximum effort. The same fix covers unadjudicated claim 01M3TYWHPH0PRZFA07N67M0VSR (the Explore walk in improve-codebase-architecture/SKILL.md) and the identical pattern in review-agent-instructions/SKILL.md.

<!-- gh-feedback:reply-to:97989 --> Fixed in 9ac36d692a528d829bcd28cb43e8f82bc62f7e88. INTERFACE-DESIGN.md now says the designers deliver a conclusion and defers model and effort to the agent's own instructions, falling back to the strongest model at maximum effort. The same fix covers unadjudicated claim 01M3TYWHPH0PRZFA07N67M0VSR (the Explore walk in improve-codebase-architecture/SKILL.md) and the identical pattern in review-agent-instructions/SKILL.md.
jercik marked this conversation as resolved
fix(skills): defer sub-agent model choice to the agent's instructions
All checks were successful
commit-msg / commitlint (pull_request) Successful in 26s
Node tests / node:test (pull_request) Successful in 1m35s
Review / Review (pull_request_target) Successful in 5m8s
9ac36d692a
@ -0,0 +6,4 @@
- `gpt-6-sol` at `medium` for work that takes trial and error to finish: implementing a scoped change, getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.
- `gpt-6-astra` at `high` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.
Route a task that sits between two tiers to the higher one, and move a task that needs more reasoning up a tier rather than raising its effort. To set `model` and `reasoning_effort`, use `fork_turns: "none"` or a positive integer string and pass the child the task context it needs. A full-history fork inherits the parent model and effort and rejects overrides; use it when preserving that context matters more than selecting a different tier.

low — Model escalation rule has no action above the highest tier
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

I examined all three model tiers in this new Codex rule. Its conclusion tier is gpt-6-astra at high, the highest tier listed. The final directive says to move any task needing more reasoning up a tier rather than raise effort. If a complex review or diagnosis already starts at that conclusion tier, there is no higher tier, so the agent has no stated way to respond when the assigned setting is insufficient. It must guess whether to keep working at high, raise effort, divide the task, or report a limit. This leaves the rule’s hardest case unresolved.

The writing standard calls for operational choices and explicit handling when a prerequisite cannot be met. Scope the “move up a tier” instruction to the first two tiers and state the intended top-tier fallback, such as a higher effort if available or splitting the task and reporting the limit. This preserves the preferred model escalation for tasks below the top tier. This is a logical gap from static inspection; an explicit top-tier policy elsewhere in the selected instruction set would refute it.

claim 01M3TZSY4HTZWX7TN0JYCTM7CS of review 01M3TZKC5QYK787RHC0KAVA8XF

<!-- review:claim:01M3TZSY4HTZWX7TN0JYCTM7CS --> **low** — Model escalation rule has no action above the highest tier lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > I examined all three model tiers in this new Codex rule. Its conclusion tier is gpt-6-astra at high, the highest tier listed. The final directive says to move any task needing more reasoning up a tier rather than raise effort. If a complex review or diagnosis already starts at that conclusion tier, there is no higher tier, so the agent has no stated way to respond when the assigned setting is insufficient. It must guess whether to keep working at high, raise effort, divide the task, or report a limit. This leaves the rule’s hardest case unresolved. > > The writing standard calls for operational choices and explicit handling when a prerequisite cannot be met. Scope the “move up a tier” instruction to the first two tiers and state the intended top-tier fallback, such as a higher effort if available or splitting the task and reporting the limit. This preserves the preferred model escalation for tasks below the top tier. This is a logical gap from static inspection; an explicit top-tier policy elsewhere in the selected instruction set would refute it. claim `01M3TZSY4HTZWX7TN0JYCTM7CS` of review `01M3TZKC5QYK787RHC0KAVA8XF`
Author
Owner

The gap is real: nothing says what to do when a task already on gpt-6-astra needs more reasoning. Which way it goes is the rule owner's policy call, so I've asked @jercik: keep astra at high as the ceiling (split the task or report the limit), or allow xhigh at the top tier only?

<!-- gh-feedback:reply-to:98015 --> The gap is real: nothing says what to do when a task already on gpt-6-astra needs more reasoning. Which way it goes is the rule owner's policy call, so I've asked @jercik: keep astra at high as the ceiling (split the task or report the limit), or allow xhigh at the top tier only?

superseded by review 01M3V3MV82K6PQQ2G16FGY1VBT for head 1cd99e7d9154990d729c5486a52ac97f460aeb98

<!-- review:superseded:01M3V3MV82K6PQQ2G16FGY1VBT --> superseded by review `01M3V3MV82K6PQQ2G16FGY1VBT` for head `1cd99e7d9154990d729c5486a52ac97f460aeb98`
Author
Owner

Fixed in 86027ddd86. The owner chose gpt-6-astra at high as the top tier. The rule now calls it the ceiling and says to split a task that needs more or report the limit. The 'up a tier rather than raising its effort' sentence is gone.

<!-- gh-feedback:reply-to:98015 --> Fixed in 86027ddd866fe78aa3c8bac18a3002df63831cad. The owner chose gpt-6-astra at high as the top tier. The rule now calls it the ceiling and says to split a task that needs more or report the limit. The 'up a tier rather than raising its effort' sentence is gone.
jercik marked this conversation as resolved
fix(skills): keep sub-agent instructions harness-neutral (#92)
Some checks failed
commit-msg / commitlint (pull_request) Failing after 12s
Node tests / node:test (pull_request) Successful in 1m18s
Review / Review (pull_request_target) Successful in 5m38s
1cd99e7d91
Skills now leave model and effort to the per-harness rules under `rules/claude/` and `rules/codex/`. Three skill sites lose their fallback model and effort wording and the Claude Code-specific "Agent tool" and `subagent_type=Explore`. Each keeps only what the sub-agent delivers (facts or a conclusion), which is what those rules select on. Without a rule, the harness default picks the model.

In 16 test runs, agents given the Claude rule picked `opus` at `xhigh` for conclusions and `medium` for facts. Agents given the Codex rule picked `gpt-6-astra` at `high` and `gpt-6-sol` at `medium`.

Two neighboring lines also changed. `improve-codebase-architecture` now applies the deletion test with the outcomes the skill defines. In `review-agent-instructions`, the bar on editing a retest no longer ends once the retest passes.

Stacked on #91. Retarget to `main` after #91 merges.

Reviewed-on: #92
@ -0,0 +6,4 @@
- `gpt-6-sol` at `medium` for work that takes trial and error to finish: implementing a scoped change, getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.
- `gpt-6-astra` at `high` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.
Route a task that sits between two tiers to the higher one, and move a task that needs more reasoning up a tier rather than raising its effort. To set `model` and `reasoning_effort`, use `fork_turns: "none"` or a positive integer string and pass the child the task context it needs. A full-history fork inherits the parent model and effort and rejects overrides; use it when preserving that context matters more than selecting a different tier.

low — Define an escalation path for tasks already on Astra
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

I read the complete Codex model-selection rule and the changed skills that delegate architecture design and instruction review. The rule names only three tiers, ending with gpt-6-astra at high, then tells the agent to move any task needing more reasoning up a tier rather than raise effort. For a task already assigned to Astra, that action has no destination; a difficult review or design conclusion has no stated next step. This is a static decision-tree gap, not a reproduced agent failure. Narrow the sentence to say to step up when a higher tier exists and state what to do at the top tier, for example raise Astra effort where supported or split the task. That preserves the intended preference for model-tier escalation while making the last branch actionable. The writing standard calls for operational decisions with explicit boundaries; an execution trace of a top-tier task that needs extra reasoning would establish whether current agents infer a safe fallback.

claim 01M3V3SQAC4S12K4B5YKJX091A of review 01M3V3MV82K6PQQ2G16FGY1VBT

<!-- review:claim:01M3V3SQAC4S12K4B5YKJX091A --> **low** — Define an escalation path for tasks already on Astra lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > I read the complete Codex model-selection rule and the changed skills that delegate architecture design and instruction review. The rule names only three tiers, ending with gpt-6-astra at high, then tells the agent to move any task needing more reasoning up a tier rather than raise effort. For a task already assigned to Astra, that action has no destination; a difficult review or design conclusion has no stated next step. This is a static decision-tree gap, not a reproduced agent failure. Narrow the sentence to say to step up when a higher tier exists and state what to do at the top tier, for example raise Astra effort where supported or split the task. That preserves the intended preference for model-tier escalation while making the last branch actionable. The writing standard calls for operational decisions with explicit boundaries; an execution trace of a top-tier task that needs extra reasoning would establish whether current agents infer a safe fallback. claim `01M3V3SQAC4S12K4B5YKJX091A` of review `01M3V3MV82K6PQQ2G16FGY1VBT`

superseded by review 01M3V4BP30R44N7E58JT3RJKRM for head 86027ddd866fe78aa3c8bac18a3002df63831cad

<!-- review:superseded:01M3V4BP30R44N7E58JT3RJKRM --> superseded by review `01M3V4BP30R44N7E58JT3RJKRM` for head `86027ddd866fe78aa3c8bac18a3002df63831cad`
Author
Owner

Fixed in 86027ddd86: gpt-6-astra at high is now stated as the ceiling, with split-or-report beyond it.

<!-- gh-feedback:reply-to:98163 --> Fixed in 86027ddd866fe78aa3c8bac18a3002df63831cad: gpt-6-astra at high is now stated as the ceiling, with split-or-report beyond it.
jercik marked this conversation as resolved
@ -39,3 +39,3 @@
Follow its loading procedure: a root `CONTEXT-MAP.md` means the repo has multiple contexts — read the touched contexts' `CONTEXT.md` files, the root `docs/adr/` (system-wide decisions), and each touched context's own `docs/adr/` — otherwise the root `CONTEXT.md` and the ADRs in the area you're touching.
Then use the Agent tool with `subagent_type=Explore` — a fast executor model at moderate reasoning effort, since the walk delivers facts, not conclusions — to walk the codebase. The sub-agent starts with none of your context, so the brief must carry everything the walk needs: the friction questions below, the glossary terms the project docs define, and the [LANGUAGE.md](LANGUAGE.md) definitions of **shallow**, **seam**, and **locality**, with the instruction to report findings as `file:line` evidence in that vocabulary. Don't impose rigid heuristics — have it explore organically and note friction:
Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions. The sub-agent starts with none of your context, so the brief must carry everything the walk needs: the friction questions below, the glossary terms the project docs define, and the [LANGUAGE.md](LANGUAGE.md) definitions of **shallow**, **seam**, and **locality**, with the instruction to report findings as `file:line` evidence in that vocabulary. Don't impose rigid heuristics — have it explore organically and note friction:

medium — Align the exploration brief with its fact-only boundary
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

I read the full architecture skill and its LANGUAGE.md definitions. This new sentence makes the exploration sub-agent fact-only, and the later paragraph says the caller will decide what is shallow. Yet the same brief tells the explorer to answer questions such as where modules are shallow, where pure functions were extracted just for testability, and where coupled modules leak across seams. Those are precisely the architectural judgments that the skill reserves for the caller. A literal explorer can either violate the fact-only boundary or withhold the requested answers, degrading the candidate report. Ask it for observable evidence instead, such as interface surface, forwarding code, call sites, repeated logic, and test entry points, then have the caller apply the defined terms; alternatively permit explicitly labeled hypotheses with supporting evidence. That keeps the useful search targets and makes the division of work executable. This is a static instruction conflict, not an observed run; a trace showing that explorers consistently return only evidence while the caller makes the classifications would refute the practical impact. The writing standard says to put each decision in one place and make task contracts operational.

claim 01M3V3T90DPGJGDRKAK1Q56PMN of review 01M3V3MV82K6PQQ2G16FGY1VBT

<!-- review:claim:01M3V3T90DPGJGDRKAK1Q56PMN --> **medium** — Align the exploration brief with its fact-only boundary lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > I read the full architecture skill and its LANGUAGE.md definitions. This new sentence makes the exploration sub-agent fact-only, and the later paragraph says the caller will decide what is shallow. Yet the same brief tells the explorer to answer questions such as where modules are shallow, where pure functions were extracted just for testability, and where coupled modules leak across seams. Those are precisely the architectural judgments that the skill reserves for the caller. A literal explorer can either violate the fact-only boundary or withhold the requested answers, degrading the candidate report. Ask it for observable evidence instead, such as interface surface, forwarding code, call sites, repeated logic, and test entry points, then have the caller apply the defined terms; alternatively permit explicitly labeled hypotheses with supporting evidence. That keeps the useful search targets and makes the division of work executable. This is a static instruction conflict, not an observed run; a trace showing that explorers consistently return only evidence while the caller makes the classifications would refute the practical impact. The writing standard says to put each decision in one place and make task contracts operational. claim `01M3V3T90DPGJGDRKAK1Q56PMN` of review `01M3V3MV82K6PQQ2G16FGY1VBT`

low — Qualify the claim that every explorer starts without context
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

I read the whole architecture skill and the new Codex sub-agent model-selection rule. This paragraph now says to spawn a generic exploration sub-agent, then assumes it starts with none of the caller context and therefore requires the brief to repeat the glossary terms and LANGUAGE.md definitions. The Codex rule explicitly says a full-history fork inherits parent context, and the architecture skill does not require a fresh-context fork. On that supported path the premise is false, so the agent may duplicate material already loaded into every exploration brief and readers cannot tell which fork mode this procedure intends. Either require a fresh-context spawn here and keep the complete brief, or say that a fresh-context spawn needs those definitions while an inheriting spawn can refer to material already present. This preserves the needed context for isolated explorers and makes the instruction accurate across the advertised spawn modes. The writing standard calls for verifiable environment claims and conditions attached to the actions they govern. This is a static cross-file inconsistency; an actual invocation trace would establish which fork mode the workflow uses in practice.

claim 01M3V3WDCDTZSFW026ET633CXA of review 01M3V3MV82K6PQQ2G16FGY1VBT

<!-- review:claim:01M3V3WDCDTZSFW026ET633CXA --> **low** — Qualify the claim that every explorer starts without context lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > I read the whole architecture skill and the new Codex sub-agent model-selection rule. This paragraph now says to spawn a generic exploration sub-agent, then assumes it starts with none of the caller context and therefore requires the brief to repeat the glossary terms and LANGUAGE.md definitions. The Codex rule explicitly says a full-history fork inherits parent context, and the architecture skill does not require a fresh-context fork. On that supported path the premise is false, so the agent may duplicate material already loaded into every exploration brief and readers cannot tell which fork mode this procedure intends. Either require a fresh-context spawn here and keep the complete brief, or say that a fresh-context spawn needs those definitions while an inheriting spawn can refer to material already present. This preserves the needed context for isolated explorers and makes the instruction accurate across the advertised spawn modes. The writing standard calls for verifiable environment claims and conditions attached to the actions they govern. This is a static cross-file inconsistency; an actual invocation trace would establish which fork mode the workflow uses in practice. claim `01M3V3WDCDTZSFW026ET633CXA` of review `01M3V3MV82K6PQQ2G16FGY1VBT`

superseded by review 01M3V4BP30R44N7E58JT3RJKRM for head 86027ddd866fe78aa3c8bac18a3002df63831cad

<!-- review:superseded:01M3V4BP30R44N7E58JT3RJKRM --> superseded by review `01M3V4BP30R44N7E58JT3RJKRM` for head `86027ddd866fe78aa3c8bac18a3002df63831cad`
jercik marked this conversation as resolved
feat(rules): make Sonnet 5.5 and gpt-6.1-sol the regular sub-agent tier
Some checks failed
commit-msg / commitlint (pull_request) Failing after 12s
Node tests / node:test (pull_request) Successful in 1m20s
Review / Review (pull_request_target) Successful in 4m8s
86027ddd86
@ -0,0 +2,4 @@
When choosing a sub-agent's model explicitly, use one of three settings:
- `gpt-6.1-sol` at `low` is the regular tier, and most sub-agents run on it: mechanical work with a clear output and little judgment, such as running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, or applying a specified bulk edit.

medium — Regular sub-agent tier names a model absent from the available Codex overrides
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

I read the new Codex selection rule, the README selection guidance, and the sub-agent tool contract available to this run. The rule names gpt-6.1-sol at both low and high effort and says most sub-agents use the low tier. The current spawn_agent model override list includes gpt-6-sol and gpt-6-astra but no gpt-6.1-sol. Following the rule with an explicit override therefore sends an unsupported model name and blocks routine delegated work. The writing standard asks for verifiable, version-scoped claims and prefers an authoritative tool lookup over copied environment facts. Name a currently supported model, or direct the agent to select from the tool's available overrides before setting one. An available-model list containing gpt-6.1-sol in the intended deployment would refute the current mismatch; I could verify only the tool contract exposed here, not other installations.

claim 01M3V4FE3C18FAP6RZAN8JBYWF of review 01M3V4BP30R44N7E58JT3RJKRM

<!-- review:claim:01M3V4FE3C18FAP6RZAN8JBYWF --> **medium** — Regular sub-agent tier names a model absent from the available Codex overrides lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > I read the new Codex selection rule, the README selection guidance, and the sub-agent tool contract available to this run. The rule names `gpt-6.1-sol` at both low and high effort and says most sub-agents use the low tier. The current `spawn_agent` model override list includes `gpt-6-sol` and `gpt-6-astra` but no `gpt-6.1-sol`. Following the rule with an explicit override therefore sends an unsupported model name and blocks routine delegated work. The writing standard asks for verifiable, version-scoped claims and prefers an authoritative tool lookup over copied environment facts. Name a currently supported model, or direct the agent to select from the tool's available overrides before setting one. An available-model list containing `gpt-6.1-sol` in the intended deployment would refute the current mismatch; I could verify only the tool contract exposed here, not other installations. claim `01M3V4FE3C18FAP6RZAN8JBYWF` of review `01M3V4BP30R44N7E58JT3RJKRM`
Author
Owner

gpt-6.1-sol is available. Codex 0.159.2's model catalog, fetched 2026-10-01T07:08Z in a local session (models_cache.json), lists gpt-6.1-sol with visibility 'list' and reasoning levels low through ultra, beside gpt-6-astra. The review environment's spawn_agent list predates that model.

<!-- gh-feedback:reply-to:98177 --> gpt-6.1-sol is available. Codex 0.159.2's model catalog, fetched 2026-10-01T07:08Z in a local session (models_cache.json), lists gpt-6.1-sol with visibility 'list' and reasoning levels low through ultra, beside gpt-6-astra. The review environment's spawn_agent list predates that model.

superseded by review 01M3V67FK226HY7EDBZ5F16S4T for head 4bfca904c46127293ec67581ebc0cbb7856cbc10

<!-- review:superseded:01M3V67FK226HY7EDBZ5F16S4T --> superseded by review `01M3V67FK226HY7EDBZ5F16S4T` for head `4bfca904c46127293ec67581ebc0cbb7856cbc10`
feat(rules): run scoped implementation on the regular sub-agent tier
Some checks failed
commit-msg / commitlint (pull_request) Failing after 13s
Node tests / node:test (pull_request) Successful in 1m35s
Review / Review (pull_request_target) Successful in 3m38s
4bfca904c4
@ -0,0 +2,4 @@
Launch every sub-agent at one of three settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier, and most sub-agents run on it: routine work with a clear goal, such as running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, applying a specified bulk edit, or implementing a scoped change.

low — Claude sub-agent rule says the sonnet alias is "Sonnet 5.5", a model version the alias does not resolve to
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: rules/claude/subagent-model-selection.md, the new Claude-specific rule. It labels its tiers "sonnet (Sonnet 5.5)" and "opus (Opus 5.5)".

As far as the reviewer's knowledge of current Claude models goes, the Claude 5 family has Opus 5.5 (claude-opus-5-5) and Sonnet 5 (claude-sonnet-5), with no Sonnet 5.5 release. The sonnet alias passed to Claude Code's Agent tool or a workflow agent() call therefore resolves to Sonnet 5. The parenthetical names a model that doesn't exist.

Why it matters: the rule is delivered verbatim into users' Claude instruction files. An agent, or a human tuning the tiers, who reads "Sonnet 5.5" may try a model ID like claude-sonnet-5-5, which would be rejected. They may also take the regular tier to be a newer model than it is, which skews the trade-off the rule asks them to make between the regular tier and the two opus tiers. The alias itself (sonnet) is correct, so the impact is limited to this mislabel.

This claim rests on the reviewer's model catalog, not on anything in the repository, and I did not check it against a live model list. Anthropic's current models page would confirm or refute it. Fix: change the label to "Sonnet 5", or drop the version parentheticals so the rule doesn't go stale when the aliases move.

claim 01M3V6A8V3D15B0H225V83XGNM of review 01M3V67FK226HY7EDBZ5F16S4T

<!-- review:claim:01M3V6A8V3D15B0H225V83XGNM --> **low** — Claude sub-agent rule says the `sonnet` alias is "Sonnet 5.5", a model version the alias does not resolve to lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: rules/claude/subagent-model-selection.md, the new Claude-specific rule. It labels its tiers "`sonnet` (Sonnet 5.5)" and "`opus` (Opus 5.5)". > > As far as the reviewer's knowledge of current Claude models goes, the Claude 5 family has Opus 5.5 (`claude-opus-5-5`) and Sonnet 5 (`claude-sonnet-5`), with no Sonnet 5.5 release. The `sonnet` alias passed to Claude Code's Agent tool or a workflow `agent()` call therefore resolves to Sonnet 5. The parenthetical names a model that doesn't exist. > > Why it matters: the rule is delivered verbatim into users' Claude instruction files. An agent, or a human tuning the tiers, who reads "Sonnet 5.5" may try a model ID like `claude-sonnet-5-5`, which would be rejected. They may also take the regular tier to be a newer model than it is, which skews the trade-off the rule asks them to make between the regular tier and the two `opus` tiers. The alias itself (`sonnet`) is correct, so the impact is limited to this mislabel. > > This claim rests on the reviewer's model catalog, not on anything in the repository, and I did not check it against a live model list. Anthropic's current models page would confirm or refute it. Fix: change the label to "Sonnet 5", or drop the version parentheticals so the rule doesn't go stale when the aliases move. claim `01M3V6A8V3D15B0H225V83XGNM` of review `01M3V67FK226HY7EDBZ5F16S4T`
Author
Owner

Sonnet 5.5 exists and is what the alias resolves to. The installed Claude Code 2.1.286 binary maps the aliases as sonnet: claude-sonnet-5-5 and opus: claude-opus-5-5, and its 2.1.284 changelog reads 'Added Claude Sonnet 5.5 (claude-sonnet-5-5), now the default Sonnet model on the Anthropic API'.

<!-- gh-feedback:reply-to:98337 --> Sonnet 5.5 exists and is what the alias resolves to. The installed Claude Code 2.1.286 binary maps the aliases as sonnet: claude-sonnet-5-5 and opus: claude-opus-5-5, and its 2.1.284 changelog reads 'Added Claude Sonnet 5.5 (claude-sonnet-5-5), now the default Sonnet model on the Anthropic API'.

superseded by review 01M3V6JPN0S0M7BTX1S93K2BAG for head 83f5156c0a43b2bf85c398c980aa7456c1c33fc1

<!-- review:superseded:01M3V6JPN0S0M7BTX1S93K2BAG --> superseded by review `01M3V6JPN0S0M7BTX1S93K2BAG` for head `83f5156c0a43b2bf85c398c980aa7456c1c33fc1`
@ -0,0 +6,4 @@
- `opus` (Opus 5.5) at `medium` for work that takes trial and error to finish: getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.
- `opus` at `xhigh` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.
Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set the model and effort explicitly wherever they can be set: workflow `agent()` calls and agent definition frontmatter.

medium — Claude model-selection rule requires a setting on "every sub-agent" but omits the Agent tool, where effort cannot be set
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: the new rules/claude/subagent-model-selection.md in full, and the Claude Code Agent tool contract as exposed in this harness.

The rule opens with an absolute: "Launch every sub-agent at one of three settings:" and each setting is a model plus an effort (sonnet at low, opus at medium, opus at xhigh). Its closing sentence then lists where to apply them: "Set the model and effort explicitly wherever they can be set: workflow agent() calls and agent definition frontmatter."

The most common spawn path, a direct Agent tool call, is missing from that list. In Claude Code the Agent tool takes an optional model override (sonnet/opus/haiku/…), but its description says an agent type's "model, reasoning effort, and tools come from its definition (.claude/agents/*.md frontmatter or SDK agents)". A direct call can therefore set the model but not the effort. Built-in types such as Explore or general-purpose keep whatever effort their definitions carry.

The two sentences conflict for a literal reader. "every sub-agent" demands a model-and-effort pair the agent often cannot express. The enumerated list implies that a plain Agent tool call is outside the rule, so the agent may skip even the model override it could set. Either reading defeats the rule's purpose of keeping most sub-agents on the regular tier. The Codex sibling avoids this by scoping itself: "When choosing a sub-agent's model explicitly, use one of three settings". (writing-for-agents: "Make claims verifiable"; give boundaries the reader can act on.)

Proposed correction: change the opener to "Choose each sub-agent's setting from these three:". Replace the last sentence with: "Set both model and effort in workflow agent() calls and agent-definition frontmatter. On a direct Agent tool call, pass the tier's model; the effort comes from the agent type's definition, so use or define an agent type whose effort matches the tier." This keeps the tier table and makes clear what each spawn path can enforce.

Proof gap: I checked the Agent tool schema available in this review sandbox's Claude Code build. Other Claude Code versions may expose an effort parameter on direct calls.

claim 01M3V69XR8W2X9ASA3YPEZK2BA of review 01M3V67FK226HY7EDBZ5F16S4T

<!-- review:claim:01M3V69XR8W2X9ASA3YPEZK2BA --> **medium** — Claude model-selection rule requires a setting on "every sub-agent" but omits the Agent tool, where effort cannot be set lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: the new rules/claude/subagent-model-selection.md in full, and the Claude Code Agent tool contract as exposed in this harness. > > The rule opens with an absolute: "Launch every sub-agent at one of three settings:" and each setting is a model plus an effort (`sonnet` at `low`, `opus` at `medium`, `opus` at `xhigh`). Its closing sentence then lists where to apply them: "Set the model and effort explicitly wherever they can be set: workflow `agent()` calls and agent definition frontmatter." > > The most common spawn path, a direct Agent tool call, is missing from that list. In Claude Code the Agent tool takes an optional `model` override (sonnet/opus/haiku/…), but its description says an agent type's "model, reasoning effort, and tools come from its definition (`.claude/agents/*.md` frontmatter or SDK `agents`)". A direct call can therefore set the model but not the effort. Built-in types such as `Explore` or `general-purpose` keep whatever effort their definitions carry. > > The two sentences conflict for a literal reader. "every sub-agent" demands a model-and-effort pair the agent often cannot express. The enumerated list implies that a plain Agent tool call is outside the rule, so the agent may skip even the `model` override it could set. Either reading defeats the rule's purpose of keeping most sub-agents on the regular tier. The Codex sibling avoids this by scoping itself: "When choosing a sub-agent's model explicitly, use one of three settings". (writing-for-agents: "Make claims verifiable"; give boundaries the reader can act on.) > > Proposed correction: change the opener to "Choose each sub-agent's setting from these three:". Replace the last sentence with: "Set both model and effort in workflow `agent()` calls and agent-definition frontmatter. On a direct Agent tool call, pass the tier's `model`; the effort comes from the agent type's definition, so use or define an agent type whose effort matches the tier." This keeps the tier table and makes clear what each spawn path can enforce. > > Proof gap: I checked the Agent tool schema available in this review sandbox's Claude Code build. Other Claude Code versions may expose an effort parameter on direct calls. claim `01M3V69XR8W2X9ASA3YPEZK2BA` of review `01M3V67FK226HY7EDBZ5F16S4T`
Author
Owner

Fixed in 83f5156c0a. The rule now says workflow agent() calls and agent definition frontmatter set both, and a direct Agent tool call sets only the model, with effort from the agent type's definition.

<!-- gh-feedback:reply-to:98335 --> Fixed in 83f5156c0a43b2bf85c398c980aa7456c1c33fc1. The rule now says workflow agent() calls and agent definition frontmatter set both, and a direct Agent tool call sets only the model, with effort from the agent type's definition.
jercik marked this conversation as resolved
@ -47,3 +47,3 @@
- Which parts of the codebase are untested, or hard to test through their current interface?
The judgment calls stay with you: from the returned facts, decide what is shallow and apply the **deletion test** to anything you suspect — would deleting it concentrate complexity, or just move it? A "yes, concentrates" is the signal you want.
The judgment calls stay with you: from the returned facts, decide what is shallow and apply the **deletion test** to anything you suspect. Complexity that vanishes marks a pass-through, the candidate you want; complexity that reappears across callers means the module is earning its keep.

low — Explore step re-defines the deletion test that the glossary already defines a few paragraphs earlier
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: skills/improve-codebase-architecture/SKILL.md in full, including the glossary's "Key principles" list and step 1 (Explore).

The glossary already defines the test: "Deletion test: imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep." The changed sentence in step 1 states the same definition again: "Complexity that vanishes marks a pass-through, the candidate you want; complexity that reappears across callers means the module is earning its keep."

writing-for-agents ("One Idea, One Place") says to give each instruction one home and not to restate what an earlier sentence already says. The definition now exists in two places, and their wording has already drifted ("across N callers" vs "across callers"). A later edit to one can make the two disagree. The only new information is that a pass-through is "the candidate you want".

Proposed correction: "...decide what is shallow and apply the deletion test to anything you suspect; a pass-through is the candidate you want." This keeps the step's link from the test to a candidate, and leaves the glossary as the only definition.

claim 01M3V6AEK1SYCHEPFXER5AZZF6 of review 01M3V67FK226HY7EDBZ5F16S4T

<!-- review:claim:01M3V6AEK1SYCHEPFXER5AZZF6 --> **low** — Explore step re-defines the deletion test that the glossary already defines a few paragraphs earlier lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: skills/improve-codebase-architecture/SKILL.md in full, including the glossary's "Key principles" list and step 1 (Explore). > > The glossary already defines the test: "**Deletion test**: imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep." The changed sentence in step 1 states the same definition again: "Complexity that vanishes marks a pass-through, the candidate you want; complexity that reappears across callers means the module is earning its keep." > > writing-for-agents ("One Idea, One Place") says to give each instruction one home and not to restate what an earlier sentence already says. The definition now exists in two places, and their wording has already drifted ("across N callers" vs "across callers"). A later edit to one can make the two disagree. The only new information is that a pass-through is "the candidate you want". > > Proposed correction: "...decide what is shallow and apply the **deletion test** to anything you suspect; a pass-through is the candidate you want." This keeps the step's link from the test to a candidate, and leaves the glossary as the only definition. claim `01M3V6AEK1SYCHEPFXER5AZZF6` of review `01M3V67FK226HY7EDBZ5F16S4T`
Author
Owner

Fixed in 83f5156c0a with the proposed wording; the glossary keeps the only definition.

<!-- gh-feedback:reply-to:98338 --> Fixed in 83f5156c0a43b2bf85c398c980aa7456c1c33fc1 with the proposed wording; the glossary keeps the only definition.
jercik marked this conversation as resolved
@ -32,3 +32,3 @@
## Judge independently
For a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:
For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:

medium — review-agent-instructions drops its reviewer-strength requirement and leaves only an implicit cue for an optional, agent-specific rule
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: skills/review-agent-instructions/SKILL.md in full, the two new rules rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, the parallel edits in skills/improve-codebase-architecture/SKILL.md and INTERFACE-DESIGN.md, and README.md's description of rule selection.

The diff replaced "use a fresh reviewer with the strongest available model at its deepest reasoning setting" with the anchored sentence. The only remaining signal of reviewer strength is the appositive "which delivers a conclusion". That phrase matters only as a match for the new rules' top tier ("work whose deliverable is a conclusion: reviewing work, ..."). The same pattern appears in the other skills: "Each delivers a design conclusion" (INTERFACE-DESIGN.md) and "The walk delivers facts, not conclusions" (improve-codebase-architecture/SKILL.md).

The skill does not name the rule or say what tier to choose, and the rule is not guaranteed to be present. README.md says rules must be selected separately: "Registering this source makes its rules available; it does not select them". It also describes the new directories as "selected by an agent override". No rule exists for Cursor, Gemini, OpenCode, or Grok. The Codex rule applies only "When choosing a sub-agent's model explicitly". Without the rule, a literal reader sees "which delivers a conclusion" as a description that changes no behavior. The reviewer then runs at the inherited or default setting, which can be weaker than the agent under test, and the independent-judgment step loses its intended strength.

writing-for-agents says: "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have." A label that changes no behavior also fails its no-op test.

Proposed correction: "For a clean-context review, use a fresh reviewer at your top reasoning tier; it need not be the model under test." This keeps the change's harness neutrality, because it names no model, and works without the rule. When the rule is selected, its explicit tiers decide the setting. A similar short phrase in INTERFACE-DESIGN.md ("Spawn 3+ sub-agents in parallel at your top reasoning tier") would keep that skill portable as well.

Proof gap: I cannot see whether every intended deployment of these skills also selects the matching rule. If that selection is guaranteed, the impact is lower.

claim 01M3V6A850BGN12969VJ4DHC8K of review 01M3V67FK226HY7EDBZ5F16S4T

<!-- review:claim:01M3V6A850BGN12969VJ4DHC8K --> **medium** — review-agent-instructions drops its reviewer-strength requirement and leaves only an implicit cue for an optional, agent-specific rule lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: skills/review-agent-instructions/SKILL.md in full, the two new rules rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, the parallel edits in skills/improve-codebase-architecture/SKILL.md and INTERFACE-DESIGN.md, and README.md's description of rule selection. > > The diff replaced "use a fresh reviewer with the strongest available model at its deepest reasoning setting" with the anchored sentence. The only remaining signal of reviewer strength is the appositive "which delivers a conclusion". That phrase matters only as a match for the new rules' top tier ("work whose deliverable is a conclusion: reviewing work, ..."). The same pattern appears in the other skills: "Each delivers a design conclusion" (INTERFACE-DESIGN.md) and "The walk delivers facts, not conclusions" (improve-codebase-architecture/SKILL.md). > > The skill does not name the rule or say what tier to choose, and the rule is not guaranteed to be present. README.md says rules must be selected separately: "Registering this source makes its rules available; it does not select them". It also describes the new directories as "selected by an agent override". No rule exists for Cursor, Gemini, OpenCode, or Grok. The Codex rule applies only "When choosing a sub-agent's model explicitly". Without the rule, a literal reader sees "which delivers a conclusion" as a description that changes no behavior. The reviewer then runs at the inherited or default setting, which can be weaker than the agent under test, and the independent-judgment step loses its intended strength. > > writing-for-agents says: "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have." A label that changes no behavior also fails its no-op test. > > Proposed correction: "For a clean-context review, use a fresh reviewer at your top reasoning tier; it need not be the model under test." This keeps the change's harness neutrality, because it names no model, and works without the rule. When the rule is selected, its explicit tiers decide the setting. A similar short phrase in INTERFACE-DESIGN.md ("Spawn 3+ sub-agents in parallel at your top reasoning tier") would keep that skill portable as well. > > Proof gap: I cannot see whether every intended deployment of these skills also selects the matching rule. If that selection is guaranteed, the impact is lower. claim `01M3V6A850BGN12969VJ4DHC8K` of review `01M3V67FK226HY7EDBZ5F16S4T`
Author
Owner

This is deliberate. The repository owner decided that skills name no model, effort, or tier, and that model choice belongs to the per-harness rules. Without a selected rule, the harness default picks the reviewer. 'Top reasoning tier' is still a model-strength instruction, which is exactly what the owner removed from skills.

<!-- gh-feedback:reply-to:98336 --> This is deliberate. The repository owner decided that skills name no model, effort, or tier, and that model choice belongs to the per-harness rules. Without a selected rule, the harness default picks the reviewer. 'Top reasoning tier' is still a model-strength instruction, which is exactly what the owner removed from skills.

superseded by review 01M3V6JPN0S0M7BTX1S93K2BAG for head 83f5156c0a43b2bf85c398c980aa7456c1c33fc1

<!-- review:superseded:01M3V6JPN0S0M7BTX1S93K2BAG --> superseded by review `01M3V6JPN0S0M7BTX1S93K2BAG` for head `83f5156c0a43b2bf85c398c980aa7456c1c33fc1`
fix(rules): say where Claude sub-agent effort can be set
Some checks failed
commit-msg / commitlint (pull_request) Failing after 19s
Node tests / node:test (pull_request) Successful in 1m36s
Review / Review (pull_request_target) Successful in 4m24s
83f5156c0a
@ -0,0 +2,4 @@
Launch every sub-agent at one of three settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier, and most sub-agents run on it: routine work with a clear goal, such as running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, applying a specified bulk edit, or implementing a scoped change.

low — Claude sub-agent rule names a nonexistent "Sonnet 5.5" as the model behind the sonnet alias
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the new rule rules/claude/subagent-model-selection.md, which tells the agent to run most sub-agents on "sonnet (Sonnet 5.5) at low" and the conclusion tier on "opus (Opus 5.5)".

What goes wrong: the current Claude 5 family is Fable 5.1, Opus 5.5, Sonnet 5 (claude-sonnet-5), and Haiku 4.5. No Sonnet 5.5 release exists, so the sonnet alias resolves to Sonnet 5. The Opus parenthetical is correct; the Sonnet one is not. The rule is delivered into the agent's user instructions. An agent that trusts the parenthetical will report the wrong model in its reasoning or summaries. Worse, it may try to pin a full model ID it derives from the name (e.g. claude-sonnet-5-5) where a full ID is accepted, such as agent-definition frontmatter, and that ID does not exist. The alias itself (sonnet) is still a valid Agent-tool model value, so a run that uses only the alias is unaffected. That is why this is low severity.

Evidence and proof gap: this is static reasoning against the published model lineup, not something I could check in the repository. Nothing else in the tree names a Sonnet version. A maintainer can confirm by checking the current Claude model list. The fix is to change the parenthetical to "Sonnet 5" or drop the version so the alias is not tied to one release.

claim 01M3V6MMA3WSG7PBB2QPCFZH0N of review 01M3V6JPN0S0M7BTX1S93K2BAG

<!-- review:claim:01M3V6MMA3WSG7PBB2QPCFZH0N --> **low** — Claude sub-agent rule names a nonexistent "Sonnet 5.5" as the model behind the `sonnet` alias lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the new rule `rules/claude/subagent-model-selection.md`, which tells the agent to run most sub-agents on "`sonnet` (Sonnet 5.5) at `low`" and the conclusion tier on "`opus` (Opus 5.5)". > > What goes wrong: the current Claude 5 family is Fable 5.1, Opus 5.5, Sonnet 5 (`claude-sonnet-5`), and Haiku 4.5. No Sonnet 5.5 release exists, so the `sonnet` alias resolves to Sonnet 5. The Opus parenthetical is correct; the Sonnet one is not. The rule is delivered into the agent's user instructions. An agent that trusts the parenthetical will report the wrong model in its reasoning or summaries. Worse, it may try to pin a full model ID it derives from the name (e.g. `claude-sonnet-5-5`) where a full ID is accepted, such as agent-definition frontmatter, and that ID does not exist. The alias itself (`sonnet`) is still a valid Agent-tool model value, so a run that uses only the alias is unaffected. That is why this is low severity. > > Evidence and proof gap: this is static reasoning against the published model lineup, not something I could check in the repository. Nothing else in the tree names a Sonnet version. A maintainer can confirm by checking the current Claude model list. The fix is to change the parenthetical to "Sonnet 5" or drop the version so the alias is not tied to one release. claim `01M3V6MMA3WSG7PBB2QPCFZH0N` of review `01M3V6JPN0S0M7BTX1S93K2BAG`

low — Tier bullets break parallel form, and "most sub-agents run on it" repeats the closing default sentence (both model-selection rules)
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, which share this structure word for word apart from model names.

In both rules the first bullet is a full sentence ending in a colon: "... at low is the regular tier, and most sub-agents run on it: routine work with a clear goal, such as ...". The other two bullets are fragments: "opus (Opus 5.5) at medium for work that takes trial and error to finish: ...". Because the shapes differ, the scan for "which work maps to which setting" is harder than it needs to be. In the first bullet, "most sub-agents run on it" sits between the setting and the work it covers. The default is then stated again in the closing paragraph: "Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion." The writing-for-agents skill asks for one home per instruction ("Do not restate what ... an earlier sentence ... already says") and for bullets used for flat choices, which work best in parallel form.

Proposed correction, applied to both files: "- sonnet at low (regular tier) for routine work with a clear goal: running a given command ...". Keep the default only in the closing sentence. This keeps the tier name and the full list of examples, removes the duplicate default, and gives all three bullets the same shape: setting, then "for", then the work it covers.

Separately, the parenthetical "(Sonnet 5.5)" pins what the sonnet alias resolves to. That mapping changes with Claude Code releases, and the writing-for-agents skill says to scope version-dependent claims. Dropping the parenthetical or adding a version scope would keep the rule true over time. I could not confirm the current alias mapping from inside the sandbox.

claim 01M3V6QDBE4638D64J876D64ET of review 01M3V6JPN0S0M7BTX1S93K2BAG

<!-- review:claim:01M3V6QDBE4638D64J876D64ET --> **low** — Tier bullets break parallel form, and "most sub-agents run on it" repeats the closing default sentence (both model-selection rules) lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, which share this structure word for word apart from model names. > > In both rules the first bullet is a full sentence ending in a colon: "... at `low` is the regular tier, and most sub-agents run on it: routine work with a clear goal, such as ...". The other two bullets are fragments: "`opus` (Opus 5.5) at `medium` for work that takes trial and error to finish: ...". Because the shapes differ, the scan for "which work maps to which setting" is harder than it needs to be. In the first bullet, "most sub-agents run on it" sits between the setting and the work it covers. The default is then stated again in the closing paragraph: "Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion." The writing-for-agents skill asks for one home per instruction ("Do not restate what ... an earlier sentence ... already says") and for bullets used for flat choices, which work best in parallel form. > > Proposed correction, applied to both files: "- `sonnet` at `low` (regular tier) for routine work with a clear goal: running a given command ...". Keep the default only in the closing sentence. This keeps the tier name and the full list of examples, removes the duplicate default, and gives all three bullets the same shape: setting, then "for", then the work it covers. > > Separately, the parenthetical "(Sonnet 5.5)" pins what the `sonnet` alias resolves to. That mapping changes with Claude Code releases, and the writing-for-agents skill says to scope version-dependent claims. Dropping the parenthetical or adding a version scope would keep the rule true over time. I could not confirm the current alias mapping from inside the sandbox. claim `01M3V6QDBE4638D64J876D64ET` of review `01M3V6JPN0S0M7BTX1S93K2BAG`
Author
Owner

Same claim as #98337, already refuted: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Claude Sonnet 5.5 as the default Sonnet model.

<!-- gh-feedback:reply-to:98377 --> Same claim as #98337, already refuted: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Claude Sonnet 5.5 as the default Sonnet model.

superseded by review 01M3V78FSFQ0HV1WGSCMRF0T04 for head d7530a344b372b4d9ee54c3201341d9d40c7cb93

<!-- review:superseded:01M3V78FSFQ0HV1WGSCMRF0T04 --> superseded by review `01M3V78FSFQ0HV1WGSCMRF0T04` for head `d7530a344b372b4d9ee54c3201341d9d40c7cb93`
@ -0,0 +6,4 @@
- `opus` (Opus 5.5) at `medium` for work that takes trial and error to finish: getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.
- `opus` at `xhigh` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.
Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both explicitly in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; its effort comes from the agent type's definition.

medium — Claude model-selection rule requires a model+effort pair for every sub-agent, then admits the Agent tool cannot set effort, without saying what to do
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: rules/claude/subagent-model-selection.md (new), together with rules/general/subagent-delegation.md ("Spawn Sub-Agents Liberally"), which makes the Agent tool the main way sub-agents get launched.

The rule opens with an unconditional mandate: "Launch every sub-agent at one of three settings:". Each setting is a model plus an effort (sonnet at low, opus at medium, opus at xhigh). The last paragraph then says: "A direct Agent tool call sets only the model; its effort comes from the agent type's definition." So the most common launch path cannot meet the mandate, and the rule never says what the agent should do in that case. A literal reader has several plausible responses. It could set only model and accept whatever effort the agent type carries. It could hunt for an agent type whose definition has the right effort. It could switch to a workflow agent() call, although Claude Code allows workflows only when the user explicitly opts in. It could also write a new agent definition. The writing-for-agents skill says to keep an exception "when the trigger is silent or several responses are plausible" and to pair a constraint with the safe behavior. Here the constraint and the limitation appear side by side, with no resolution between them. The Codex sibling rule avoids the problem by scoping itself to "When choosing a sub-agent's model explicitly".

Proposed correction: scope the opening to settings the launch path can actually control, and give the Agent-tool fallback. For example: "Choose each sub-agent's tier from the three below. Set both model and effort in workflow agent() calls and agent-definition frontmatter. A direct Agent tool call sets only model, so pick the tier's model and an agent type whose definition carries the tier's effort; when none does, [the chosen fallback]." The author has to pick the fallback, because it is a policy choice. This keeps the tier table and the ceiling, and it removes the contradiction between "every sub-agent" and the Agent tool's limitation.

Evidence: static reading only. I did not test a Claude Code session to see which response a model actually picks. The Agent tool's inability to set effort is taken from the rule's own text.

claim 01M3V6NXGTF7797DV8N86DMAVT of review 01M3V6JPN0S0M7BTX1S93K2BAG

<!-- review:claim:01M3V6NXGTF7797DV8N86DMAVT --> **medium** — Claude model-selection rule requires a model+effort pair for every sub-agent, then admits the Agent tool cannot set effort, without saying what to do lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: rules/claude/subagent-model-selection.md (new), together with rules/general/subagent-delegation.md ("Spawn Sub-Agents Liberally"), which makes the Agent tool the main way sub-agents get launched. > > The rule opens with an unconditional mandate: "Launch every sub-agent at one of three settings:". Each setting is a model plus an effort (`sonnet` at `low`, `opus` at `medium`, `opus` at `xhigh`). The last paragraph then says: "A direct Agent tool call sets only the `model`; its effort comes from the agent type's definition." So the most common launch path cannot meet the mandate, and the rule never says what the agent should do in that case. A literal reader has several plausible responses. It could set only `model` and accept whatever effort the agent type carries. It could hunt for an agent type whose definition has the right effort. It could switch to a workflow `agent()` call, although Claude Code allows workflows only when the user explicitly opts in. It could also write a new agent definition. The writing-for-agents skill says to keep an exception "when the trigger is silent or several responses are plausible" and to pair a constraint with the safe behavior. Here the constraint and the limitation appear side by side, with no resolution between them. The Codex sibling rule avoids the problem by scoping itself to "When choosing a sub-agent's model explicitly". > > Proposed correction: scope the opening to settings the launch path can actually control, and give the Agent-tool fallback. For example: "Choose each sub-agent's tier from the three below. Set both model and effort in workflow `agent()` calls and agent-definition frontmatter. A direct Agent tool call sets only `model`, so pick the tier's model and an agent type whose definition carries the tier's effort; when none does, [the chosen fallback]." The author has to pick the fallback, because it is a policy choice. This keeps the tier table and the ceiling, and it removes the contradiction between "every sub-agent" and the Agent tool's limitation. > > Evidence: static reading only. I did not test a Claude Code session to see which response a model actually picks. The Agent tool's inability to set effort is taken from the rule's own text. claim `01M3V6NXGTF7797DV8N86DMAVT` of review `01M3V6JPN0S0M7BTX1S93K2BAG`
jercik marked this conversation as resolved
@ -32,3 +32,3 @@
## Judge independently
For a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:
For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:

low — "a fresh reviewer, which delivers a conclusion" hides a model-tier signal in a descriptive clause that changes no behavior for a reader without the tier rule
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: skills/review-agent-instructions/SKILL.md, and the new rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. Those rules sort sub-agents into tiers, and their top tier is "work whose deliverable is a conclusion".

The change removed the explicit reviewer requirement ("the strongest available model at its deepest reasoning setting") and put this in its place: "use a fresh reviewer, which delivers a conclusion and need not be the model under test." The phrase "delivers a conclusion" exists to trigger the rules' conclusion tier. In this skill, though, it sits in a non-restrictive relative clause that reads as a description of the reviewer, and it is joined with "and" to an unrelated fact about model identity. The skill is portable. The README says agent-specific rules are "selected by an agent override" and that registering a source "does not select" its rules. A reader without one of those rules gets a label that changes no behavior. That fails the writing-for-agents no-op test ("A label that changes no behavior ... fails the no-op test"). A reader with the rule has to notice that a passing clause is a tier selector. The skill also calls the reviewer's output "findings" and a pass/fail/inconclusive verdict, not a "conclusion", so the cue uses a different term from the rest of the skill ("use one term for one concept").

Proposed correction: state the requirement in its own sentence. For example: "For a clean-context review, launch a fresh reviewer; it need not be the model under test. Its deliverable is a judgment, so run it at the setting you use for conclusion work." Alternatively, delete "which delivers a conclusion" if tier selection is meant to come only from the rule. Either version keeps the "need not be the model under test" permission and makes the model-choice intent explicit rather than incidental.

Static reading only; I have no change history explaining the intent beyond the diff.

claim 01M3V6PG5MQH0Z4XKF332THC64 of review 01M3V6JPN0S0M7BTX1S93K2BAG

<!-- review:claim:01M3V6PG5MQH0Z4XKF332THC64 --> **low** — "a fresh reviewer, which delivers a conclusion" hides a model-tier signal in a descriptive clause that changes no behavior for a reader without the tier rule lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: skills/review-agent-instructions/SKILL.md, and the new rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. Those rules sort sub-agents into tiers, and their top tier is "work whose deliverable is a conclusion". > > The change removed the explicit reviewer requirement ("the strongest available model at its deepest reasoning setting") and put this in its place: "use a fresh reviewer, which delivers a conclusion and need not be the model under test." The phrase "delivers a conclusion" exists to trigger the rules' conclusion tier. In this skill, though, it sits in a non-restrictive relative clause that reads as a description of the reviewer, and it is joined with "and" to an unrelated fact about model identity. The skill is portable. The README says agent-specific rules are "selected by an agent override" and that registering a source "does not select" its rules. A reader without one of those rules gets a label that changes no behavior. That fails the writing-for-agents no-op test ("A label that changes no behavior ... fails the no-op test"). A reader with the rule has to notice that a passing clause is a tier selector. The skill also calls the reviewer's output "findings" and a pass/fail/inconclusive verdict, not a "conclusion", so the cue uses a different term from the rest of the skill ("use one term for one concept"). > > Proposed correction: state the requirement in its own sentence. For example: "For a clean-context review, launch a fresh reviewer; it need not be the model under test. Its deliverable is a judgment, so run it at the setting you use for conclusion work." Alternatively, delete "which delivers a conclusion" if tier selection is meant to come only from the rule. Either version keeps the "need not be the model under test" permission and makes the model-choice intent explicit rather than incidental. > > Static reading only; I have no change history explaining the intent beyond the diff. claim `01M3V6PG5MQH0Z4XKF332THC64` of review `01M3V6JPN0S0M7BTX1S93K2BAG`
Author
Owner

Same point as #98336: the owner removed model-strength guidance from skills on purpose and left model choice to the per-harness rules.

<!-- gh-feedback:reply-to:98379 --> Same point as #98336: the owner removed model-strength guidance from skills on purpose and left model choice to the per-harness rules.

superseded by review 01M3V78FSFQ0HV1WGSCMRF0T04 for head d7530a344b372b4d9ee54c3201341d9d40c7cb93

<!-- review:superseded:01M3V78FSFQ0HV1WGSCMRF0T04 --> superseded by review `01M3V78FSFQ0HV1WGSCMRF0T04` for head `d7530a344b372b4d9ee54c3201341d9d40c7cb93`
Author
Owner

Replying to review comment #98376

Round 4: per the review round gate, I'm not pushing these, and both stay acknowledged.

  • #98376 is real but small and deferred to a follow-up. In rules/claude/subagent-model-selection.md, append: "To give a direct call its tier's effort, use an agent type whose definition sets that effort."
  • #98378 (parallel form, repeated "most sub-agents") is a style point that changes no behavior.
> Replying to review comment #98376 Round 4: per the review round gate, I'm not pushing these, and both stay `acknowledged`. - #98376 is real but small and deferred to a follow-up. In `rules/claude/subagent-model-selection.md`, append: "To give a direct call its tier's effort, use an agent type whose definition sets that effort." - #98378 (parallel form, repeated "most sub-agents") is a style point that changes no behavior.
feat(rules): describe sub-agent tiers by category, with exploration and watching on the regular tier
Some checks failed
commit-msg / commitlint (pull_request) Failing after 18s
Node tests / node:test (pull_request) Successful in 1m31s
Review / Review (pull_request_target) Successful in 3m39s
d7530a344b
@ -0,0 +2,4 @@
Launch every sub-agent at one of three settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier, and most sub-agents run on it: exploring and searching code, researching how a codebase works, running commands, watching for changes such as a CI run or deploy, and routine edits up to a scoped change.

low — Claude sub-agent rule names a nonexistent "Sonnet 5.5" as the model behind the sonnet alias
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the new rule rules/claude/subagent-model-selection.md, which defines the three sub-agent tiers. Its first bullet says the regular tier is "sonnet (Sonnet 5.5) at low", and its second says "opus (Opus 5.5) at medium".

What goes wrong: the current Claude model list provided to this run gives Opus 5.5 (claude-opus-5-5) and Sonnet 5 (claude-sonnet-5). It has no Sonnet 5.5. The parenthetical seems to copy the Opus version number onto Sonnet. The rule is delivered into agents' user instruction files, so an agent that reads it gets a wrong model name. If the agent spells the model out instead of using the sonnet alias (for example claude-sonnet-5-5 in agent definition frontmatter, which the same rule tells it to set), the launch fails or resolves to the wrong model. Users reading the rule are also told the wrong model generation.

Evidence and gap: this is static reasoning from the model list given to this run (current as of its knowledge). I did not query a live model catalog. A Sonnet 5.5 released later would refute the claim. Fix: change the parenthetical to "Sonnet 5", or remove the version numbers so the rule only names the aliases.

claim 01M3V7B2483A4CDZP4KKAJ1AV9 of review 01M3V78FSFQ0HV1WGSCMRF0T04

<!-- review:claim:01M3V7B2483A4CDZP4KKAJ1AV9 --> **low** — Claude sub-agent rule names a nonexistent "Sonnet 5.5" as the model behind the `sonnet` alias lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the new rule `rules/claude/subagent-model-selection.md`, which defines the three sub-agent tiers. Its first bullet says the regular tier is "`sonnet` (Sonnet 5.5) at `low`", and its second says "`opus` (Opus 5.5) at `medium`". > > What goes wrong: the current Claude model list provided to this run gives Opus 5.5 (`claude-opus-5-5`) and Sonnet 5 (`claude-sonnet-5`). It has no Sonnet 5.5. The parenthetical seems to copy the Opus version number onto Sonnet. The rule is delivered into agents' user instruction files, so an agent that reads it gets a wrong model name. If the agent spells the model out instead of using the `sonnet` alias (for example `claude-sonnet-5-5` in agent definition frontmatter, which the same rule tells it to set), the launch fails or resolves to the wrong model. Users reading the rule are also told the wrong model generation. > > Evidence and gap: this is static reasoning from the model list given to this run (current as of its knowledge). I did not query a live model catalog. A Sonnet 5.5 released later would refute the claim. Fix: change the parenthetical to "Sonnet 5", or remove the version numbers so the rule only names the aliases. claim `01M3V7B2483A4CDZP4KKAJ1AV9` of review `01M3V78FSFQ0HV1WGSCMRF0T04`
Author
Owner

Repeat of #98337: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5 (Sonnet 5.5).

<!-- gh-feedback:reply-to:98437 --> Repeat of #98337: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5 (Sonnet 5.5).

superseded by review 01M3V7H8W09SMH6TCFF6K8FBMP for head 714d376d368b6b4b51f1c135bafbe3bf9e4fd14b

<!-- review:superseded:01M3V7H8W09SMH6TCFF6K8FBMP --> superseded by review `01M3V7H8W09SMH6TCFF6K8FBMP` for head `714d376d368b6b4b51f1c135bafbe3bf9e4fd14b`

superseded by review 01M3V7H8W09SMH6TCFF6K8FBMP for head 714d376d368b6b4b51f1c135bafbe3bf9e4fd14b

<!-- review:superseded:01M3V7H8W09SMH6TCFF6K8FBMP --> superseded by review `01M3V7H8W09SMH6TCFF6K8FBMP` for head `714d376d368b6b4b51f1c135bafbe3bf9e4fd14b`
@ -0,0 +6,4 @@
- `opus` (Opus 5.5) at `medium` for work that takes trial and error to finish, such as getting a failing test suite green or driving a multi-step browser flow.
- `opus` at `xhigh` for work whose deliverable is a conclusion: reviews, diagnoses, designs, and decisions between approaches.
Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both explicitly in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; its effort comes from the agent type's definition.

medium — Claude sub-agent rule mandates a model+effort setting for every sub-agent, then says direct Agent calls cannot set effort, with no instruction for that case
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: the new rule rules/claude/subagent-model-selection.md in full.

The rule opens with an absolute directive: "Launch every sub-agent at one of three settings:" — sonnet at low, opus at medium, opus at xhigh. Its last paragraph then says "Set both explicitly in workflow agent() calls and agent definition frontmatter. A direct Agent tool call sets only the model; its effort comes from the agent type's definition."

The direct Agent tool call is the most common way to spawn a sub-agent, and the rule itself says effort cannot be set there. So a literal reader who needs opus at xhigh for a review via the Agent tool cannot comply with "every sub-agent at one of three settings", and the rule gives no safe behavior: choose an agent type whose definition carries the needed effort, define one, use a workflow agent() call, or accept the type's effort. The final sentence reads as a fact about the harness rather than as an instruction, so the agent is left to improvise (most likely: pass model and ignore effort, silently running reviews at whatever effort the type defines). writing-for-agents asks to pair a boundary with the safe behavior and to keep exceptions whose response is not obvious ("Keep it when the trigger is silent or several responses are plausible"). A smaller ambiguity in the same paragraph: "Set both" has no nearby antecedent; "model and effort" appears only in the title.

Proposed correction: replace the last two sentences with something like "Set model and effort explicitly in workflow agent() calls and agent definition frontmatter. A direct Agent tool call sets only model and takes effort from the agent type's definition; for a tier whose effort differs from that, pick or define an agent type with that effort [or: launch it through a workflow agent() call]." This keeps the harness fact and adds the missing action. Decisive evidence would be the author's intended handling of direct calls; I did not verify Claude Code's Agent tool behavior beyond what the rule states.

claim 01M3V7B3VF5RE7YFAJ2B8VG0JZ of review 01M3V78FSFQ0HV1WGSCMRF0T04

<!-- review:claim:01M3V7B3VF5RE7YFAJ2B8VG0JZ --> **medium** — Claude sub-agent rule mandates a model+effort setting for every sub-agent, then says direct Agent calls cannot set effort, with no instruction for that case lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: the new rule `rules/claude/subagent-model-selection.md` in full. > > The rule opens with an absolute directive: "Launch every sub-agent at one of three settings:" — `sonnet` at `low`, `opus` at `medium`, `opus` at `xhigh`. Its last paragraph then says "Set both explicitly in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; its effort comes from the agent type's definition." > > The direct Agent tool call is the most common way to spawn a sub-agent, and the rule itself says effort cannot be set there. So a literal reader who needs `opus` at `xhigh` for a review via the Agent tool cannot comply with "every sub-agent at one of three settings", and the rule gives no safe behavior: choose an agent type whose definition carries the needed effort, define one, use a workflow `agent()` call, or accept the type's effort. The final sentence reads as a fact about the harness rather than as an instruction, so the agent is left to improvise (most likely: pass `model` and ignore effort, silently running reviews at whatever effort the type defines). writing-for-agents asks to pair a boundary with the safe behavior and to keep exceptions whose response is not obvious ("Keep it when the trigger is silent or several responses are plausible"). A smaller ambiguity in the same paragraph: "Set both" has no nearby antecedent; "model and effort" appears only in the title. > > Proposed correction: replace the last two sentences with something like "Set `model` and effort explicitly in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only `model` and takes effort from the agent type's definition; for a tier whose effort differs from that, pick or define an agent type with that effort [or: launch it through a workflow `agent()` call]." This keeps the harness fact and adds the missing action. Decisive evidence would be the author's intended handling of direct calls; I did not verify Claude Code's Agent tool behavior beyond what the rule states. claim `01M3V7B3VF5RE7YFAJ2B8VG0JZ` of review `01M3V78FSFQ0HV1WGSCMRF0T04`
Author
Owner

Fixed in 714d376d36: a direct Agent tool call gets the tier's effort through an agent type whose definition sets it.

<!-- gh-feedback:reply-to:98435 --> Fixed in 714d376d368b6b4b51f1c135bafbe3bf9e4fd14b: a direct Agent tool call gets the tier's effort through an agent type whose definition sets it.
jercik marked this conversation as resolved
@ -32,3 +32,3 @@
## Judge independently
For a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:
For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:

medium — Skills swap explicit model/effort guidance for coded "delivers a conclusion" / "facts, not conclusions" hints that only mean something to readers with the Claude/Codex sub-agent rule
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: the three skill edits in this change and the two new rules they now lean on (rules/claude/subagent-model-selection.md, rules/codex/subagent-model-selection.md), plus README.md's Layout and How-it-works sections.

What changed: three skills dropped their explicit model guidance and kept only a classification phrase:

  • skills/review-agent-instructions/SKILL.md went from "use a fresh reviewer with the strongest available model at its deepest reasoning setting" to "use a fresh reviewer, which delivers a conclusion and need not be the model under test."
  • skills/improve-codebase-architecture/SKILL.md now says "Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions."
  • skills/improve-codebase-architecture/INTERFACE-DESIGN.md now says "Each delivers a design conclusion, not a fact list, and must produce...".

These phrases only work as cues for the new rules, which map "work whose deliverable is a conclusion" to opus at xhigh / gpt-6-astra at high. Those rules live under rules/claude/ and rules/codex/, and README.md describes them as "selected by an agent override". The README says delivery also supports Cursor, Gemini, OpenCode, and Grok, and none of them gets a matching rule. Claude and Codex users who have not selected the override don't get it either. For those readers, "which delivers a conclusion" no longer tells them anything to do: it fails the no-op test, and the clean-context reviewer now runs at whatever default the harness picks instead of the deepest setting. The relative clause also reads oddly, because it attaches an unrelated fact to "need not be the model under test." writing-for-agents says: "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have."

Proposed correction: say the tier in harness-neutral words wherever the skill depends on it, e.g. "use a fresh reviewer on your strongest model at its deepest reasoning setting; it need not be the model under test", and "spawn a read-only exploration sub-agent on your regular (low-cost) tier; the walk returns facts, and the judgment calls stay with you." That keeps the facts-versus-conclusion distinction the rules key on and still gives a directive to readers without the rule. To refute this, show that every supported agent always receives an equivalent rule. Nothing in the tree shows that.

claim 01M3V7BP6VTCA5C7S6KW35WQ98 of review 01M3V78FSFQ0HV1WGSCMRF0T04

<!-- review:claim:01M3V7BP6VTCA5C7S6KW35WQ98 --> **medium** — Skills swap explicit model/effort guidance for coded "delivers a conclusion" / "facts, not conclusions" hints that only mean something to readers with the Claude/Codex sub-agent rule lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: the three skill edits in this change and the two new rules they now lean on (`rules/claude/subagent-model-selection.md`, `rules/codex/subagent-model-selection.md`), plus README.md's Layout and How-it-works sections. > > What changed: three skills dropped their explicit model guidance and kept only a classification phrase: > - `skills/review-agent-instructions/SKILL.md` went from "use a fresh reviewer with the strongest available model at its deepest reasoning setting" to "use a fresh reviewer, which delivers a conclusion and need not be the model under test." > - `skills/improve-codebase-architecture/SKILL.md` now says "Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions." > - `skills/improve-codebase-architecture/INTERFACE-DESIGN.md` now says "Each delivers a design conclusion, not a fact list, and must produce...". > > These phrases only work as cues for the new rules, which map "work whose deliverable is a conclusion" to `opus` at `xhigh` / `gpt-6-astra` at `high`. Those rules live under `rules/claude/` and `rules/codex/`, and README.md describes them as "selected by an agent override". The README says delivery also supports Cursor, Gemini, OpenCode, and Grok, and none of them gets a matching rule. Claude and Codex users who have not selected the override don't get it either. For those readers, "which delivers a conclusion" no longer tells them anything to do: it fails the no-op test, and the clean-context reviewer now runs at whatever default the harness picks instead of the deepest setting. The relative clause also reads oddly, because it attaches an unrelated fact to "need not be the model under test." writing-for-agents says: "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have." > > Proposed correction: say the tier in harness-neutral words wherever the skill depends on it, e.g. "use a fresh reviewer on your strongest model at its deepest reasoning setting; it need not be the model under test", and "spawn a read-only exploration sub-agent on your regular (low-cost) tier; the walk returns facts, and the judgment calls stay with you." That keeps the facts-versus-conclusion distinction the rules key on and still gives a directive to readers without the rule. To refute this, show that every supported agent always receives an equivalent rule. Nothing in the tree shows that. claim `01M3V7BP6VTCA5C7S6KW35WQ98` of review `01M3V78FSFQ0HV1WGSCMRF0T04`
Author
Owner

Repeat of #98336: the owner removed model guidance from skills on purpose; the per-harness rules own it.

<!-- gh-feedback:reply-to:98436 --> Repeat of #98336: the owner removed model guidance from skills on purpose; the per-harness rules own it.

superseded by review 01M3V7H8W09SMH6TCFF6K8FBMP for head 714d376d368b6b4b51f1c135bafbe3bf9e4fd14b

<!-- review:superseded:01M3V7H8W09SMH6TCFF6K8FBMP --> superseded by review `01M3V7H8W09SMH6TCFF6K8FBMP` for head `714d376d368b6b4b51f1c135bafbe3bf9e4fd14b`

superseded by review 01M3V7H8W09SMH6TCFF6K8FBMP for head 714d376d368b6b4b51f1c135bafbe3bf9e4fd14b

<!-- review:superseded:01M3V7H8W09SMH6TCFF6K8FBMP --> superseded by review `01M3V7H8W09SMH6TCFF6K8FBMP` for head `714d376d368b6b4b51f1c135bafbe3bf9e4fd14b`
feat(rules): reduce sub-agent selection to a regular and a higher tier
Some checks failed
commit-msg / commitlint (pull_request) Failing after 18s
Node tests / node:test (pull_request) Successful in 1m35s
Review / Review (pull_request_target) Successful in 5m23s
714d376d36
@ -0,0 +2,4 @@
Run every sub-agent at one of two settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.

low — Claude sub-agent rule says the sonnet alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "sonnet (Sonnet 5.5) at low" and higher-tier ones as "opus (Opus 5.5) at xhigh".

What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's model enum is sonnet/opus/haiku/fable, so the sonnet alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions.

Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the sonnet alias points to it on the user's install, there is no defect. Running claude with --model sonnet and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale.

claim 01M3V7KWGWXB9CJ1291P0CYN6K of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7KWGWXB9CJ1291P0CYN6K --> **low** — Claude sub-agent rule says the `sonnet` alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "`sonnet` (Sonnet 5.5) at `low`" and higher-tier ones as "`opus` (Opus 5.5) at `xhigh`". > > What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's `model` enum is sonnet/opus/haiku/fable, so the `sonnet` alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions. > > Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the `sonnet` alias points to it on the user's install, there is no defect. Running `claude` with `--model sonnet` and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale. claim `01M3V7KWGWXB9CJ1291P0CYN6K` of review `01M3V7H8W09SMH6TCFF6K8FBMP`

low — Claude sub-agent rule says the sonnet alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "sonnet (Sonnet 5.5) at low" and higher-tier ones as "opus (Opus 5.5) at xhigh".

What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's model enum is sonnet/opus/haiku/fable, so the sonnet alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions.

Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the sonnet alias points to it on the user's install, there is no defect. Running claude with --model sonnet and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale.

claim 01M3V7KWGWXB9CJ1291P0CYN6K of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7KWGWXB9CJ1291P0CYN6K --> **low** — Claude sub-agent rule says the `sonnet` alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "`sonnet` (Sonnet 5.5) at `low`" and higher-tier ones as "`opus` (Opus 5.5) at `xhigh`". > > What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's `model` enum is sonnet/opus/haiku/fable, so the `sonnet` alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions. > > Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the `sonnet` alias points to it on the user's install, there is no defect. Running `claude` with `--model sonnet` and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale. claim `01M3V7KWGWXB9CJ1291P0CYN6K` of review `01M3V7H8W09SMH6TCFF6K8FBMP`
Author
Owner

Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.

<!-- gh-feedback:reply-to:98473 --> Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.
Author
Owner

Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.

<!-- gh-feedback:reply-to:98474 --> Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.

superseded by review 01M3VA1JHFYRT7H3A16SX436B9 for head 17e17788f269a2a307d94a1f33e7636bc17d7b7e

<!-- review:superseded:01M3VA1JHFYRT7H3A16SX436B9 --> superseded by review `01M3VA1JHFYRT7H3A16SX436B9` for head `17e17788f269a2a307d94a1f33e7636bc17d7b7e`
@ -0,0 +3,4 @@
Run every sub-agent at one of two settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.
- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.

low — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "sonnet (Sonnet 5.5) at low is the regular tier for most work: ...". The second bullet has no verb and no name: "opus (Opus 5.5) at xhigh for work that needs original thought: ..." (Codex: "gpt-6-astra at high for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on."

A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment.

Suggested fix, applied to both files: "- opus (Opus 5.5) at xhigh is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined.

claim 01M3V7Q640RTJ7EP4J40EPHPT2 of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7Q640RTJ7EP4J40EPHPT2 --> **low** — Second tier bullet has no verb or name, so the later "higher tier" term is never defined lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "`sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: ...". The second bullet has no verb and no name: "`opus` (Opus 5.5) at `xhigh` for work that needs original thought: ..." (Codex: "`gpt-6-astra` at `high` for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on." > > A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment. > > Suggested fix, applied to both files: "- `opus` (Opus 5.5) at `xhigh` is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined. claim `01M3V7Q640RTJ7EP4J40EPHPT2` of review `01M3V7H8W09SMH6TCFF6K8FBMP`

low — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "sonnet (Sonnet 5.5) at low is the regular tier for most work: ...". The second bullet has no verb and no name: "opus (Opus 5.5) at xhigh for work that needs original thought: ..." (Codex: "gpt-6-astra at high for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on."

A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment.

Suggested fix, applied to both files: "- opus (Opus 5.5) at xhigh is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined.

claim 01M3V7Q640RTJ7EP4J40EPHPT2 of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7Q640RTJ7EP4J40EPHPT2 --> **low** — Second tier bullet has no verb or name, so the later "higher tier" term is never defined lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "`sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: ...". The second bullet has no verb and no name: "`opus` (Opus 5.5) at `xhigh` for work that needs original thought: ..." (Codex: "`gpt-6-astra` at `high` for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on." > > A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment. > > Suggested fix, applied to both files: "- `opus` (Opus 5.5) at `xhigh` is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined. claim `01M3V7Q640RTJ7EP4J40EPHPT2` of review `01M3V7H8W09SMH6TCFF6K8FBMP`
jercik marked this conversation as resolved
@ -0,0 +5,4 @@
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.
- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.
Use the higher tier only when the deliverable is a judgment the caller will rely on. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort.

medium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (grep -rn -i "effort\|xhigh"). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (sonnet at low, opus at xhigh). For the most common spawn path it then says: "A direct Agent tool call sets only the model; to give it the tier's effort, use an agent type whose definition sets that effort."

This repository ships no agent definitions. The search finds no .claude/agents files and no frontmatter that sets low or xhigh; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only model set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place").

Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set model alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork.

Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them.

claim 01M3V7PKW97BMNCVHYZQ2MW1PY of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7PKW97BMNCVHYZQ2MW1PY --> **medium** — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (`grep -rn -i "effort\|xhigh"`). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (`sonnet` at `low`, `opus` at `xhigh`). For the most common spawn path it then says: "A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort." > > This repository ships no agent definitions. The search finds no `.claude/agents` files and no frontmatter that sets `low` or `xhigh`; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only `model` set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place"). > > Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set `model` alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork. > > Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them. claim `01M3V7PKW97BMNCVHYZQ2MW1PY` of review `01M3V7H8W09SMH6TCFF6K8FBMP`

medium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (grep -rn -i "effort\|xhigh"). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (sonnet at low, opus at xhigh). For the most common spawn path it then says: "A direct Agent tool call sets only the model; to give it the tier's effort, use an agent type whose definition sets that effort."

This repository ships no agent definitions. The search finds no .claude/agents files and no frontmatter that sets low or xhigh; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only model set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place").

Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set model alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork.

Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them.

claim 01M3V7PKW97BMNCVHYZQ2MW1PY of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7PKW97BMNCVHYZQ2MW1PY --> **medium** — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (`grep -rn -i "effort\|xhigh"`). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (`sonnet` at `low`, `opus` at `xhigh`). For the most common spawn path it then says: "A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort." > > This repository ships no agent definitions. The search finds no `.claude/agents` files and no frontmatter that sets `low` or `xhigh`; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only `model` set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place"). > > Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set `model` alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork. > > Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them. claim `01M3V7PKW97BMNCVHYZQ2MW1PY` of review `01M3V7H8W09SMH6TCFF6K8FBMP`
jercik marked this conversation as resolved
@ -0,0 +2,4 @@
Run every sub-agent at one of two settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.

low — Claude sub-agent rule says the sonnet alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "sonnet (Sonnet 5.5) at low" and higher-tier ones as "opus (Opus 5.5) at xhigh".

What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's model enum is sonnet/opus/haiku/fable, so the sonnet alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions.

Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the sonnet alias points to it on the user's install, there is no defect. Running claude with --model sonnet and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale.

claim 01M3V7KWGWXB9CJ1291P0CYN6K of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7KWGWXB9CJ1291P0CYN6K --> **low** — Claude sub-agent rule says the `sonnet` alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "`sonnet` (Sonnet 5.5) at `low`" and higher-tier ones as "`opus` (Opus 5.5) at `xhigh`". > > What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's `model` enum is sonnet/opus/haiku/fable, so the `sonnet` alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions. > > Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the `sonnet` alias points to it on the user's install, there is no defect. Running `claude` with `--model sonnet` and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale. claim `01M3V7KWGWXB9CJ1291P0CYN6K` of review `01M3V7H8W09SMH6TCFF6K8FBMP`

low — Claude sub-agent rule says the sonnet alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "sonnet (Sonnet 5.5) at low" and higher-tier ones as "opus (Opus 5.5) at xhigh".

What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's model enum is sonnet/opus/haiku/fable, so the sonnet alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions.

Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the sonnet alias points to it on the user's install, there is no defect. Running claude with --model sonnet and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale.

claim 01M3V7KWGWXB9CJ1291P0CYN6K of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7KWGWXB9CJ1291P0CYN6K --> **low** — Claude sub-agent rule says the `sonnet` alias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnet lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the new rule rules/claude/subagent-model-selection.md, plus the model catalog that the Claude Code harness running this review declares. The rule tells Claude Code to run regular-tier sub-agents as "`sonnet` (Sonnet 5.5) at `low`" and higher-tier ones as "`opus` (Opus 5.5) at `xhigh`". > > What goes wrong: according to the harness environment for this run, the most recent Claude models are Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), Sonnet 5 (claude-sonnet-5) and Haiku 4.5. The Agent tool's `model` enum is sonnet/opus/haiku/fable, so the `sonnet` alias resolves to Sonnet 5. No Sonnet 5.5 is listed. The Opus 5.5 label matches the harness, but the Sonnet label does not. A reader of the rule, or an agent checking whether the regular tier is strong enough for a task, is told the wrong model generation. This rule is the only place that now names the model, because this same change removed the per-skill model guidance from improve-codebase-architecture and review-agent-instructions. > > Evidence type: static comparison against the harness's declared model list. I ran nothing. Proof gap: if Sonnet 5.5 shipped and the `sonnet` alias points to it on the user's install, there is no defect. Running `claude` with `--model sonnet` and checking which model it resolves to would settle it. Fix: change the parenthetical to the model the alias actually resolves to, or drop the version so the label cannot go stale. claim `01M3V7KWGWXB9CJ1291P0CYN6K` of review `01M3V7H8W09SMH6TCFF6K8FBMP`
Author
Owner

Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.

<!-- gh-feedback:reply-to:98473 --> Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.
Author
Owner

Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.

<!-- gh-feedback:reply-to:98474 --> Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.

superseded by review 01M3VA1JHFYRT7H3A16SX436B9 for head 17e17788f269a2a307d94a1f33e7636bc17d7b7e

<!-- review:superseded:01M3VA1JHFYRT7H3A16SX436B9 --> superseded by review `01M3VA1JHFYRT7H3A16SX436B9` for head `17e17788f269a2a307d94a1f33e7636bc17d7b7e`
@ -0,0 +3,4 @@
Run every sub-agent at one of two settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.
- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.

low — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "sonnet (Sonnet 5.5) at low is the regular tier for most work: ...". The second bullet has no verb and no name: "opus (Opus 5.5) at xhigh for work that needs original thought: ..." (Codex: "gpt-6-astra at high for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on."

A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment.

Suggested fix, applied to both files: "- opus (Opus 5.5) at xhigh is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined.

claim 01M3V7Q640RTJ7EP4J40EPHPT2 of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7Q640RTJ7EP4J40EPHPT2 --> **low** — Second tier bullet has no verb or name, so the later "higher tier" term is never defined lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "`sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: ...". The second bullet has no verb and no name: "`opus` (Opus 5.5) at `xhigh` for work that needs original thought: ..." (Codex: "`gpt-6-astra` at `high` for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on." > > A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment. > > Suggested fix, applied to both files: "- `opus` (Opus 5.5) at `xhigh` is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined. claim `01M3V7Q640RTJ7EP4J40EPHPT2` of review `01M3V7H8W09SMH6TCFF6K8FBMP`

low — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "sonnet (Sonnet 5.5) at low is the regular tier for most work: ...". The second bullet has no verb and no name: "opus (Opus 5.5) at xhigh for work that needs original thought: ..." (Codex: "gpt-6-astra at high for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on."

A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment.

Suggested fix, applied to both files: "- opus (Opus 5.5) at xhigh is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined.

claim 01M3V7Q640RTJ7EP4J40EPHPT2 of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7Q640RTJ7EP4J40EPHPT2 --> **low** — Second tier bullet has no verb or name, so the later "higher tier" term is never defined lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: both new rule files, rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md. The first bullet names its tier: "`sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: ...". The second bullet has no verb and no name: "`opus` (Opus 5.5) at `xhigh` for work that needs original thought: ..." (Codex: "`gpt-6-astra` at `high` for work that needs original thought: ..."). The closing paragraph then relies on a term neither bullet defines: "Use the higher tier only when the deliverable is a judgment the caller will rely on." > > A reader can work out that "higher tier" means the second bullet, but the rule uses "regular tier" and "higher tier" as its two key terms and defines only one. The writing-for-agents skill says to "use one term for one concept, and define project-local terms on first use" ("Use Precise Language"). The broken parallel structure also makes the second bullet read like a fragment. > > Suggested fix, applied to both files: "- `opus` (Opus 5.5) at `xhigh` is the higher tier, for work that needs original thought: ...". This keeps every example and makes the closing sentence's term refer to something defined. claim `01M3V7Q640RTJ7EP4J40EPHPT2` of review `01M3V7H8W09SMH6TCFF6K8FBMP`
jercik marked this conversation as resolved
@ -0,0 +5,4 @@
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.
- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.
Use the higher tier only when the deliverable is a judgment the caller will rely on. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort.

medium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (grep -rn -i "effort\|xhigh"). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (sonnet at low, opus at xhigh). For the most common spawn path it then says: "A direct Agent tool call sets only the model; to give it the tier's effort, use an agent type whose definition sets that effort."

This repository ships no agent definitions. The search finds no .claude/agents files and no frontmatter that sets low or xhigh; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only model set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place").

Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set model alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork.

Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them.

claim 01M3V7PKW97BMNCVHYZQ2MW1PY of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7PKW97BMNCVHYZQ2MW1PY --> **medium** — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (`grep -rn -i "effort\|xhigh"`). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (`sonnet` at `low`, `opus` at `xhigh`). For the most common spawn path it then says: "A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort." > > This repository ships no agent definitions. The search finds no `.claude/agents` files and no frontmatter that sets `low` or `xhigh`; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only `model` set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place"). > > Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set `model` alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork. > > Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them. claim `01M3V7PKW97BMNCVHYZQ2MW1PY` of review `01M3V7H8W09SMH6TCFF6K8FBMP`

medium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (grep -rn -i "effort\|xhigh"). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (sonnet at low, opus at xhigh). For the most common spawn path it then says: "A direct Agent tool call sets only the model; to give it the tier's effort, use an agent type whose definition sets that effort."

This repository ships no agent definitions. The search finds no .claude/agents files and no frontmatter that sets low or xhigh; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only model set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place").

Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set model alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork.

Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them.

claim 01M3V7PKW97BMNCVHYZQ2MW1PY of review 01M3V7H8W09SMH6TCFF6K8FBMP

<!-- review:claim:01M3V7PKW97BMNCVHYZQ2MW1PY --> **medium** — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > Examined: rules/claude/subagent-model-selection.md, the README layout and rule-delivery sections, and a repository-wide search for agent definitions and effort settings (`grep -rn -i "effort\|xhigh"`). The rule starts with an absolute: "Run every sub-agent at one of two settings", each defined as a model plus an effort (`sonnet` at `low`, `opus` at `xhigh`). For the most common spawn path it then says: "A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort." > > This repository ships no agent definitions. The search finds no `.claude/agents` files and no frontmatter that sets `low` or `xhigh`; the only matches are in the two new rules. So an agent following this rule in an ordinary checkout usually has no agent type to use, and the rule doesn't say what happens next. Plausible responses include: spawn with only `model` set and accept the inherited effort; stop and report a blocker; write a new agent definition file (a configuration change the user did not ask for); or avoid delegating. Some of these change files or skip work. The writing-for-agents skill says blockers must state what would unblock them and let independent work continue ("Specify the Discipline"), and says to keep an exception when several responses are plausible ("Make Every Line Earn Its Place"). > > Suggested fix: add the fallback in one clause, e.g. "...use an agent type whose definition sets that effort; when none exists, set `model` alone and accept the inherited effort." Use the fallback the author actually wants. This keeps the existing mechanism and removes the guesswork. > > Not verified: whether the author's own environment supplies such agent types through a private source or a settings fragment. If it does, the rule still reaches other users of this public repository without them. claim `01M3V7PKW97BMNCVHYZQ2MW1PY` of review `01M3V7H8W09SMH6TCFF6K8FBMP`
jercik marked this conversation as resolved
Author
Owner

Replying to review comment #98471

Round 6: per the review round gate, I'm not pushing these, and they stay acknowledged. The wrapper posted each claim twice.

  • #98471 and #98472: the rule already names the agent-type route for effort on a direct call. A fallback for a missing agent type is a refinement, not a defect.
  • #98475 and #98476 are real and small. In both rules, the second bullet should start "is the higher tier for work that needs original thought", to name the term the closing paragraph uses. I'll include that with the next owner-requested change.
> Replying to review comment #98471 Round 6: per the review round gate, I'm not pushing these, and they stay `acknowledged`. The wrapper posted each claim twice. - #98471 and #98472: the rule already names the agent-type route for effort on a direct call. A fallback for a missing agent type is a refinement, not a defect. - #98475 and #98476 are real and small. In both rules, the second bullet should start "is the higher tier for work that needs original thought", to name the term the closing paragraph uses. I'll include that with the next owner-requested change.
feat(rules): name scoped implementation and planning in the sub-agent tiers
Some checks failed
commit-msg / commitlint (pull_request) Failing after 22s
Node tests / node:test (pull_request) Successful in 2m7s
Review / Review (pull_request_target) Successful in 3m52s
17e17788f2
@ -0,0 +2,4 @@
Run every sub-agent at one of two settings:
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and implementing scoped changes.

low — Claude sub-agent rule says the sonnet alias is "Sonnet 5.5", a model that does not exist
lens general-bug · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the new rule rules/claude/subagent-model-selection.md, which is delivered to Claude Code sessions as user-level guidance for choosing sub-agent models. Its regular tier reads "sonnet (Sonnet 5.5) at low", and the higher tier reads "opus (Opus 5.5) at xhigh".

What goes wrong: the current Claude 5 family is Fable 5.1 (claude-fable-5-1), Opus 5.5 (claude-opus-5-5), and Sonnet 5 (claude-sonnet-5), plus Haiku 4.5. No Sonnet 5.5 exists. The Opus label is correct, but the sonnet alias resolves to Sonnet 5, not 5.5. An agent that follows the rule and pins a full model ID instead of the alias (for example in agent-definition frontmatter, where the rule says to set the model) could derive claude-sonnet-5-5 from this label, and that ID would not resolve. A reader also gets the wrong idea of which model the regular tier uses. The cost is small because the rule's main instruction is to use the sonnet alias, which works.

Evidence basis: I checked the label against the current Claude model list, which I know from my own environment. I did not run anything against the Claude API from this sandbox, which has no network access. To confirm or refute the claim, check whether claude-sonnet-5-5 appears in Anthropic's model list. If it does not, change the parenthetical to "Sonnet 5" or remove it.

claim 01M3VA3YDKATBHNCMCJV6WH9X3 of review 01M3VA1JHFYRT7H3A16SX436B9

<!-- review:claim:01M3VA3YDKATBHNCMCJV6WH9X3 --> **low** — Claude sub-agent rule says the `sonnet` alias is "Sonnet 5.5", a model that does not exist lens `general-bug` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the new rule `rules/claude/subagent-model-selection.md`, which is delivered to Claude Code sessions as user-level guidance for choosing sub-agent models. Its regular tier reads "`sonnet` (Sonnet 5.5) at `low`", and the higher tier reads "`opus` (Opus 5.5) at `xhigh`". > > What goes wrong: the current Claude 5 family is Fable 5.1 (`claude-fable-5-1`), Opus 5.5 (`claude-opus-5-5`), and Sonnet 5 (`claude-sonnet-5`), plus Haiku 4.5. No Sonnet 5.5 exists. The Opus label is correct, but the `sonnet` alias resolves to Sonnet 5, not 5.5. An agent that follows the rule and pins a full model ID instead of the alias (for example in agent-definition frontmatter, where the rule says to set the model) could derive `claude-sonnet-5-5` from this label, and that ID would not resolve. A reader also gets the wrong idea of which model the regular tier uses. The cost is small because the rule's main instruction is to use the `sonnet` alias, which works. > > Evidence basis: I checked the label against the current Claude model list, which I know from my own environment. I did not run anything against the Claude API from this sandbox, which has no network access. To confirm or refute the claim, check whether `claude-sonnet-5-5` appears in Anthropic's model list. If it does not, change the parenthetical to "Sonnet 5" or remove it. claim `01M3VA3YDKATBHNCMCJV6WH9X3` of review `01M3VA1JHFYRT7H3A16SX436B9`
Author
Owner

Repeat of #98337: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5 (Sonnet 5.5).

<!-- gh-feedback:reply-to:98680 --> Repeat of #98337: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5 (Sonnet 5.5).
@ -0,0 +5,4 @@
- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and implementing scoped changes.
- `opus` (Opus 5.5) at `xhigh` is the higher tier for work that needs original thought: plans, designs, architecture, reviews, opinions, root-cause diagnoses, and choosing between approaches.
Use the higher tier only when the deliverable is a judgment the caller will rely on. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort.

low — Tier rules key on a "judgment" deliverable while the skills that feed them signal "conclusion"/"facts", splitting one concept across two terms
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the trigger sentence shared by rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, and the three skill passages this change rewrote to feed it:

  • improve-codebase-architecture/SKILL.md: "Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions."
  • improve-codebase-architecture/INTERFACE-DESIGN.md: "Spawn 3+ sub-agents in parallel. Each delivers a design conclusion, not a fact list, ..."
  • review-agent-instructions/SKILL.md: "use a fresh reviewer, which delivers a conclusion ..."

What goes wrong: the rules decide the tier on "the deliverable is a judgment the caller will rely on". The skills describe deliverables as "conclusions" or "facts". The same skill also uses "judgment" for something else: "The judgment calls stay with you" means the main agent's own work, not a sub-agent deliverable. The writing guide says "use one term for one concept". A literal reader has to infer that "design conclusion" means "judgment" before the rule applies. In the architecture skill, the nearby "judgment calls stay with you" points away from that reading.

Proposed correction: pick one term and use it in both places. One option is to change the rules to "Use the higher tier only when the deliverable is a conclusion the caller will rely on — a design, review, diagnosis, or choice — not gathered facts." Then the skills' existing "conclusion"/"facts" wording maps onto the rule directly. The other option is to keep "judgment" and change the skills to say the sub-agent "delivers a judgment". Either way the tier split and its examples stay the same, and the rule-to-skill link becomes explicit.

claim 01M3VA56GWMZDAV9MR15BBV4Z0 of review 01M3VA1JHFYRT7H3A16SX436B9

<!-- review:claim:01M3VA56GWMZDAV9MR15BBV4Z0 --> **low** — Tier rules key on a "judgment" deliverable while the skills that feed them signal "conclusion"/"facts", splitting one concept across two terms lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the trigger sentence shared by `rules/claude/subagent-model-selection.md` and `rules/codex/subagent-model-selection.md`, and the three skill passages this change rewrote to feed it: > - improve-codebase-architecture/SKILL.md: "Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions." > - improve-codebase-architecture/INTERFACE-DESIGN.md: "Spawn 3+ sub-agents in parallel. Each delivers a design conclusion, not a fact list, ..." > - review-agent-instructions/SKILL.md: "use a fresh reviewer, which delivers a conclusion ..." > > What goes wrong: the rules decide the tier on "the deliverable is a judgment the caller will rely on". The skills describe deliverables as "conclusions" or "facts". The same skill also uses "judgment" for something else: "The judgment calls stay with you" means the main agent's own work, not a sub-agent deliverable. The writing guide says "use one term for one concept". A literal reader has to infer that "design conclusion" means "judgment" before the rule applies. In the architecture skill, the nearby "judgment calls stay with you" points away from that reading. > > Proposed correction: pick one term and use it in both places. One option is to change the rules to "Use the higher tier only when the deliverable is a conclusion the caller will rely on — a design, review, diagnosis, or choice — not gathered facts." Then the skills' existing "conclusion"/"facts" wording maps onto the rule directly. The other option is to keep "judgment" and change the skills to say the sub-agent "delivers a judgment". Either way the tier split and its examples stay the same, and the rule-to-skill link becomes explicit. claim `01M3VA56GWMZDAV9MR15BBV4Z0` of review `01M3VA1JHFYRT7H3A16SX436B9`
jercik marked this conversation as resolved
@ -32,3 +32,3 @@
## Judge independently
For a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:
For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:

medium — review-agent-instructions replaces explicit strongest-model guidance with an unexplained "delivers a conclusion" aside that only works with an agent-specific rule
lens writing-quality · arm default · tally 1 valid / 0 invalid / 0 uncertain

What I examined: the changed sentence in skills/review-agent-instructions/SKILL.md ("## Judge independently"), the new rules/claude/subagent-model-selection.md and rules/codex/subagent-model-selection.md, and the README.

What changed: the diff removed "use a fresh reviewer with the strongest available model at its deepest reasoning setting" and put in "use a fresh reviewer, which delivers a conclusion and need not be the model under test". The words "which delivers a conclusion" carry no instruction of their own. They only matter to a reader who has also loaded the new tier rule ("Use the higher tier only when the deliverable is a judgment the caller will rely on"). Those rules exist only under rules/claude/ and rules/codex/. The README says they are "selected by an agent override", and it lists Claude, Codex, Cursor, Gemini, OpenCode, and Grok as delivery targets.

What goes wrong: (1) Portability. The writing guide says "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have." An agent that runs this skill without the override rule (any of the other four harnesses, or Claude/Codex without the override selected) no longer gets any model guidance for the judge. The judge is the evidence-bearing step of this skill, so it may run at a default or cheap tier. (2) Clarity. Read alone, the relative clause looks like an odd description of the reviewer, not an instruction. That fails the no-op test for those readers. (3) Terminology. The rule keys on "judgment" and this skill says "conclusion", so a literal reader has to infer that the two words mean the same thing (see the companion finding on the rule).

Proposed correction: say the decision in portable terms that don't name a model, for example "For a clean-context review, use a fresh reviewer on your strongest available model tier, since its verdict is a judgment the caller relies on; it need not be the model under test." This keeps the change's goal of keeping model names out of the skill. It restores the instruction for readers without the rule and lines up with the rule's "judgment" trigger for readers who have it.

Proof gap: I couldn't check which rules a given installation selects. The claim rests on the README's statement that these rules are selected per agent and on the rule directories covering only two of the six listed harnesses.

claim 01M3VA560TT6X25SRN4RYZ1K93 of review 01M3VA1JHFYRT7H3A16SX436B9

<!-- review:claim:01M3VA560TT6X25SRN4RYZ1K93 --> **medium** — review-agent-instructions replaces explicit strongest-model guidance with an unexplained "delivers a conclusion" aside that only works with an agent-specific rule lens `writing-quality` · arm `default` · tally 1 valid / 0 invalid / 0 uncertain > What I examined: the changed sentence in `skills/review-agent-instructions/SKILL.md` ("## Judge independently"), the new `rules/claude/subagent-model-selection.md` and `rules/codex/subagent-model-selection.md`, and the README. > > What changed: the diff removed "use a fresh reviewer with the strongest available model at its deepest reasoning setting" and put in "use a fresh reviewer, which delivers a conclusion and need not be the model under test". The words "which delivers a conclusion" carry no instruction of their own. They only matter to a reader who has also loaded the new tier rule ("Use the higher tier only when the deliverable is a judgment the caller will rely on"). Those rules exist only under `rules/claude/` and `rules/codex/`. The README says they are "selected by an agent override", and it lists Claude, Codex, Cursor, Gemini, OpenCode, and Grok as delivery targets. > > What goes wrong: (1) Portability. The writing guide says "A skill is portable: repeat short required material from external rules or undeclared skills that readers may not have." An agent that runs this skill without the override rule (any of the other four harnesses, or Claude/Codex without the override selected) no longer gets any model guidance for the judge. The judge is the evidence-bearing step of this skill, so it may run at a default or cheap tier. (2) Clarity. Read alone, the relative clause looks like an odd description of the reviewer, not an instruction. That fails the no-op test for those readers. (3) Terminology. The rule keys on "judgment" and this skill says "conclusion", so a literal reader has to infer that the two words mean the same thing (see the companion finding on the rule). > > Proposed correction: say the decision in portable terms that don't name a model, for example "For a clean-context review, use a fresh reviewer on your strongest available model tier, since its verdict is a judgment the caller relies on; it need not be the model under test." This keeps the change's goal of keeping model names out of the skill. It restores the instruction for readers without the rule and lines up with the rule's "judgment" trigger for readers who have it. > > Proof gap: I couldn't check which rules a given installation selects. The claim rests on the README's statement that these rules are selected per agent and on the rule directories covering only two of the six listed harnesses. claim `01M3VA560TT6X25SRN4RYZ1K93` of review `01M3VA1JHFYRT7H3A16SX436B9`
Author
Owner

Repeat of #98336: the owner removed model guidance from skills on purpose; the per-harness rules own it.

<!-- gh-feedback:reply-to:98679 --> Repeat of #98336: the owner removed model guidance from skills on purpose; the per-harness rules own it.
Author
Owner

Replying to review comment #98681

Round 7: per the review round gate, I'm not pushing these, and they stay acknowledged.

  • #98681: the terms differ, but routing doesn't depend on them. Both rules list the task types the skills hand off: designs, exploring code, and reviews.
  • Unadjudicated claim 01M3VA4KZZ0AJDT0SQEP2FWQPV (no effort fallback when no agent type sets it) repeats #98471, acknowledged in round 6.
> Replying to review comment #98681 Round 7: per the review round gate, I'm not pushing these, and they stay `acknowledged`. - #98681: the terms differ, but routing doesn't depend on them. Both rules list the task types the skills hand off: designs, exploring code, and reviews. - Unadjudicated claim `01M3VA4KZZ0AJDT0SQEP2FWQPV` (no effort fallback when no agent type sets it) repeats #98471, acknowledged in round 6.
jercik merged commit e969bc2216 into main 2026-10-01 09:07:11 +00:00
jercik deleted branch feat/share-subagent-model-rules 2026-10-01 09:07:11 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
j4k-oss/agent-skills!91
No description provided.