feat(rules): share the Claude and Codex sub-agent model rules #91
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/share-subagent-model-rules"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Moves the Claude and Codex sub-agent model rules here from
j4k/setup-atlasand reduces each to two tiers:sonnetatlowgpt-6.1-solatlowopusatxhighgpt-6-astraathighRule IDs are unchanged, so existing selections need no edit.
Skills now leave model and effort to these rules. Three sites in
improve-codebase-architectureandreview-agent-instructionssay only what the sub-agent delivers and drop Claude Code-specific tool names (#92).j4k/setup-atlas#134 already deleted the private copies. Until this merges, a refreshed setup-atlas checkout selects rules that no source provides, and
axskillsrefuses to run.Review
01M3VA1JHFYRT7H3A16SX436B9— head17e17788f269a2a307d94a1f33e7636bc17d7b7eReview — j4k-oss/agent-skills @
9bc616862dScope: diff against base tree
87531f305d3aStatus: dispatched — coverage complete (3/3 slots terminal)
Facts: current review-wide projection
Computed under:
Findings (3)
medium — review-agent-instructions replaces explicit strongest-model guidance with an unexplained "delivers a conclusion" aside that only works with an agent-specific rule
01M3VA560TT6X25SRN4RYZ1K93skills/review-agent-instructions/SKILL.md(snippet)low — Claude sub-agent rule says the
sonnetalias is "Sonnet 5.5", a model that does not exist01M3VA3YDKATBHNCMCJV6WH9X3rules/claude/subagent-model-selection.md(snippet)01M3VA5H888VW2937FP658QPEE(writing-quality)low — Tier rules key on a "judgment" deliverable while the skills that feed them signal "conclusion"/"facts", splitting one concept across two terms
01M3VA56GWMZDAV9MR15BBV4Z0rules/claude/subagent-model-selection.md(snippet)Other claims
01M3VA5H888VW2937FP658QPEElow — Claude tier rule pins thesonnet/opusaliases to version numbers that will go stale, and "Sonnet 5.5" doesn't match the current model lineup →01M3VA3YDKATBHNCMCJV6WH9X301M3VA4KZZ0AJDT0SQEP2FWQPVmedium — Claude model rule says every sub-agent must use one of two settings but has no fallback when no agent type sets the needed effortCoverage
Coverage pass: 01M3VA1JJXFAM3C57BK4Z5PFEF
Accounting: complete
Slot health: healthy
@ -0,0 +4,4 @@- `gpt-6-luna` at `high` for mechanical work with a clear output and little judgment: running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, applying a specified bulk edit. The small model needs `high` even for mechanical work.- `gpt-6-sol` at `medium` for work that takes trial and error to finish: implementing a scoped change, getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.- `gpt-6-astra` at `high` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.medium — The Astra effort rule conflicts with the interface-design skill maximum-effort instruction
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3TYZ7M0FFRN5CGE1P6BX2N7of review01M3TYMYHQ04WGZ7WBT0Y2MJHNFixed in
9ac36d692a. INTERFACE-DESIGN.md now says the designers deliver a conclusion and defers model and effort to the agent's own instructions, falling back to the strongest model at maximum effort. The same fix covers unadjudicated claim 01M3TYWHPH0PRZFA07N67M0VSR (the Explore walk in improve-codebase-architecture/SKILL.md) and the identical pattern in review-agent-instructions/SKILL.md.@ -0,0 +6,4 @@- `gpt-6-sol` at `medium` for work that takes trial and error to finish: implementing a scoped change, getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.- `gpt-6-astra` at `high` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.Route a task that sits between two tiers to the higher one, and move a task that needs more reasoning up a tier rather than raising its effort. To set `model` and `reasoning_effort`, use `fork_turns: "none"` or a positive integer string and pass the child the task context it needs. A full-history fork inherits the parent model and effort and rejects overrides; use it when preserving that context matters more than selecting a different tier.low — Model escalation rule has no action above the highest tier
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3TZSY4HTZWX7TN0JYCTM7CSof review01M3TZKC5QYK787RHC0KAVA8XFThe gap is real: nothing says what to do when a task already on gpt-6-astra needs more reasoning. Which way it goes is the rule owner's policy call, so I've asked @jercik: keep astra at high as the ceiling (split the task or report the limit), or allow xhigh at the top tier only?
superseded by review
01M3V3MV82K6PQQ2G16FGY1VBTfor head1cd99e7d9154990d729c5486a52ac97f460aeb98Fixed in
86027ddd86. The owner chose gpt-6-astra at high as the top tier. The rule now calls it the ceiling and says to split a task that needs more or report the limit. The 'up a tier rather than raising its effort' sentence is gone.@ -0,0 +6,4 @@- `gpt-6-sol` at `medium` for work that takes trial and error to finish: implementing a scoped change, getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.- `gpt-6-astra` at `high` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.Route a task that sits between two tiers to the higher one, and move a task that needs more reasoning up a tier rather than raising its effort. To set `model` and `reasoning_effort`, use `fork_turns: "none"` or a positive integer string and pass the child the task context it needs. A full-history fork inherits the parent model and effort and rejects overrides; use it when preserving that context matters more than selecting a different tier.low — Define an escalation path for tasks already on Astra
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V3SQAC4S12K4B5YKJX091Aof review01M3V3MV82K6PQQ2G16FGY1VBTsuperseded by review
01M3V4BP30R44N7E58JT3RJKRMfor head86027ddd866fe78aa3c8bac18a3002df63831cadFixed in
86027ddd86: gpt-6-astra at high is now stated as the ceiling, with split-or-report beyond it.@ -39,3 +39,3 @@Follow its loading procedure: a root `CONTEXT-MAP.md` means the repo has multiple contexts — read the touched contexts' `CONTEXT.md` files, the root `docs/adr/` (system-wide decisions), and each touched context's own `docs/adr/` — otherwise the root `CONTEXT.md` and the ADRs in the area you're touching.Then use the Agent tool with `subagent_type=Explore` — a fast executor model at moderate reasoning effort, since the walk delivers facts, not conclusions — to walk the codebase. The sub-agent starts with none of your context, so the brief must carry everything the walk needs: the friction questions below, the glossary terms the project docs define, and the [LANGUAGE.md](LANGUAGE.md) definitions of **shallow**, **seam**, and **locality**, with the instruction to report findings as `file:line` evidence in that vocabulary. Don't impose rigid heuristics — have it explore organically and note friction:Then spawn a read-only exploration sub-agent to walk the codebase. The walk delivers facts, not conclusions. The sub-agent starts with none of your context, so the brief must carry everything the walk needs: the friction questions below, the glossary terms the project docs define, and the [LANGUAGE.md](LANGUAGE.md) definitions of **shallow**, **seam**, and **locality**, with the instruction to report findings as `file:line` evidence in that vocabulary. Don't impose rigid heuristics — have it explore organically and note friction:medium — Align the exploration brief with its fact-only boundary
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V3T90DPGJGDRKAK1Q56PMNof review01M3V3MV82K6PQQ2G16FGY1VBTlow — Qualify the claim that every explorer starts without context
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V3WDCDTZSFW026ET633CXAof review01M3V3MV82K6PQQ2G16FGY1VBTsuperseded by review
01M3V4BP30R44N7E58JT3RJKRMfor head86027ddd866fe78aa3c8bac18a3002df63831cad@ -0,0 +2,4 @@When choosing a sub-agent's model explicitly, use one of three settings:- `gpt-6.1-sol` at `low` is the regular tier, and most sub-agents run on it: mechanical work with a clear output and little judgment, such as running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, or applying a specified bulk edit.medium — Regular sub-agent tier names a model absent from the available Codex overrides
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V4FE3C18FAP6RZAN8JBYWFof review01M3V4BP30R44N7E58JT3RJKRMgpt-6.1-sol is available. Codex 0.159.2's model catalog, fetched 2026-10-01T07:08Z in a local session (models_cache.json), lists gpt-6.1-sol with visibility 'list' and reasoning levels low through ultra, beside gpt-6-astra. The review environment's spawn_agent list predates that model.
superseded by review
01M3V67FK226HY7EDBZ5F16S4Tfor head4bfca904c46127293ec67581ebc0cbb7856cbc10@ -0,0 +2,4 @@Launch every sub-agent at one of three settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier, and most sub-agents run on it: routine work with a clear goal, such as running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, applying a specified bulk edit, or implementing a scoped change.low — Claude sub-agent rule says the
sonnetalias is "Sonnet 5.5", a model version the alias does not resolve tolens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V6A8V3D15B0H225V83XGNMof review01M3V67FK226HY7EDBZ5F16S4TSonnet 5.5 exists and is what the alias resolves to. The installed Claude Code 2.1.286 binary maps the aliases as sonnet: claude-sonnet-5-5 and opus: claude-opus-5-5, and its 2.1.284 changelog reads 'Added Claude Sonnet 5.5 (claude-sonnet-5-5), now the default Sonnet model on the Anthropic API'.
superseded by review
01M3V6JPN0S0M7BTX1S93K2BAGfor head83f5156c0a43b2bf85c398c980aa7456c1c33fc1@ -0,0 +6,4 @@- `opus` (Opus 5.5) at `medium` for work that takes trial and error to finish: getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.- `opus` at `xhigh` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set the model and effort explicitly wherever they can be set: workflow `agent()` calls and agent definition frontmatter.medium — Claude model-selection rule requires a setting on "every sub-agent" but omits the Agent tool, where effort cannot be set
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V69XR8W2X9ASA3YPEZK2BAof review01M3V67FK226HY7EDBZ5F16S4TFixed in
83f5156c0a. The rule now says workflow agent() calls and agent definition frontmatter set both, and a direct Agent tool call sets only the model, with effort from the agent type's definition.@ -47,3 +47,3 @@- Which parts of the codebase are untested, or hard to test through their current interface?The judgment calls stay with you: from the returned facts, decide what is shallow and apply the **deletion test** to anything you suspect — would deleting it concentrate complexity, or just move it? A "yes, concentrates" is the signal you want.The judgment calls stay with you: from the returned facts, decide what is shallow and apply the **deletion test** to anything you suspect. Complexity that vanishes marks a pass-through, the candidate you want; complexity that reappears across callers means the module is earning its keep.low — Explore step re-defines the deletion test that the glossary already defines a few paragraphs earlier
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V6AEK1SYCHEPFXER5AZZF6of review01M3V67FK226HY7EDBZ5F16S4TFixed in
83f5156c0awith the proposed wording; the glossary keeps the only definition.@ -32,3 +32,3 @@## Judge independentlyFor a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:medium — review-agent-instructions drops its reviewer-strength requirement and leaves only an implicit cue for an optional, agent-specific rule
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V6A850BGN12969VJ4DHC8Kof review01M3V67FK226HY7EDBZ5F16S4TThis is deliberate. The repository owner decided that skills name no model, effort, or tier, and that model choice belongs to the per-harness rules. Without a selected rule, the harness default picks the reviewer. 'Top reasoning tier' is still a model-strength instruction, which is exactly what the owner removed from skills.
superseded by review
01M3V6JPN0S0M7BTX1S93K2BAGfor head83f5156c0a43b2bf85c398c980aa7456c1c33fc1@ -0,0 +2,4 @@Launch every sub-agent at one of three settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier, and most sub-agents run on it: routine work with a clear goal, such as running a given command and reporting its output, finding every call site of a symbol, extracting fields from files, applying a specified bulk edit, or implementing a scoped change.low — Claude sub-agent rule names a nonexistent "Sonnet 5.5" as the model behind the
sonnetaliaslens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V6MMA3WSG7PBB2QPCFZH0Nof review01M3V6JPN0S0M7BTX1S93K2BAGlow — Tier bullets break parallel form, and "most sub-agents run on it" repeats the closing default sentence (both model-selection rules)
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V6QDBE4638D64J876D64ETof review01M3V6JPN0S0M7BTX1S93K2BAGSame claim as #98337, already refuted: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Claude Sonnet 5.5 as the default Sonnet model.
superseded by review
01M3V78FSFQ0HV1WGSCMRF0T04for headd7530a344b372b4d9ee54c3201341d9d40c7cb93@ -0,0 +6,4 @@- `opus` (Opus 5.5) at `medium` for work that takes trial and error to finish: getting a failing test suite green, browsing and scraping pages, collecting evidence against given criteria.- `opus` at `xhigh` for work whose deliverable is a conclusion: reviewing work, synthesizing findings into a report, diagnosing a root cause, choosing an approach, weighing tradeoffs, designing a plan.Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both explicitly in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; its effort comes from the agent type's definition.medium — Claude model-selection rule requires a model+effort pair for every sub-agent, then admits the Agent tool cannot set effort, without saying what to do
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V6NXGTF7797DV8N86DMAVTof review01M3V6JPN0S0M7BTX1S93K2BAG@ -32,3 +32,3 @@## Judge independentlyFor a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:low — "a fresh reviewer, which delivers a conclusion" hides a model-tier signal in a descriptive clause that changes no behavior for a reader without the tier rule
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V6PG5MQH0Z4XKF332THC64of review01M3V6JPN0S0M7BTX1S93K2BAGSame point as #98336: the owner removed model-strength guidance from skills on purpose and left model choice to the per-harness rules.
superseded by review
01M3V78FSFQ0HV1WGSCMRF0T04for headd7530a344b372b4d9ee54c3201341d9d40c7cb93Round 4: per the review round gate, I'm not pushing these, and both stay
acknowledged.rules/claude/subagent-model-selection.md, append: "To give a direct call its tier's effort, use an agent type whose definition sets that effort."@ -0,0 +2,4 @@Launch every sub-agent at one of three settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier, and most sub-agents run on it: exploring and searching code, researching how a codebase works, running commands, watching for changes such as a CI run or deploy, and routine edits up to a scoped change.low — Claude sub-agent rule names a nonexistent "Sonnet 5.5" as the model behind the
sonnetaliaslens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7B2483A4CDZP4KKAJ1AV9of review01M3V78FSFQ0HV1WGSCMRF0T04Repeat of #98337: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5 (Sonnet 5.5).
superseded by review
01M3V7H8W09SMH6TCFF6K8FBMPfor head714d376d368b6b4b51f1c135bafbe3bf9e4fd14bsuperseded by review
01M3V7H8W09SMH6TCFF6K8FBMPfor head714d376d368b6b4b51f1c135bafbe3bf9e4fd14b@ -0,0 +6,4 @@- `opus` (Opus 5.5) at `medium` for work that takes trial and error to finish, such as getting a failing test suite green or driving a multi-step browser flow.- `opus` at `xhigh` for work whose deliverable is a conclusion: reviews, diagnoses, designs, and decisions between approaches.Stay on the regular tier unless the task clearly needs trial and error or delivers a conclusion. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both explicitly in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; its effort comes from the agent type's definition.medium — Claude sub-agent rule mandates a model+effort setting for every sub-agent, then says direct Agent calls cannot set effort, with no instruction for that case
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7B3VF5RE7YFAJ2B8VG0JZof review01M3V78FSFQ0HV1WGSCMRF0T04Fixed in
714d376d36: a direct Agent tool call gets the tier's effort through an agent type whose definition sets it.@ -32,3 +32,3 @@## Judge independentlyFor a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:medium — Skills swap explicit model/effort guidance for coded "delivers a conclusion" / "facts, not conclusions" hints that only mean something to readers with the Claude/Codex sub-agent rule
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7BP6VTCA5C7S6KW35WQ98of review01M3V78FSFQ0HV1WGSCMRF0T04Repeat of #98336: the owner removed model guidance from skills on purpose; the per-harness rules own it.
superseded by review
01M3V7H8W09SMH6TCFF6K8FBMPfor head714d376d368b6b4b51f1c135bafbe3bf9e4fd14bsuperseded by review
01M3V7H8W09SMH6TCFF6K8FBMPfor head714d376d368b6b4b51f1c135bafbe3bf9e4fd14b@ -0,0 +2,4 @@Run every sub-agent at one of two settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.low — Claude sub-agent rule says the
sonnetalias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnetlens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7KWGWXB9CJ1291P0CYN6Kof review01M3V7H8W09SMH6TCFF6K8FBMPlow — Claude sub-agent rule says the
sonnetalias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnetlens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7KWGWXB9CJ1291P0CYN6Kof review01M3V7H8W09SMH6TCFF6K8FBMPRepeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.
Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.
superseded by review
01M3VA1JHFYRT7H3A16SX436B9for head17e17788f269a2a307d94a1f33e7636bc17d7b7e@ -0,0 +3,4 @@Run every sub-agent at one of two settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.low — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7Q640RTJ7EP4J40EPHPT2of review01M3V7H8W09SMH6TCFF6K8FBMPlow — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7Q640RTJ7EP4J40EPHPT2of review01M3V7H8W09SMH6TCFF6K8FBMP@ -0,0 +5,4 @@- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.Use the higher tier only when the deliverable is a judgment the caller will rely on. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort.medium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7PKW97BMNCVHYZQ2MW1PYof review01M3V7H8W09SMH6TCFF6K8FBMPmedium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7PKW97BMNCVHYZQ2MW1PYof review01M3V7H8W09SMH6TCFF6K8FBMP@ -0,0 +2,4 @@Run every sub-agent at one of two settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.low — Claude sub-agent rule says the
sonnetalias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnetlens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7KWGWXB9CJ1291P0CYN6Kof review01M3V7H8W09SMH6TCFF6K8FBMPlow — Claude sub-agent rule says the
sonnetalias is Sonnet 5.5, but the Claude Code harness lists Sonnet 5 as the current Sonnetlens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7KWGWXB9CJ1291P0CYN6Kof review01M3V7H8W09SMH6TCFF6K8FBMPRepeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.
Repeat of #98337: the installed Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5, and its 2.1.284 changelog adds Sonnet 5.5 as the default Sonnet. The review harness's model list is older.
superseded by review
01M3VA1JHFYRT7H3A16SX436B9for head17e17788f269a2a307d94a1f33e7636bc17d7b7e@ -0,0 +3,4 @@Run every sub-agent at one of two settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.low — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7Q640RTJ7EP4J40EPHPT2of review01M3V7H8W09SMH6TCFF6K8FBMPlow — Second tier bullet has no verb or name, so the later "higher tier" term is never defined
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7Q640RTJ7EP4J40EPHPT2of review01M3V7H8W09SMH6TCFF6K8FBMP@ -0,0 +5,4 @@- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and edits up to a scoped change.- `opus` (Opus 5.5) at `xhigh` for work that needs original thought: reviews, opinions, root-cause diagnoses, designs, architecture, and choosing between approaches.Use the higher tier only when the deliverable is a judgment the caller will rely on. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort.medium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7PKW97BMNCVHYZQ2MW1PYof review01M3V7H8W09SMH6TCFF6K8FBMPmedium — Claude rule requires every sub-agent to run at a fixed effort but gives no fallback when no agent type sets that effort
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3V7PKW97BMNCVHYZQ2MW1PYof review01M3V7H8W09SMH6TCFF6K8FBMPRound 6: per the review round gate, I'm not pushing these, and they stay
acknowledged. The wrapper posted each claim twice.@ -0,0 +2,4 @@Run every sub-agent at one of two settings:- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and implementing scoped changes.low — Claude sub-agent rule says the
sonnetalias is "Sonnet 5.5", a model that does not existlens
general-bug· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3VA3YDKATBHNCMCJV6WH9X3of review01M3VA1JHFYRT7H3A16SX436B9Repeat of #98337: Claude Code 2.1.286 maps sonnet to claude-sonnet-5-5 (Sonnet 5.5).
@ -0,0 +5,4 @@- `sonnet` (Sonnet 5.5) at `low` is the regular tier for most work: exploring and searching code, researching how a codebase or tool works, running commands and test suites, getting a failing test suite green, watching for changes such as a CI run or deploy, browsing and scraping pages, collecting evidence against given criteria, and implementing scoped changes.- `opus` (Opus 5.5) at `xhigh` is the higher tier for work that needs original thought: plans, designs, architecture, reviews, opinions, root-cause diagnoses, and choosing between approaches.Use the higher tier only when the deliverable is a judgment the caller will rely on. `opus` at `xhigh` is the ceiling: split a task that needs more, or report the limit. Set both in workflow `agent()` calls and agent definition frontmatter. A direct Agent tool call sets only the `model`; to give it the tier's effort, use an agent type whose definition sets that effort.low — Tier rules key on a "judgment" deliverable while the skills that feed them signal "conclusion"/"facts", splitting one concept across two terms
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3VA56GWMZDAV9MR15BBV4Z0of review01M3VA1JHFYRT7H3A16SX436B9@ -32,3 +32,3 @@## Judge independentlyFor a clean-context review, use a fresh reviewer with the strongest available model at its deepest reasoning setting. This reviewer need not be the model under test. Give it, by path and reading order:For a clean-context review, use a fresh reviewer, which delivers a conclusion and need not be the model under test. Give it, by path and reading order:medium — review-agent-instructions replaces explicit strongest-model guidance with an unexplained "delivers a conclusion" aside that only works with an agent-specific rule
lens
writing-quality· armdefault· tally 1 valid / 0 invalid / 0 uncertainclaim
01M3VA560TT6X25SRN4RYZ1K93of review01M3VA1JHFYRT7H3A16SX436B9Repeat of #98336: the owner removed model guidance from skills on purpose; the per-harness rules own it.
Round 7: per the review round gate, I'm not pushing these, and they stay
acknowledged.01M3VA4KZZ0AJDT0SQEP2FWQPV(no effort fallback when no agent type sets it) repeats #98471, acknowledged in round 6.