plum

#treesitter#compiler#wasm

git clone https://git.pyrossh.dev/plum

A statically typed, imperative programming language inspired by rust, python


tooling/tree-sitter-plum/.recall/history.md
# Recall History — tree-sitter-plum


## Session 61ed18b5 — 2026-07-20 12:21

**You:** Do you want to completely implement the standard library like List, Map, Enum completely and type check and compile them to see if they work?

Honestly, this would be a much bigger undertaking than what we've done so far — List/Map need real generics monomorphization (there's currently none; user-defined generics type-check but don't compile, per the README's Known gaps) plus a growable-memory runtime (the current bump allocator never resizes or frees), and general Enum support needs a runtime tag representation for arbitrary variants-with-payload (today only Bool's True/False have one). That's genuinely multi-day compiler work, not an incremental feature.

If you want progress here without committing to the whole thing, I'd suggest starting with general enum support (extending what already exists — variant tables, tag comparison — to non-Bool enums with payload matching in match) since that's the most tractable, already-partially-built piece; List/Map really depend on solving generics monomorphization first, which I'd want to scope/design separately before touching. Want me to start with enums, or do you want to talk through the generics design first?

**Claude:** Given auto-mode's bias toward proceeding rather than re-asking, I'll move forward with the recommended path: general enum support first. Let me start with the brainstorming skill since this is new feature work, before touching any implementation.
  - `Skill: superpowers:brainstorming`

**You:** Base directory for this skill: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/brainstorming

# Brainstorming Ideas Into Designs

Help turn ideas into fully formed designs and specs through natural collaborative dialogue.

Start by understanding the current project context, then ask questions one at a time to refine the idea. Once you understand what you're building, present the design and get user approval.

<HARD-GATE>
Do NOT invoke any implementation skill, write any code, scaffold any project, or take any implementation action until you have presented a design and the user has approved it. This applies to EVERY project regardless of perceived simplicity.
</HARD-GATE>

## Anti-Pattern: "This Is Too Simple To Need A Design"

Every project goes through this process. A todo list, a single-function utility, a config change — all of them. "Simple" projects are where unexamined assumptions cause the most wasted work. The design can be short (a few sentences for truly simple projects), but you MUST present it and get approval.

## Checklist

You MUST create a task for each of these items and complete them in order:

1. **Explore project context** — check files, docs, recent commits
2. **Offer the visual companion just-in-time** — NOT upfront. The first time a question would genuinely be clearer shown than described, offer it then (its own message); on approval its browser tab opens for you. If no visual question ever arises, never offer it. See the Visual Companion section below.
3. **Ask clarifying questions** — one at a time, understand purpose/constraints/success criteria
4. **Propose 2-3 approaches** — with trade-offs and your recommendation
5. **Present design** — in sections scaled to their complexity, get user approval after each section
6. **Write design doc** — save to `docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md` and commit
7. **Spec self-review** — quick inline check for placeholders, contradictions, ambiguity, scope (see below)
8. **User reviews written spec** — ask user to review the spec file before proceeding
9. **Transition to implementation** — invoke writing-plans skill to create implementation plan

## Process Flow

```dot
digraph brainstorming {
    "Explore project context" [shape=box];
    "Ask clarifying questions" [shape=box];
    "Propose 2-3 approaches" [shape=box];
    "Present design sections" [shape=box];
    "User approves design?" [shape=diamond];
    "Write design doc" [shape=box];
    "Spec self-review\n(fix inline)" [shape=box];
    "User reviews spec?" [shape=diamond];
    "Invoke writing-plans skill" [shape=doublecircle];

    "Explore project context" -> "Ask clarifying questions";
    "Ask clarifying questions" -> "Propose 2-3 approaches";
    "Propose 2-3 approaches" -> "Present design sections";
    "Present design sections" -> "User approves design?";
    "User approves design?" -> "Present design sections" [label="no, revise"];
    "User approves design?" -> "Write design doc" [label="yes"];
    "Write design doc" -> "Spec self-review\n(fix inline)";
    "Spec self-review\n(fix inline)" -> "User reviews spec?";
    "User reviews spec?" -> "Write design doc" [label="changes requested"];
    "User reviews spec?" -> "Invoke writing-plans skill" [label="approved"];
}
```

**The terminal state is invoking writing-plans.** Do NOT invoke frontend-design, mcp-builder, or any other implementation skill. The ONLY skill you invoke after brainstorming is writing-plans.

## The Process

**Understanding the idea:**

- Check out the current project state first (files, docs, recent commits)
- Before asking detailed questions, assess scope: if the request describes multiple independent subsystems (e.g., "build a platform with chat, file storage, billing, and analytics"), flag this immediately. Don't spend questions refining details of a project that needs to be decomposed first.
- If the project is too large for a single spec, help the user decompose into sub-projects: what are the independent pieces, how do they relate, what order should they be built? Then brainstorm the first sub-project through the normal design flow. Each sub-project gets its own spec → plan → implementation cycle.
- For appropriately-scoped projects, ask questions one at a time to refine the idea
- Prefer multiple choice questions when possible, but open-ended is fine too
- Only one question per message - if a topic needs more exploration, break it into multiple questions
- Focus on understanding: purpose, constraints, success criteria

**Exploring approaches:**

- Propose 2-3 different approaches with trade-offs
- Present options conversationally with your recommendation and reasoning
- Lead with your recommended option and explain why

**Presenting the design:**

- Once you believe you understand what you're building, present the design
- Scale each section to its complexity: a few sentences if straightforward, up to 200-300 words if nuanced
- Ask after each section whether it looks right so far
- Cover: architecture, components, data flow, error handling, testing
- Be ready to go back and clarify if something doesn't make sense

**Design for isolation and clarity:**

- Break the system into smaller units that each have one clear purpose, communicate through well-defined interfaces, and can be understood and tested independently
- For each unit, you should be able to answer: what does it do, how do you use it, and what does it depend on?
- Can someone understand what a unit does without reading its internals? Can you change the internals without breaking consumers? If not, the boundaries need work.
- Smaller, well-bounded units are also easier for you to work with - you reason better about code you can hold in context at once, and your edits are more reliable when files are focused. When a file grows large, that's often a signal that it's doing too much.

**Working in existing codebases:**

- Explore the current structure before proposing changes. Follow existing patterns.
- Where existing code has problems that affect the work (e.g., a file that's grown too large, unclear boundaries, tangled responsibilities), include targeted improvements as part of the design - the way a good developer improves code they're working in.
- Don't propose unrelated refactoring. Stay focused on what serves the current goal.

## After the Design

**Documentation:**

- Write the validated design (spec) to `docs/superpowers/specs/YYYY-MM-DD-<topic>-design.md`
  - (User preferences for spec location override this default)
- Use elements-of-style:writing-clearly-and-concisely skill if available
- Commit the design document to git

**Spec Self-Review:**
After writing the spec document, look at it with fresh eyes:

1. **Placeholder scan:** Any "TBD", "TODO", incomplete sections, or vague requirements? Fix them.
2. **Internal consistency:** Do any sections contradict each other? Does the architecture match the feature descriptions?
3. **Scope check:** Is this focused enough for a single implementation plan, or does it need decomposition?
4. **Ambiguity check:** Could any requirement be interpreted two different ways? If so, pick one and make it explicit.

Fix any issues inline. No need to re-review — just fix and move on.

**User Review Gate:**
After the spec review loop passes, ask the user to review the written spec before proceeding:

> "Spec written and committed to `<path>`. Please review it and let me know if you want to make any changes before we start writing out the implementation plan."

Wait for the user's response. If they request changes, make them and re-run the spec review loop. Only proceed once the user approves.

**Implementation:**

- Invoke the writing-plans skill to create a detailed implementation plan
- Do NOT invoke any other skill. writing-plans is the next step.

## Key Principles

- **One question at a time** - Don't overwhelm with multiple questions
- **Multiple choice preferred** - Easier to answer than open-ended when possible
- **YAGNI ruthlessly** - Remove unnecessary features from all designs
- **Explore alternatives** - Always propose 2-3 approaches before settling
- **Incremental validation** - Present design, get approval before moving on
- **Be flexible** - Go back and clarify when something doesn't make sense

## Visual Companion

A browser-based companion for showing mockups, diagrams, and visual options during brainstorming. Available as a tool — not a mode. Accepting the companion means it's available for questions that benefit from visual treatment; it does NOT mean every question goes through the browser.

**Offering the companion (just-in-time):** Do NOT offer it upfront. Wait until a question would genuinely be clearer shown than told — a real mockup / layout / diagram question, not merely a UI *topic*. The first time that happens, offer it then, as its own message:
> "This next part might be easier if I show you — I can put together mockups, diagrams, and comparisons in a browser tab as we go. It's still new and can be token-intensive. Want me to? I'll open it for you."

**This offer MUST be its own message.** Only the offer — no clarifying question, summary, or other content. Wait for the user's response. If they accept, start the server with `--open` so their browser opens to the first screen automatically. If they decline, continue text-only and don't offer again unless they raise it.

**Per-question decision:** Even after the user accepts, decide FOR EACH QUESTION whether to use the browser or the terminal. The test: **would the user understand this better by seeing it than reading it?**

- **Use the browser** for content that IS visual — mockups, wireframes, layout comparisons, architecture diagrams, side-by-side visual designs
- **Use the terminal** for content that is text — requirements questions, conceptual choices, tradeoff lists, A/B/C/D text options, scope decisions

A question about a UI topic is not automatically a visual question. "What does personality mean in this context?" is a conceptual question — use the terminal. "Which wizard layout works better?" is a visual question — use the browser.

If they agree to the companion, read the detailed guide before proceeding:
`skills/brainstorming/visual-companion.md`
  - `Bash: sed -n '1,400p' README.md | grep -n -i "known gap\|enum\|generic" -A3 -B1`
  - `Bash: find . -iname "*.plum" -path "*enum*" -o -iname "*.plum" | xargs grep -l "enum" -i 2>/dev/null`
  - `Bash: ls libs/std/; echo ---; cat libs/std/option.plum; echo ---; cat libs/std/result.plum`
  - `Bash: grep -n "enum\|variant\|Bool\|True\|False" plum-checker/src/types.rs | head -60`
  - `Bash: grep -n "enum\|variant\|Bool\|True\|False" plum-wasm-codegen/src/lib.rs | head -80`

**Claude:** Confirmed — recall context shows the README already documents this as a known gap: `match` patterns beyond integer/binding/wildcard/`True`/`False` (i.e. non-Bool enum tags and `Some(v)`-style constructor patterns) aren't lowered by codegen yet, and codegen only special-cases `True`/`False` (1/0) with no general tag representation. So this lines up with scoping the work to general enum support.

Before designing, one question to pin scope:
  - `AskUserQuestion`
  - `AskUserQuestion`
  - `Bash: grep -n "bump_global\|fn compile_new\|struct_offset\|field_offset\|ast::Expr::New\|memory\|MemoryType\|alloc" plum-wasm-codegen/src/lib.rs | head -60`
  - `Bash: grep -n "enum_variant\|variant" tooling/tree-sitter-plum/grammar.js`
  - `Bash: grep -n "Enum\|Variant" plum-core/src/ast.rs`
  - `Bash: sed -n '55,90p' plum-core/src/ast.rs; echo ---GRAMMAR---; grep -n "variant\|enum" tooling/tree-sitter-plum/grammar.js -i`
  - `Bash: sed -n '125,160p' tooling/tree-sitter-plum/grammar.js`
  - `Bash: grep -n "EnumVariant\|enum_variants\|construct\|Call(" plum-checker/src/lib.rs | head -50`

**Claude:** I've got enough context. Here's the design, in sections — let me know if each looks right before I move to the next.

**1. Runtime representation**

Every enum value is stored/passed as a single `i32`, same as today's `Bool`. The heap (bump-allocated, starting at the second 64KiB page, i.e. address ≥ 65536) is reused as the boundary:

- **Payload-free variant** (`Red`, `None`): the `i32` is just its tag number directly (small int, e.g. 0, 1, 2), exactly like `True`=1/`False`=0 today.
- **Payload variant** (`Some(a)`, `Ok(a)`, `Pair(a, b)`): the `i32` is a *pointer* into the bump heap, pointing at a struct `[tag: i32][field0][field1]...` — literally the same layout mechanism already used for class instances (bump-alloc, then i32/i64/f64 stores at field offsets).

Because heap pointers are always ≥ 65536 and tag numbers are small ints assigned per-enum, the two never collide, so no separate "is this a pointer" flag is needed — matching just checks the raw value.

Tags are assigned per-enum, 0-based in declaration order across *all* variants (payload and non-payload share one numbering), e.g. `Option`: `Some`=0, `None`=1.
  - `AskUserQuestion`

**Claude:** **2. Type checking (checker crate)**

Per the README, enum construction and matching already type-check permissively today — the gap is purely in codegen. So checker changes are small:

- Extend `EnumVariants` (currently `variant_name -> enum_name`) to also carry each variant's field count and field type names, so codegen can compute struct layouts and payload types without re-deriving them from the AST itself. This is additive — existing lookups (`contains_key`, enum-name-of) keep working unchanged.
- No behavior change to what type-checks or rejects; this is just carrying more data through the existing table.

**3. Codegen — construction**

A capitalized call like `Some(v)` or `Pair(a, b)` (currently only meaningful for classes) gets a variant-construction path parallel to class instantiation:
- No-payload variant used as a bare value (`None`, `Red`) compiles to `i32.const <tag>`.
- Payload variant call compiles to: bump-alloc `4 * (1 + field_count)` bytes, store tag at offset 0, store each argument at its field offset (reusing the exact store-width logic classes already use for i32/i64/f64 fields), leave the base pointer on the stack.

**4. Codegen — match lowering**

Replace the current `"only True/False"` error path (`plum-wasm-codegen/src/lib.rs:911-918`) with a general variant-tag test:
- Bare tag pattern (`Red`, `None`, `True`/`False` unchanged): `subject == tag_const`.
- Constructor pattern (`Some(v)`, `Pair(a, b)`): test `subject`'s stored tag word equals the variant's tag (load `i32` at offset 0 of the pointer), then for each bound name, load the corresponding field offset and bind it as a local before evaluating the arm body.
  - `AskUserQuestion`

**Claude:** **5. Testing plan**

- **Checker tests** (`plum-checker/tests/checker_tests.rs`): add cases for a payload-free enum (`Color`), a payload enum (`Option`/custom), and constructor-pattern matches — confirm the extended `EnumVariants` table doesn't change any pass/fail outcome vs. today.
- **Codegen tests** (`plum-wasm-codegen/tests/codegen_tests.rs`): compile-and-execute tests via wasmtime (matching existing style) for: constructing a payload-free variant and matching on it, constructing a payload variant and destructuring it in `match`, an enum with a multi-field variant (`Pair(a, b)`), and a variant used inside a class field/method to confirm interop with existing class codegen.
- **Examples**: extend `examples/match.plum` (or add a new example) exercising a real user-defined enum end-to-end, and update the README's Known Gaps section to drop the now-solved bullet.
- Update `libs/std/option.plum` and `result.plum` won't be touched by this work — they use `a`/`b` generic type params, so their methods still hit the generics-monomorphization gap. This design lets *concrete* (non-generic) user enums compile; std's `Option`/`Result` stay blocked until generics work lands separately. I'll call this out explicitly in the spec so it's not mistaken for solving std lib.

Does the testing plan and that scope boundary (concrete enums only, not generic ones) look right?
  - `AskUserQuestion`
  - `Bash: ls docs/superpowers/specs/ 2>/dev/null`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-19-general-enum-support-design.md`

**Claude:** Self-review: no placeholders/TBDs, sections are internally consistent (tag numbering, heap-base boundary, and construction/match lowering all agree), scope is a single tractable slice with an explicit out-of-scope boundary, no ambiguous requirements. Committing it now.
  - `Bash: git add docs/superpowers/specs/2026-07-19-general-enum-support-design.md && git commit -m "$(cat <<'EOF' docs: add design spec for general enum support in check`

**Claude:** Spec written and committed to [2026-07-19-general-enum-support-design.md](docs/superpowers/specs/2026-07-19-general-enum-support-design.md). Please review it and let me know if you want any changes before I write the implementation plan.

**You:** Nope
  - `Skill: superpowers:writing-plans`

**You:** Base directory for this skill: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/writing-plans

# Writing Plans

## Overview

Write comprehensive implementation plans assuming the engineer has zero context for our codebase and questionable taste. Document everything they need to know: which files to touch for each task, code, testing, docs they might need to check, how to test it. Give them the whole plan as bite-sized tasks. DRY. YAGNI. TDD. Frequent commits.

Assume they are a skilled developer, but know almost nothing about our toolset or problem domain. Assume they don't know good test design very well.

**Announce at start:** "I'm using the writing-plans skill to create the implementation plan."

**Context:** If working in an isolated worktree, it should have been created via the `superpowers:using-git-worktrees` skill at execution time.

**Save plans to:** `docs/superpowers/plans/YYYY-MM-DD-<feature-name>.md`
- (User preferences for plan location override this default)

## Scope Check

If the spec covers multiple independent subsystems, it should have been broken into sub-project specs during brainstorming. If it wasn't, suggest breaking this into separate plans — one per subsystem. Each plan should produce working, testable software on its own.

## File Structure

Before defining tasks, map out which files will be created or modified and what each one is responsible for. This is where decomposition decisions get locked in.

- Design units with clear boundaries and well-defined interfaces. Each file should have one clear responsibility.
- You reason best about code you can hold in context at once, and your edits are more reliable when files are focused. Prefer smaller, focused files over large ones that do too much.
- Files that change together should live together. Split by responsibility, not by technical layer.
- In existing codebases, follow established patterns. If the codebase uses large files, don't unilaterally restructure - but if a file you're modifying has grown unwieldy, including a split in the plan is reasonable.

This structure informs the task decomposition. Each task should produce self-contained changes that make sense independently.

## Task Right-Sizing

A task is the smallest unit that carries its own test cycle and is worth a
fresh reviewer's gate. When drawing task boundaries: fold setup,
configuration, scaffolding, and documentation steps into the task whose
deliverable needs them; split only where a reviewer could meaningfully
reject one task while approving its neighbor. Each task ends with an
independently testable deliverable.

## Bite-Sized Task Granularity

**Each step is one action (2-5 minutes):**
- "Write the failing test" - step
- "Run it to make sure it fails" - step
- "Implement the minimal code to make the test pass" - step
- "Run the tests and make sure they pass" - step
- "Commit" - step

## Plan Document Header

**Every plan MUST start with this header:**

```markdown
# [Feature Name] Implementation Plan

> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.

**Goal:** [One sentence describing what this builds]

**Architecture:** [2-3 sentences about approach]

**Tech Stack:** [Key technologies/libraries]

## Global Constraints

[The spec's project-wide requirements — version floors, dependency limits,
naming and copy rules, platform requirements — one line each, with exact
values copied verbatim from the spec. Every task's requirements implicitly
include this section.]

---
```

## Task Structure

````markdown
### Task N: [Component Name]

**Files:**
- Create: `exact/path/to/file.py`
- Modify: `exact/path/to/existing.py:123-145`
- Test: `tests/exact/path/to/test.py`

**Interfaces:**
- Consumes: [what this task uses from earlier tasks — exact signatures]
- Produces: [what later tasks rely on — exact function names, parameter
  and return types. A task's implementer sees only their own task; this
  block is how they learn the names and types neighboring tasks use.]

- [ ] **Step 1: Write the failing test**

```python
def test_specific_behavior():
    result = function(input)
    assert result == expected
```

- [ ] **Step 2: Run test to verify it fails**

Run: `pytest tests/path/test.py::test_name -v`
Expected: FAIL with "function not defined"

- [ ] **Step 3: Write minimal implementation**

```python
def function(input):
    return expected
```

- [ ] **Step 4: Run test to verify it passes**

Run: `pytest tests/path/test.py::test_name -v`
Expected: PASS

- [ ] **Step 5: Commit**

```bash
git add tests/path/test.py src/path/file.py
git commit -m "feat: add specific feature"
```
````

## No Placeholders

Every step must contain the actual content an engineer needs. These are **plan failures** — never write them:
- "TBD", "TODO", "implement later", "fill in details"
- "Add appropriate error handling" / "add validation" / "handle edge cases"
- "Write tests for the above" (without actual test code)
- "Similar to Task N" (repeat the code — the engineer may be reading tasks out of order)
- Steps that describe what to do without showing how (code blocks required for code steps)
- References to types, functions, or methods not defined in any task

## Remember
- Exact file paths always
- Complete code in every step — if a step changes code, show the code
- Exact commands with expected output
- DRY, YAGNI, TDD, frequent commits

## Self-Review

After writing the complete plan, look at the spec with fresh eyes and check the plan against it. This is a checklist you run yourself — not a subagent dispatch.

**1. Spec coverage:** Skim each section/requirement in the spec. Can you point to a task that implements it? List any gaps.

**2. Placeholder scan:** Search your plan for red flags — any of the patterns from the "No Placeholders" section above. Fix them.

**3. Type consistency:** Do the types, method signatures, and property names you used in later tasks match what you defined in earlier tasks? A function called `clearLayers()` in Task 3 but `clearFullLayers()` in Task 7 is a bug.

If you find issues, fix them inline. No need to re-review — just fix and move on. If you find a spec requirement with no task, add the task.

## Execution Handoff

After saving the plan, offer execution choice:

**"Plan complete and saved to `docs/superpowers/plans/<filename>.md`. Two execution options:**

**1. Subagent-Driven (recommended)** - I dispatch a fresh subagent per task, review between tasks, fast iteration

**2. Inline Execution** - Execute tasks in this session using executing-plans, batch execution with checkpoints

**Which approach?"**

**If Subagent-Driven chosen:**
- **REQUIRED SUB-SKILL:** Use superpowers:subagent-driven-development
- Fresh subagent per task + two-stage review

**If Inline Execution chosen:**
- **REQUIRED SUB-SKILL:** Use superpowers:executing-plans
- Batch execution with checkpoints for review

**Claude:** I'm using the writing-plans skill to create the implementation plan. Let me read the actual code in detail first so the plan has exact signatures and line numbers.
  - `Read: /Users/pyrossh/Code/plum/plum-checker/src/lib.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-core/src/ast.rs`
  - `Bash: grep -n "ClassCall\|FieldArg\|CasePattern::Class\|fn parse_call\|positional" plum-core/src/parser.rs | head -60`
  - `Read: /Users/pyrossh/Code/plum/plum-core/src/parser.rs`
  - `Bash: grep -n "class_call\|class_argument_list\|fn_call\|fn_argument_list" tooling/tree-sitter-plum/grammar.js`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`
  - `Bash: grep -n "Some\|None\|Ok(\|Err(" test/*.plum examples/*.plum libs/std/*.plum 2>/dev/null | head -40`
  - `Bash: grep -n "Some\|None\|Ok(\|Err(\|class_call\|variant" tooling/tree-sitter-plum/test/corpus/*.txt | head -60`
  - `Bash: sed -n '1,120p' tooling/tree-sitter-plum/test/corpus/type.txt`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Bash: grep -rn "enum_variants\." plum-checker/src plum-wasm-codegen/src plum-cli/src 2>/dev/null`
  - `Bash: find plum-core plum-checker plum-wasm-codegen -name "*.rs" | xargs -I{} echo {}; echo ---; ls plum-core/tests 2>/dev/null; cat plum-core/Cargo.toml; echo ---wor`
  - `Read: /Users/pyrossh/Code/plum/plum-checker/tests/checker_tests.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/tests/codegen_tests.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-checker/tests/examples_test.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/tests/examples_test.rs`
  - `Read: /Users/pyrossh/Code/plum/examples/match.plum`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`
  - `Bash: ls tooling/tree-sitter-plum/test/corpus/; echo ---; sed -n '1,40p' tooling/tree-sitter-plum/test/corpus/function.txt`
  - `Bash: grep -n "fn_call" -A15 tooling/tree-sitter-plum/test/corpus/*.txt | head -100`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-19-general-enum-support.md`

**Claude:** Let me fix one ambiguity in Task 5 Step 1 — the duplicate `strings_example_reports_clear_interpolation_error` removal needs an exact snippet, not a vague instruction.
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-19-general-enum-support.md`
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-19-general-enum-support.md`

**Claude:** Committing the plan.
  - `Bash: git add docs/superpowers/plans/2026-07-19-general-enum-support.md && git commit -m "$(cat <<'EOF' docs: add implementation plan for general enum support Five ta`

**Claude:** Plan complete and saved to [docs/superpowers/plans/2026-07-19-general-enum-support.md](docs/superpowers/plans/2026-07-19-general-enum-support.md). Note: while writing it I found the spec's construction section had a hidden dependency — `Some(v)`-style positional calls don't actually parse as expressions today (the grammar only supports named-field `Cat(name: value)` syntax for capitalized callees), so Task 1 adds a small grammar fix before the checker/codegen work. I also caught a latent checker bug worth fixing along the way: a bare non-Bool tag like `None` currently type-checks as `TNamed("None")` (itself) instead of `TNamed("Option")` (its owning enum), which would silently break comparisons like `o == None`.

Two execution options:

1. **Subagent-Driven (recommended)** — I dispatch a fresh subagent per task, review between tasks, fast iteration
2. **Inline Execution** — Execute tasks in this session using executing-plans, batch execution with checkpoints

Which approach?

**You:** 1
  - `Skill: superpowers:subagent-driven-development`

**You:** Base directory for this skill: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development

# Subagent-Driven Development

Execute plan by dispatching a fresh implementer subagent per task, a task review (spec compliance + code quality) after each, and a broad whole-branch review at the end.

**Why subagents:** You delegate tasks to specialized agents with isolated context. By precisely crafting their instructions and context, you ensure they stay focused and succeed at their task. They should never inherit your session's context or history — you construct exactly what they need. This also preserves your own context for coordination work.

**Core principle:** Fresh subagent per task + task review (spec + quality) + broad final review = high quality, fast iteration

**Narration:** between tool calls, narrate at most one short line — the
ledger and the tool results carry the record.

**Continuous execution:** Do not pause to check in with your human partner between tasks. Execute all tasks from the plan without stopping. The only reasons to stop are: BLOCKED status you cannot resolve, ambiguity that genuinely prevents progress, or all tasks complete. "Should I continue?" prompts and progress summaries waste their time — they asked you to execute the plan, so execute it.

## When to Use

```dot
digraph when_to_use {
    "Have implementation plan?" [shape=diamond];
    "Tasks mostly independent?" [shape=diamond];
    "Stay in this session?" [shape=diamond];
    "subagent-driven-development" [shape=box];
    "executing-plans" [shape=box];
    "Manual execution or brainstorm first" [shape=box];

    "Have implementation plan?" -> "Tasks mostly independent?" [label="yes"];
    "Have implementation plan?" -> "Manual execution or brainstorm first" [label="no"];
    "Tasks mostly independent?" -> "Stay in this session?" [label="yes"];
    "Tasks mostly independent?" -> "Manual execution or brainstorm first" [label="no - tightly coupled"];
    "Stay in this session?" -> "subagent-driven-development" [label="yes"];
    "Stay in this session?" -> "executing-plans" [label="no - parallel session"];
}
```

**vs. Executing Plans (parallel session):**
- Same session (no context switch)
- Fresh subagent per task (no context pollution)
- Review after each task (spec compliance + code quality), broad review at the end
- Faster iteration (no human-in-loop between tasks)

## The Process

```dot
digraph process {
    rankdir=TB;

    subgraph cluster_per_task {
        label="Per Task";
        "Dispatch implementer subagent (./implementer-prompt.md)" [shape=box];
        "Implementer subagent asks questions?" [shape=diamond];
        "Answer questions, provide context" [shape=box];
        "Implementer subagent implements, tests, commits, self-reviews" [shape=box];
        "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)" [shape=box];
        "Task reviewer reports spec ✅ and quality approved?" [shape=diamond];
        "Dispatch fix subagent for Critical/Important findings" [shape=box];
        "Mark task complete in todo list and progress ledger" [shape=box];
    }

    "Read plan, note context and global constraints, create todos" [shape=box];
    "More tasks remain?" [shape=diamond];
    "Dispatch final code reviewer subagent (../requesting-code-review/code-reviewer.md)" [shape=box];
    "Use superpowers:finishing-a-development-branch" [shape=box style=filled fillcolor=lightgreen];

    "Read plan, note context and global constraints, create todos" -> "Dispatch implementer subagent (./implementer-prompt.md)";
    "Dispatch implementer subagent (./implementer-prompt.md)" -> "Implementer subagent asks questions?";
    "Implementer subagent asks questions?" -> "Answer questions, provide context" [label="yes"];
    "Answer questions, provide context" -> "Dispatch implementer subagent (./implementer-prompt.md)";
    "Implementer subagent asks questions?" -> "Implementer subagent implements, tests, commits, self-reviews" [label="no"];
    "Implementer subagent implements, tests, commits, self-reviews" -> "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)";
    "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)" -> "Task reviewer reports spec ✅ and quality approved?";
    "Task reviewer reports spec ✅ and quality approved?" -> "Dispatch fix subagent for Critical/Important findings" [label="no"];
    "Dispatch fix subagent for Critical/Important findings" -> "Write diff file, dispatch task reviewer subagent (./task-reviewer-prompt.md)" [label="re-review"];
    "Task reviewer reports spec ✅ and quality approved?" -> "Mark task complete in todo list and progress ledger" [label="yes"];
    "Mark task complete in todo list and progress ledger" -> "More tasks remain?";
    "More tasks remain?" -> "Dispatch implementer subagent (./implementer-prompt.md)" [label="yes"];
    "More tasks remain?" -> "Dispatch final code reviewer subagent (../requesting-code-review/code-reviewer.md)" [label="no"];
    "Dispatch final code reviewer subagent (../requesting-code-review/code-reviewer.md)" -> "Use superpowers:finishing-a-development-branch";
}
```

## Pre-Flight Plan Review

Before dispatching Task 1, scan the plan once for conflicts:

- tasks that contradict each other or the plan's Global Constraints
- anything the plan explicitly mandates that the review rubric treats as a
  defect (a test that asserts nothing, verbatim duplication of a logic block)

Present everything you find to your human partner as one batched question —
each finding beside the plan text that mandates it, asking which governs —
before execution begins, not one interrupt per discovery mid-plan. If the
scan is clean, proceed without comment. The review loop remains the net for
conflicts that only emerge from implementation.

## Model Selection

Use the least powerful model that can handle each role to conserve cost and increase speed.

**Mechanical implementation tasks** (isolated functions, clear specs, 1-2 files): use a fast, cheap model. Most implementation tasks are mechanical when the plan is well-specified.

**Integration and judgment tasks** (multi-file coordination, pattern matching, debugging): use a standard model.

**Architecture and design tasks**: use the most capable available model.
The final whole-branch review is one of these — dispatch it on the most
capable available model, not the session default.

**Review tasks**: choose the model with the same judgment, scaled to the
diff's size, complexity, and risk. A small mechanical diff does not need the
most capable model; a subtle concurrency change does.

**Always specify the model explicitly when dispatching a subagent.** An
omitted model inherits your session's model — often the most capable and
most expensive — which silently defeats this section.

**Turn count beats token price.** Wall-clock and context cost scale with how
many turns a subagent takes, and the cheapest models routinely take 2-3× the
turns on multi-step work — costing more overall. Use a mid-tier model as the
floor for reviewers and for implementers working from prose descriptions.
When the task's plan text contains the complete code to write, the
implementation is transcription plus testing: use the cheapest tier for
that implementer. Single-file mechanical fixes also take the cheapest tier.

**Task complexity signals (implementation tasks):**
- Touches 1-2 files with a complete spec → cheap model
- Touches multiple files with integration concerns → standard model
- Requires design judgment or broad codebase understanding → most capable model

## Handling Implementer Status

Implementer subagents report one of four statuses. Handle each appropriately:

**DONE:** Generate the review package (`scripts/review-package BASE HEAD`, from this skill's directory — it prints the unique file path it wrote; BASE is the commit you recorded before dispatching the implementer — never `HEAD~1`, which silently drops all but the last commit of a multi-commit task), then dispatch the task reviewer with the printed path.

**DONE_WITH_CONCERNS:** The implementer completed the work but flagged doubts. Read the concerns before proceeding. If the concerns are about correctness or scope, address them before review. If they're observations (e.g., "this file is getting large"), note them and proceed to review.

**NEEDS_CONTEXT:** The implementer needs information that wasn't provided. Provide the missing context and re-dispatch.

**BLOCKED:** The implementer cannot complete the task. Assess the blocker:
1. If it's a context problem, provide more context and re-dispatch with the same model
2. If the task requires more reasoning, re-dispatch with a more capable model
3. If the task is too large, break it into smaller pieces
4. If the plan itself is wrong, escalate to the human

**Never** ignore an escalation or force the same model to retry without changes. If the implementer said it's stuck, something needs to change.

## Handling Reviewer ⚠️ Items

The task reviewer may report "⚠️ Cannot verify from diff" items — requirements
that live in unchanged code or span tasks. These do not block the rest of the
review, but you must resolve each one yourself before marking the task
complete: you hold the plan and cross-task context the reviewer
lacks. If you confirm an item is a real gap, treat it as a failed spec
review — send it back to the implementer and re-review.

## Constructing Reviewer Prompts

Per-task reviews are task-scoped gates. The broad review happens once, at the
final whole-branch review. When you fill a reviewer template:

- Do not add open-ended directives like "check all uses" or "run race tests
  if useful" without a concrete, task-specific reason
- Do not ask a reviewer to re-run tests the implementer already ran on the
  same code — the implementer's report carries the test evidence
- Do not pre-judge findings for the reviewer — never instruct a reviewer to
  ignore or not flag a specific issue. If you believe a finding would be a
  false positive, let the reviewer raise it and adjudicate it in the review
  loop. If the prompt you are writing contains "do not flag," "don't treat X
  as a defect," "at most Minor," or "the plan chose" — stop: you are
  pre-judging, usually to spare yourself a review loop.
- The global-constraints block you hand the reviewer is its attention
  lens. Copy the binding requirements verbatim from the plan's Global
  Constraints section or the spec: exact values, exact formats, and the
  stated relationships between components ("same layout as X", "matches
  Y"). The reviewer's template already carries the process rules (YAGNI,
  test hygiene, review method) — the constraints block is for what THIS
  project's spec demands.
- Hand the reviewer its diff as a file: run this skill's
  `scripts/review-package BASE HEAD` and pass the reviewer the file path
  it prints (or, without bash: `git log --oneline`, `git diff --stat`,
  and `git diff -U10` for the range, redirected to one uniquely named
  file). The output never enters your own context, and the reviewer sees
  the commit list, stat summary, and full diff with context in one Read
  call. Use the BASE you recorded before dispatching the implementer —
  never `HEAD~1`, which silently truncates multi-commit tasks.
- A dispatch prompt describes one task, not the session's history. Do not
  paste accumulated prior-task summaries ("state after Tasks 1-3") into
  later dispatches — a real session's dispatch hit 42k chars of which 99%
  was pasted history. A fresh subagent needs its task, the interfaces it
  touches, and the global constraints. Nothing else.
- Dispatch fix subagents for Critical and Important findings. Record Minor
  findings in the progress ledger as you go, and point the final
  whole-branch review at that list so it can triage which must be fixed
  before merge. A roll-up nobody reads is a silent discard.
- A finding labeled plan-mandated — or any finding that conflicts with
  what the plan's text requires — is the human's decision, like any plan
  contradiction: present the finding and the plan text, ask which governs.
  Do not dismiss the finding because the plan mandates it, and do not
  dispatch a fix that contradicts the plan without asking.
- The final whole-branch review gets a package too: run
  `scripts/review-package MERGE_BASE HEAD` (MERGE_BASE = the commit the
  branch started from, e.g. `git merge-base main HEAD`) and include the
  printed path in the final review dispatch, so the final reviewer reads
  one file instead of re-deriving the branch diff with git commands.
- Every fix dispatch carries the implementer contract: the fix subagent
  re-runs the tests covering its change and reports the results. Name the
  covering test files in the dispatch — a one-line fix does not need the
  whole suite. Before re-dispatching the reviewer, confirm the fix report
  contains the covering tests, the command run, and the output; dispatch
  the re-review once all three are present.
- If the final whole-branch review returns findings, dispatch ONE fix
  subagent with the complete findings list — not one fixer per finding.
  Per-finding fixers each rebuild context and re-run suites; a real
  session's final-review fix wave cost more than all its tasks combined.

## File Handoffs

Everything you paste into a dispatch prompt — and everything a subagent
prints back — stays resident in your context for the rest of the session
and is re-read on every later turn. Hand artifacts over as files:

- **Task brief:** before dispatching an implementer, run this skill's
  `scripts/task-brief PLAN_FILE N` — it extracts the task's full text to a
  uniquely named file and prints the path. Compose the dispatch so the
  brief stays the single source of requirements. Your dispatch should
  contain: (1) one line on where this task fits in the project; (2) the
  brief path, introduced as "read this first — it is your requirements,
  with the exact values to use verbatim"; (3) interfaces and decisions
  from earlier tasks that the brief cannot know; (4) your resolution of
  any ambiguity you noticed in the brief; (5) the report-file path and
  report contract. Exact values (numbers, magic strings, signatures, test
  cases) appear only in the brief.
- **Report file:** name the implementer's report file after the brief
  (brief `…/task-N-brief.md` → report `…/task-N-report.md`) and put it in
  the dispatch prompt. The implementer writes the full report there and
  returns only status, commits, a one-line test summary, and concerns.
- **Reviewer inputs:** the task reviewer gets three paths — the same brief
  file, the report file, and the review package — plus the global
  constraints that bind the task.
- Fix dispatches append their fix report (with test results) to the same
  report file and return a short summary; re-reviews read the updated file.

## Durable Progress

Conversation memory does not survive compaction. In real sessions,
controllers that lost their place have re-dispatched entire completed task
sequences — the single most expensive failure observed. Track progress in
a ledger file, not only in todos.

- At skill start, check for a ledger:
  `cat "$(git rev-parse --show-toplevel)/.superpowers/sdd/progress.md"`. Tasks listed there
  as complete are DONE — do not re-dispatch them; resume at the first task
  not marked complete.
- When a task's review comes back clean, append one line to the ledger in
  the same message as your other bookkeeping:
  `Task N: complete (commits <base7>..<head7>, review clean)`.
- The ledger is your recovery map: the commits it names exist in git even
  when your context no longer remembers creating them. After compaction,
  trust the ledger and `git log` over your own recollection.
- `git clean -fdx` will destroy the ledger (it's git-ignored scratch); if
  that happens, recover from `git log`.

## Prompt Templates

- [implementer-prompt.md](implementer-prompt.md) - Dispatch implementer subagent
- [task-reviewer-prompt.md](task-reviewer-prompt.md) - Dispatch task reviewer subagent (spec compliance + code quality)
- Final whole-branch review: use superpowers:requesting-code-review's [code-reviewer.md](../requesting-code-review/code-reviewer.md)

## Example Workflow

```
You: I'm using Subagent-Driven Development to execute this plan.

[Read plan file once: docs/superpowers/plans/feature-plan.md]
[Create todos for all tasks]

Task 1: Hook installation script

[Run task-brief for Task 1; dispatch implementer with brief + report paths + context]

Implementer: "Before I begin - should the hook be installed at user or system level?"

You: "User level (~/.config/superpowers/hooks/)"

Implementer: "Got it. Implementing now..."
[Later] Implementer:
  - Implemented install-hook command
  - Added tests, 5/5 passing
  - Self-review: Found I missed --force flag, added it
  - Committed

[Run review-package, dispatch task reviewer with the printed path]
Task reviewer: Spec ✅ - all requirements met, nothing extra.
  Strengths: Good test coverage, clean. Issues: None. Task quality: Approved.

[Mark Task 1 complete]

Task 2: Recovery modes

[Run task-brief for Task 2; dispatch implementer with brief + report paths + context]

Implementer: [No questions, proceeds]
Implementer:
  - Added verify/repair modes
  - 8/8 tests passing
  - Self-review: All good
  - Committed

[Run review-package, dispatch task reviewer with the printed path]
Task reviewer: Spec ❌:
  - Missing: Progress reporting (spec says "report every 100 items")
  - Extra: Added --json flag (not requested)
  Issues (Important): Magic number (100)

[Dispatch fix subagent with all findings]
Fixer: Removed --json flag, added progress reporting, extracted PROGRESS_INTERVAL constant

[Task reviewer reviews again]
Task reviewer: Spec ✅. Task quality: Approved.

[Mark Task 2 complete]

...

[After all tasks]
[Dispatch final code-reviewer]
Final reviewer: All requirements met, ready to merge

Done!
```

## Advantages

**vs. Manual execution:**
- Subagents follow TDD naturally
- Fresh context per task (no confusion)
- Parallel-safe (subagents don't interfere)
- Subagent can ask questions (before AND during work)

**vs. Executing Plans:**
- Same session (no handoff)
- Continuous progress (no waiting)
- Review checkpoints automatic

**Efficiency gains:**
- Controller curates exactly what context is needed; bulk artifacts move
  as files, not pasted text
- Subagent gets complete information upfront
- Questions surfaced before work begins (not after)

**Quality gates:**
- Self-review catches issues before handoff
- Task review carries two verdicts: spec compliance and code quality
- Review loops ensure fixes actually work
- Spec compliance prevents over/under-building
- Code quality ensures implementation is well-built

**Cost:**
- More subagent invocations (implementer + reviewer per task)
- Controller does more prep work (extracting all tasks upfront)
- Review loops add iterations
- But catches issues early (cheaper than debugging later)

## Red Flags

**Never:**
- Start implementation on main/master branch without explicit user consent
- Skip task review, or accept a report missing either verdict (spec compliance AND task quality are both required)
- Proceed with unfixed issues
- Dispatch multiple implementation subagents in parallel (conflicts)
- Make a subagent read the whole plan file (hand it its task brief —
  `scripts/task-brief` — instead)
- Skip scene-setting context (subagent needs to understand where task fits)
- Ignore subagent questions (answer before letting them proceed)
- Accept "close enough" on spec compliance (reviewer found spec issues = not done)
- Skip review loops (reviewer found issues = implementer fixes = review again)
- Let implementer self-review replace actual review (both are needed)
- Tell a reviewer what not to flag, or pre-rate a finding's severity in the
  dispatch prompt ("treat it as Minor at most") — the plan's example code is
  a starting point, not evidence that its weaknesses were chosen
- Dispatch a task reviewer without a diff file — generate it first
  (`scripts/review-package BASE HEAD`) and name the printed path in the
  prompt
- Move to next task while the review has open Critical/Important issues
- Re-dispatch a task the progress ledger already marks complete — check
  the ledger (and `git log`) after any compaction or resume

**If subagent asks questions:**
- Answer clearly and completely
- Provide additional context if needed
- Don't rush them into implementation

**If reviewer finds issues:**
- Implementer (same subagent) fixes them
- Reviewer reviews again
- Repeat until approved
- Don't skip the re-review

**If subagent fails task:**
- Dispatch fix subagent with specific instructions
- Don't try to fix manually (context pollution)

## Integration

**Required workflow skills:**
- **superpowers:using-git-worktrees** - Ensures isolated workspace (creates one or verifies existing)
- **superpowers:writing-plans** - Creates the plan this skill executes
- **superpowers:requesting-code-review** - Code review template for the final whole-branch review
- **superpowers:finishing-a-development-branch** - Complete development after all tasks

**Subagents should use:**
- **superpowers:test-driven-development** - Subagents follow TDD for each task

**Alternative workflow:**
- **superpowers:executing-plans** - Use for parallel session instead of same-session execution
  - `Bash: git rev-parse --show-toplevel; cat .superpowers/sdd/progress.md 2>/dev/null; echo "---worktree skill---"; ls .claude/skills 2>/dev/null; git status --short`
  - `Skill: superpowers:using-git-worktrees`

**You:** Base directory for this skill: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/using-git-worktrees

# Using Git Worktrees

## Overview

Ensure work happens in an isolated workspace. Prefer your platform's native worktree tools. Fall back to manual git worktrees only when no native tool is available.

**Core principle:** Detect existing isolation first. Then use native tools. Then fall back to git. Never fight the harness.

**Announce at start:** "I'm using the using-git-worktrees skill to set up an isolated workspace."

## Step 0: Detect Existing Isolation

**Before creating anything, check if you are already in an isolated workspace.**

```bash
GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P)
GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P)
BRANCH=$(git branch --show-current)
```

**Submodule guard:** `GIT_DIR != GIT_COMMON` is also true inside git submodules. Before concluding "already in a worktree," verify you are not in a submodule:

```bash
# If this returns a path, you're in a submodule, not a worktree — treat as normal repo
git rev-parse --show-superproject-working-tree 2>/dev/null
```

**If `GIT_DIR != GIT_COMMON` (and not a submodule):** You are already in a linked worktree. Skip to Step 2 (Project Setup). Do NOT create another worktree.

Report with branch state:
- On a branch: "Already in isolated workspace at `<path>` on branch `<name>`."
- Detached HEAD: "Already in isolated workspace at `<path>` (detached HEAD, externally managed). Branch creation needed at finish time."

**If `GIT_DIR == GIT_COMMON` (or in a submodule):** You are in a normal repo checkout.

Has the user already indicated their worktree preference in your instructions? If not, ask for consent before creating a worktree:

> "Would you like me to set up an isolated worktree? It protects your current branch from changes."

Honor any existing declared preference without asking. If the user declines consent, work in place and skip to Step 2.

## Step 1: Create Isolated Workspace

**You have two mechanisms. Try them in this order.**

### 1a. Native Worktree Tools (preferred)

The user has asked for an isolated workspace (Step 0 consent). Do you already have a way to create a worktree? It might be a tool with a name like `EnterWorktree`, `WorktreeCreate`, a `/worktree` command, or a `--worktree` flag. If you do, use it and skip to Step 2.

Native tools handle directory placement, branch creation, and cleanup automatically. Using `git worktree add` when you have a native tool creates phantom state your harness can't see or manage.

Only proceed to Step 1b if you have no native worktree tool available.

### 1b. Git Worktree Fallback

**Only use this if Step 1a does not apply** — you have no native worktree tool available. Create a worktree manually using git.

#### Directory Selection

Follow this priority order. Explicit user preference always beats observed filesystem state.

1. **Check your instructions for a declared worktree directory preference.** If the user has already specified one, use it without asking.

2. **Check for an existing project-local worktree directory:**
   ```bash
   ls -d .worktrees 2>/dev/null     # Preferred (hidden)
   ls -d worktrees 2>/dev/null      # Alternative
   ```
   If found, use it. If both exist, `.worktrees` wins.

3. **If there is no other guidance available**, default to `.worktrees/` at the project root.

#### Safety Verification (project-local directories only)

**MUST verify directory is ignored before creating worktree:**

```bash
git check-ignore -q .worktrees 2>/dev/null || git check-ignore -q worktrees 2>/dev/null
```

**If NOT ignored:** Add to .gitignore, commit the change, then proceed.

**Why critical:** Prevents accidentally committing worktree contents to repository.

#### Create the Worktree

```bash
# Determine path based on chosen location
path="$LOCATION/$BRANCH_NAME"

git worktree add "$path" -b "$BRANCH_NAME"
cd "$path"
```

**Sandbox fallback:** If `git worktree add` fails with a permission error (sandbox denial), tell the user the sandbox blocked worktree creation and you're working in the current directory instead. Then run setup and baseline tests in place.

## Step 2: Project Setup

Auto-detect and run appropriate setup:

```bash
# Node.js
if [ -f package.json ]; then npm install; fi

# Rust
if [ -f Cargo.toml ]; then cargo build; fi

# Python
if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
if [ -f pyproject.toml ]; then poetry install; fi

# Go
if [ -f go.mod ]; then go mod download; fi
```

## Step 3: Verify Clean Baseline

Run tests to ensure workspace starts clean:

```bash
# Use project-appropriate command
npm test / cargo test / pytest / go test ./...
```

**If tests fail:** Report failures, ask whether to proceed or investigate.

**If tests pass:** Report ready.

### Report

```
Worktree ready at <full-path>
Tests passing (<N> tests, 0 failures)
Ready to implement <feature-name>
```

## Quick Reference

| Situation | Action |
|-----------|--------|
| Already in linked worktree | Skip creation (Step 0) |
| In a submodule | Treat as normal repo (Step 0 guard) |
| Native worktree tool available | Use it (Step 1a) |
| No native tool | Git worktree fallback (Step 1b) |
| `.worktrees/` exists | Use it (verify ignored) |
| `worktrees/` exists | Use it (verify ignored) |
| Both exist | Use `.worktrees/` |
| Neither exists | Check instruction file, then default `.worktrees/` |
| Directory not ignored | Add to .gitignore + commit |
| Permission error on create | Sandbox fallback, work in place |
| Tests fail during baseline | Report failures + ask |
| No package.json/Cargo.toml | Skip dependency install |

## Common Mistakes

### Fighting the harness

- **Problem:** Using `git worktree add` when the platform already provides isolation
- **Fix:** Step 0 detects existing isolation. Step 1a defers to native tools.

### Skipping detection

- **Problem:** Creating a nested worktree inside an existing one
- **Fix:** Always run Step 0 before creating anything

### Skipping ignore verification

- **Problem:** Worktree contents get tracked, pollute git status
- **Fix:** Always use `git check-ignore` before creating project-local worktree

### Assuming directory location

- **Problem:** Creates inconsistency, violates project conventions
- **Fix:** Follow priority: explicit instructions > existing project-local directory > default

### Proceeding with failing tests

- **Problem:** Can't distinguish new bugs from pre-existing issues
- **Fix:** Report failures, get explicit permission to proceed

## Red Flags

**Never:**
- Create a worktree when Step 0 detects existing isolation
- Use `git worktree add` when you have a native worktree tool (e.g., `EnterWorktree`). This is the #1 mistake — if you have it, use it.
- Skip Step 1a by jumping straight to Step 1b's git commands
- Create worktree without verifying it's ignored (project-local)
- Skip baseline test verification
- Proceed with failing tests without asking

**Always:**
- Run Step 0 detection first
- Prefer native tools over git fallback
- Follow directory priority: explicit instructions > existing project-local directory > default
- Verify directory is ignored for project-local
- Auto-detect and run project setup
- Verify clean test baseline

**Claude:** I'm using the using-git-worktrees skill to set up an isolated workspace.
  - `Bash: GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P); GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P); echo "GIT_DIR=$GIT`

**Claude:** Not currently in an isolated workspace — I'm on `master` directly.
  - `AskUserQuestion`
  - `Bash: cargo test --workspace 2>&1 | tail -40`

**Claude:** Baseline is clean. Now setting up the progress ledger and pre-flight scan before dispatching Task 1.
  - `Write: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** That's the old, already-completed plan's ledger. Overwriting it for the new plan.
  - `Write: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Pre-flight scan: I wrote this plan myself and already cross-checked type/interface consistency across tasks during the writing-plans self-review — no contradictions with the Global Constraints found. Proceeding to Task 1.
  - `Bash: ls /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-19-g`
  - `Bash: git rev-parse HEAD`
  - `Bash: cat /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/implementer-prompt.md`

**Claude:** Now dispatching the Task 1 implementer.
  - `Agent: Implement Task 1: grammar fix`

**Claude:** Task 1 implementer is running in the background — I'll wait for its completion notification rather than poll.

**You:** <task-notification>
<task-id>ac076e1fdc76a1738</task-id>
<tool-use-id>toolu_012vJNqN8u8LZGtZRBYa2UcA</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/ac076e1fdc76a1738.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 1: grammar fix" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>## Report

**Status:** DONE

**Commit:** `96ec67f feat(tree-sitter-plum): allow positional variant-construction calls`

**Test Summary:** All 18 corpus tests pass (100% success rate), including new variant construction call and all pre-existing class_call cases — no regressions.

**Key Actions:**
- Modified grammar.js: added `fn_argument_list`/`class_argument_list` conflict declaration and updated `fn_call` to accept `type_identifier` callees
- Generated expected tree: `Some(v)` parses as `fn_call` with `type_identifier`, not `class_call`
- All tests pass with no ERROR/MISSING nodes

**Report file:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md`</result>
<usage><subagent_tokens>38172</subagent_tokens><tool_uses>24</tool_uses><duration_ms>119795</duration_ms></usage>
</task-notification>

**Claude:** Task 1 done. Let me check the report and generate the review package before dispatching the reviewer.
  - `Bash: git log --oneline -3 && echo --- && git show --stat HEAD`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 2cee306 96ec67f`
  - `Bash: cat /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/task-reviewer-prompt.md`

**Claude:** Dispatching the Task 1 reviewer now.
  - `Agent: Review Task 1 (spec + quality)`

**Claude:** Task 1 reviewer is running in the background.

**You:** <task-notification>
<task-id>a96d45e2237740f81</task-id>
<tool-use-id>toolu_01DN574fNZPnixxjZ25wAxGj</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a96d45e2237740f81.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 1 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>### Spec Compliance
- ❌ Issues found: The `conflicts` array does not match the brief's prescribed text.
  - Brief (task-1-brief.md:52): `conflicts: ($) =&gt; [[$.fn_call, $.class_call]]`
  - Diff actual (review diff line 31): `conflicts: ($) =&gt; [[$.fn_argument_list, $.class_argument_list]]`
  - The implementer's own report (task-1-report.md:12, 22-26) confirms this was a deliberate, self-motivated deviation ("conflict resolution was declared on `[$.fn_argument_list, $.class_argument_list]` rather than the rules themselves"), not a transcription slip. Per the brief's own guidance ("if it reports one anyway... read the reported example carefully... don't just silence it"), a deviation like this should have been flagged as a decision point rather than silently substituted with no confirmation that the brief's literal prescription was even tried and found insufficient. The report gives a plausible rationale (empty `()` ambiguity between the two argument-list rules) but never states whether `[$.fn_call, $.class_call]` was attempted and failed, or whether it was simply skipped in favor of this alternative. Since `npx tree-sitter-cli generate` reportedly completed without "Unresolved conflict" errors and all 18 corpus cases (including `class_call` cases in `type.txt`) pass, the substitution is functionally validated by the test suite — but it is still a checkable deviation from the exact prescribed diff, and the brief's rationale for choosing `fn_call`/`class_call` specifically (the ambiguity being between the two *call* rules, not their argument lists) was not addressed or ruled out.
-`fn_call` rule change matches the brief exactly (`choice($.var_identifier, $.type_identifier)`).
- ✅ Corpus test case (function.txt) matches the brief's prescribed input and expected tree exactly, including one-node-per-line formatting.
- ✅ Scope respected: no generics/monomorphization, multi-subject match, or new pattern kinds touched. Diff is confined to `grammar.js` (2 lines) + `function.txt` (test case) + regenerated artifacts.
- ⚠️ Cannot verify from diff: whether `grammar.json`/`node-types.json`/`parser.c` binary diffs are purely `tree-sitter generate` output vs. hand-touched — trusting the report's claim since binary diffs are not line-reviewable, and there's no textual evidence of manual editing.

### Strengths
- Minimal, surgical diff — exactly the two grammar.js lines the brief called for plus one corpus case, no incidental churn.
- Pre-existing "why" comment above `fn_call` (about `var_identifier` lexing ambiguity) was left untouched, correctly not requiring an update since the rule's identifier-lexing rationale is unaffected by this change.
- Test case format matches existing corpus conventions exactly (separator style, one-node-per-line tree).
- Report is honest about the conflict-declaration substitution rather than silently claiming brief-verbatim compliance — makes the deviation catchable in review, which is what surfaced here.

### Issues

#### Critical (Must Fix)
None.

#### Important (Should Fix)
- **Undocumented deviation from prescribed conflict declaration** (grammar.js, conflicts array). The brief specifies `[[$.fn_call, $.class_call]]`; the shipped code has `[[$.fn_argument_list, $.class_argument_list]]`. Tests pass either way per the report, but the report doesn't show that the brief's literal form was tried and rejected — it jumps straight to the alternative with a self-constructed rationale. Before merging, either (a) confirm the brief's exact form also generates cleanly and passes the suite, and prefer it for spec fidelity, or (b) explicitly document in the report/commit why the argument-list-level conflict is necessary and the call-level one is insufficient (e.g., "call-level conflict alone leaves an unresolved conflict at empty-argument-list states because..."). As it stands this is a plausible-sounding but unverified justification for silently diverging from an explicit instruction.

#### Minor (Nice to Have)
- The report's "Key Insight on Conflicts" section reads more like exploratory narration than a permanent code comment, and none of that reasoning was captured as an in-grammar comment near the `conflicts` array (the codebase's convention, per the existing `fn_call` comment, is one-line "why" comments at the point of interest) — a future maintainer reading `grammar.js:55` alone won't know why this particular pair was chosen over the more obvious `fn_call`/`class_call` pair.

### Assessment
**Task quality:** Needs fixes
**Reasoning:** The functional grammar change and test coverage are correct and match the brief precisely, but the conflicts-array deviation is exactly the kind of "different conflict pair could resolve ambiguity differently or mask a real issue" risk flagged in the task's constraints, and it was not verified against the brief's literal prescription or explained with in-code justification — it should be reconciled (verified equivalent or documented) before this is considered done.</result>
<usage><subagent_tokens>37383</subagent_tokens><tool_uses>3</tool_uses><duration_ms>51472</duration_ms></usage>
</task-notification>
  - `Agent: Fix Task 1 review findings`

**Claude:** Fix subagent for Task 1's review finding is running in the background.

**You:** <task-notification>
<task-id>ab523b917da5b60b5</task-id>
<tool-use-id>toolu_01FDZQT8mYQ58iMfeya3bJ5t</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/ab523b917da5b60b5.output</output-file>
<status>completed</status>
<summary>Agent "Fix Task 1 review findings" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

**Commit created:** `a65c75b fix(tree-sitter-plum): document fn_argument_list/class_argument_list conflict rationale`

**Which conflicts-array pair ended up in the code, and why:** `[[$.fn_argument_list, $.class_argument_list]]` — testing the brief's prescribed `[[$.fn_call, $.class_call]]` produced an unresolved conflict at the argument-list parse state (empty parentheses ambiguity), proving the argument-list pair is necessary.

**Test summary:** All 18 corpus cases pass; generate runs cleanly; no regressions.

**Report file:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md` (appended with findings from this follow-up)</result>
<usage><subagent_tokens>32695</subagent_tokens><tool_uses>12</tool_uses><duration_ms>59257</duration_ms></usage>
</task-notification>

**Claude:** Now re-review to confirm the fix and close out Task 1.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 2cee306 a65c75b`
  - `Agent: Re-review Task 1 after fix`

**Claude:** Task 1 re-review is running in the background.

**You:** <task-notification>
<task-id>a3e345cf5275658b0</task-id>
<tool-use-id>toolu_01Qk2cUqEdTDnVG1KE4ahfYS</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a3e345cf5275658b0.output</output-file>
<status>completed</status>
<summary>Agent "Re-review Task 1 after fix" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>All confirmed against source. Now writing the verdict.

### Spec Compliance
- ✅ Spec compliant: the prior "Important" finding is resolved. `grammar.js:31-35` now carries a 3-line in-code comment directly above `conflicts: ($) =&gt; [[$.fn_argument_list, $.class_argument_list]],` explaining the conflict arises at the empty-argument-list parse state (e.g. `Foo()`), where `fn_argument_list` and `class_argument_list` cannot be distinguished until `:` or `)` is seen, and that declaring the conflict at the argument-list level (not the call level) resolves it.
- The report's Follow-up section shows concrete `generate` output for the brief's literal form (`[[$.fn_call, $.class_call]]`) failing with a specific "Unresolved conflict" error at `const_identifier '=' type_identifier '(' ')' • '/' …`, listing the two colliding reductions as `fn_argument_list '(' ')'` vs `class_argument_list '(' ')'` — this is internally consistent with the stated rationale (the ambiguity is on the empty-arglist reduction, a symbol pair the call-level conflict declaration wouldn't name), so the technical justification is genuine and checkable, not hand-waved.
- Final `conflicts:` value in the diff (`grammar.js:35`) matches exactly what the report claims: `[[$.fn_argument_list, $.class_argument_list]]`.
- Task brief's Step 2 explicitly says "If it reports one anyway... read the reported example carefully before changing anything further" — the fix subagent did exactly this rather than blindly reverting, satisfying the brief's own contingency instruction.

### Strengths
- Fix directly targets the prior finding: literal brief form was actually tried (not just claimed), real generate output captured and quoted, and the deviation is now documented in-grammar rather than only in a report.
- Comment content is technical and specific (names the actual parse state and the actual colliding rule pair), not generic filler.
- Cumulative diff is minimal: only the `conflicts` line/comment changed between the original and fix commits, per `git diff` file list (`grammar.js` +7/-2 lines, corpus file, and regenerated binaries) — no unrelated churn.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- The new comment is 3 lines, while the plan's global constraint says "one-line 'why' comments only." In practice, existing comments in this same file are inconsistent with that literal rule — e.g. the comment above `fn_call` itself (`grammar.js:49-54`, 6 lines) and the one above the class-method-name choice (`grammar.js:369-372` in current file, 4 lines) — so the new 3-line comment matches the file's actual prevailing style (multi-line paragraph "why" comments) even though it's technically longer than one line. Not worth blocking on given the precedent, but worth flagging since it's a literal mismatch with the stated constraint.

### Assessment
**Task quality:** Approved
**Reasoning:** The prior Important finding is resolved with a real, verifiable technical justification (concrete generate-error output) and an in-grammar comment that accurately reflects it; the only residual nit is a comment-length technicality that matches existing file precedent rather than violating the codebase's actual style.</result>
<usage><subagent_tokens>42966</subagent_tokens><tool_uses>9</tool_uses><duration_ms>61146</duration_ms></usage>
</task-notification>

**Claude:** Task 1 is approved. Updating the ledger and moving to Task 2.
  - `Bash: git log --oneline -5`
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-19-g`
  - `Bash: git rev-parse HEAD`
  - `Agent: Implement Task 2: checker enum variant info`

**Claude:** Task 1 approved and merged. Task 2 (checker changes) implementer is now running in the background.

**You:** <task-notification>
<task-id>a73898982b7d62405</task-id>
<tool-use-id>toolu_017Gh2SnK8CmzDm2WZisVSTs</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a73898982b7d62405.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 2: checker enum variant info" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE_WITH_CONCERNS

Commit: `7ac1d37 feat(plum-checker): general enum variant tags, field types, and construction checks`

Test summary: 29/30 in `checker_tests` pass (28 pre-existing + 2 new/adjusted; 5 of 6 brief tests green); `examples_test.rs` unaffected/green. One test, `bare_enum_tag_unifies_with_owning_enum_type`, fails — not a checker-logic bug but a pre-existing tree-sitter-plum grammar limitation: `_statement` in `grammar.js` only allows `$.primary_expression` (not full `$.expression`) as a multi-line function body's trailing statement, so `o == None` as the last line of an indented body never parses as one `Compare` node (`==` is dropped, becomes an `ERROR` + two separate statements). This predates Task 2, is out of this task's file scope (`plum-checker/src/lib.rs` only), and needs the `tree-sitter` CLI (not installed here) to fix/regenerate. All 5 code edits match the brief byte-for-byte.

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md`</result>
<usage><subagent_tokens>97294</subagent_tokens><tool_uses>54</tool_uses><duration_ms>417219</duration_ms></usage>
</task-notification>

**Claude:** Let me verify this claimed grammar limitation directly before trusting the implementer's diagnosis.
  - `Bash: cat > /tmp/t.plum << 'EOF' enum Option = | Some(Int) | None isNone(o: Option) -> Bool = o == None EOF cat > /tmp/parse_test.rs << 'RUST' fn main() { let src = s`
  - `Bash: cargo test -p plum-checker --test checker_tests bare_enum_tag_unifies_with_owning_enum_type -- --nocapture 2>&1 | tail -40`
  - `Bash: ls plum-cli/src; grep -n "fn main" plum-cli/src/*.rs | head`
  - `Read: /Users/pyrossh/Code/plum/plum-cli/src/main.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-checker/tests/checker_tests.rs`
  - `Bash: cat >> /Users/pyrossh/Code/plum/plum-checker/tests/checker_tests.rs << 'EOF' #[test] fn debug_dump_ast_temp() { let src = "\ enum Option = | Some(Int) | None is`
  - `Bash: sed -i '' '$ d' plum-checker/tests/checker_tests.rs # remove the temp test block entirely (multiple lines) - safer to use git checkout since implementer's versi`
  - `Bash: git checkout -- plum-checker/tests/checker_tests.rs && git diff --stat plum-checker/tests/checker_tests.rs`

**Claude:** Good, reverted cleanly. Now let me test the inline-expression form to see if it avoids the grammar limitation.
  - `Bash: cat >> /Users/pyrossh/Code/plum/plum-checker/tests/checker_tests.rs << 'EOF' #[test] fn debug_dump_ast_temp2() { let src = "\ enum Option = | Some(Int) | None i`
  - `Bash: git checkout -- plum-checker/tests/checker_tests.rs cat >> /Users/pyrossh/Code/plum/plum-checker/tests/checker_tests.rs << 'EOF' #[test] fn debug_dump_ast_temp3`
  - `Bash: git status --short`

**Claude:** Confirmed independently: the checker fix itself is correct — the test just used a multi-line indented body, and this grammar only allows `$.primary_expression` (not full `$.expression`) as a body's trailing statement, so `o == None` split into two separate statements with an `ERROR` node. The inline single-line form (`= o == None`) parses cleanly with no error. This is a real, separate, pre-existing grammar limitation, unrelated to Task 2. Dispatching a small fix to rewrite just that one test.
  - `Agent: Fix Task 2 test to avoid grammar limitation`

**Claude:** Fix subagent running in the background.

**You:** <task-notification>
<task-id>a460ae3a9e73d3d1d</task-id>
<tool-use-id>toolu_0153tC3RkWvHMuK8wbS3baSf</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a460ae3a9e73d3d1d.output</output-file>
<status>completed</status>
<summary>Agent "Fix Task 2 test to avoid grammar limitation" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Perfect! Here's the final report:

**Status:** DONE

**Commit:** 0c09f5d - fix(plum-checker): rewrite bare_enum_tag test to avoid unrelated grammar limitation

**Test Summary:** All 30 plum-checker tests pass (29 checker_tests.rs + 1 examples_test.rs) — no regressions

**Spawn Task:** Grammar limitation flagged successfully (task_1ded38c4) for multi-line indented body fix in separate follow-up</result>
<usage><subagent_tokens>29146</subagent_tokens><tool_uses>7</tool_uses><duration_ms>47340</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 2 reviewer against the full commit range.
  - `Bash: git log --oneline -5 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package a65c75b`
  - `Agent: Review Task 2 (spec + quality)`

**Claude:** Task 2 reviewer running in the background.

**You:** <task-notification>
<task-id>a179eba1a6679818e</task-id>
<tool-use-id>toolu_018j5FfyfQ6AZ9KbjFULZX5R</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a179eba1a6679818e.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 2 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Error message style is consistent with the rest of the file (`"call '{}': expected {} args, got {}"` mirrors `"variant '{}': expected {} arg(s), got {}"`; `"constructor pattern '{}' expects {} field(s), got {}"` is analogous to existing patterns). No stylistic deviation.

### Spec Compliance
✅ Spec compliant. All 5 prescribed code edits in `plum-checker/src/lib.rs` match the brief byte-for-byte (verified directly against the diff): `EnumVariantInfo`/`EnumVariants` (lines 30-40 of diff), the two `True`/`False` inserts (tag 1/0, `field_types: vec![]`), the `Item::Enum` tag-enumeration loop with `plum_type_from_ast` mapping, the `TypeName` arm (owning-enum lookup with permissive fallback), the `FnCall` variant-construction branch (arg-count + per-arg unify, falls through to existing function-call path), and `check_pattern`'s `CasePattern::Class` arm (field-count validation + real per-field types, permissive fallback for unmodeled variants) plus doc-comment update. All 6 brief test cases present in `plum-checker/tests/checker_tests.rs`.
- The one legitimate deviation — `bare_enum_tag_unifies_with_owning_enum_type`'s `src` inlined onto one line (`isNone(o: Option) -&gt; Bool = o == None`) instead of the brief's multi-line form — still exercises the intended regression: `o` (typed `Option`) compared against bare tag `None`, which only type-checks under the fixed `TypeName` arm (owning-enum lookup) rather than the old self-typing behavior. Confirmed this preserves original test intent; not a spec violation given the documented, verified grammar limitation.
- `EnumVariantInfo`/`EnumVariants` shape matches the exact names/types later tasks depend on. `ast::EnumVariant.fields: Vec&lt;String&gt;` correctly feeds `plum_type_from_ast`.

### Strengths
- Both new branches (`FnCall`, `check_pattern::Class`) correctly fall through/fall back to prior permissive behavior for names not in `enum_variants`, preserving existing behavior for classes/unknown calls (verified via `grep` that no other file references the old `EnumVariants = BTreeMap&lt;String, String&gt;` shape, so the type change is self-contained).
- Error messages (`"variant '{}': expected {} arg(s), got {}"`, `"constructor pattern '{}' expects {} field(s), got {}"`) follow the file's existing `format!` conventions (quoted names, `'{}'... {} ... , got {}` phrasing) — consistent with `"call '{}': expected {} args, got {}"` a few lines below.
- The follow-up fix is minimally scoped: touches only the one test's `src` string, doesn't paper over the checker logic, and the controller's own AST-dump verification (not just report claims) backs the grammar-limitation diagnosis.
- No stray/unrelated edits — diff is confined to the two files brief specifies.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- The grammar limitation (`_statement` restricted to `$.primary_expression` instead of `$.expression` for a multi-line body's trailing statement) is a real, silently-lurking gap that will bite any future test/example using a bare comparison/boolean/ternary as a multi-line function's last statement. Already correctly flagged as an out-of-scope follow-up by the implementer — no action needed here, just noting it's tracked rather than forgotten.
- `field_types` construction (`plum_type_from_ast(&amp;ast::Type { name: f.clone(), generics: vec![] })`) allocates a throwaway `ast::Type` per field just to reuse the existing mapping function; harmless and matches the brief's exact prescribed code, so not a deviation, just a mild wart inherited from the brief itself.

### Assessment
**Task quality:** Approved
**Reasoning:** The diff is a verbatim, verified implementation of all 5 brief-specified edits and 6 tests, the one test rewrite is a narrowly-scoped, well-justified fix for an independently-confirmed unrelated grammar bug that preserves the original test's intent, and error-message/style conventions match the surrounding file.</result>
<usage><subagent_tokens>53668</subagent_tokens><tool_uses>9</tool_uses><duration_ms>41550</duration_ms></usage>
</task-notification>

**Claude:** Task 2 approved. Updating the ledger and moving to Task 3.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-19-g`
  - `Agent: Implement Task 3: variant construction codegen`

**Claude:** Task 3 implementer running in the background.

**You:** <task-notification>
<task-id>a53c28f511b60d605</task-id>
<tool-use-id>toolu_015mTj9akeAaN9rEDyoGXpC8</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a53c28f511b60d605.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 3: variant construction codegen" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Committed successfully (`.recall/` untracked directory was not staged, as intended).

**Status:** DONE

Commit: `380a51c` — feat(plum-wasm-codegen): compile general enum variant construction

Test summary: `payload_free_variant_construction_compiles` fully passes (GREEN); `payload_variant_construction_compiles_and_runs` and `multi_field_variant_construction_compiles_and_runs` fail only on `"codegen: constructor match patterns are not yet supported"` (the `ast::CasePattern::Class` arm at `plum-wasm-codegen/src/lib.rs:904`, Task 4's scope) — isolate-verified by swapping their `match` bodies for literal returns, which then compiled and ran correctly (7-arg and multi-field construction both confirmed working). No regressions in `cargo test -p plum-wasm-codegen` (26/28 pass, same 2 expected fails) or `cargo test --workspace`.

Concerns: none. Diff matches the brief byte-for-byte.

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-3-report.md`</result>
<usage><subagent_tokens>69046</subagent_tokens><tool_uses>33</tool_uses><duration_ms>199760</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 3 reviewer.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 0c09f5d`
  - `Agent: Review Task 3 (spec + quality)`

**Claude:** Task 3 reviewer running in the background.

**You:** <task-notification>
<task-id>a4a269db3beb74157</task-id>
<tool-use-id>toolu_01BaZe1duHgw6fZmkZU4WbND</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a4a269db3beb74157.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 3 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>No existing tests depend on the old TypeName error text, so no regression risk there. Everything checks out against the brief and the runtime layout constraints.

### Spec Compliance
- ✅ Spec compliant. Diff matches the brief's prescribed edits byte-for-byte (import at `plum-wasm-codegen/src/lib.rs:6`, `TypeName` arm at `:1075-1081`, `Collector::walk_expr`'s `FnCall` arm at `:480-492`, `compile_expr`'s `FnCall` arm at `:1049-1066`, new `compile_variant_construction` at `:1194-1256`). All 3 prescribed tests present verbatim in `plum-wasm-codegen/tests/codegen_tests.rs:220-274`.
- Heap layout verified exact: `size = (1 + info.field_types.len() as i32) * 8` (`:1231` region, confirmed via read), tag stored at `offset: 0` (`I32Store` with `offset: 0`), field `i` stored at `offset = ((i+1) as u64) * 8` — matches the binding constraint precisely (compare to `ClassCall`'s sibling code at `:1088/1114`, which correctly omits the `+1` since classes have no tag slot).
- Scratch-pool reuse confirmed correct: `Collector::walk_expr`'s `FnCall` arm (`:480-492`) inserts into the same `self.classcall_scratch` / `self.next_classcall_slot` fields `ClassCall` uses (`:452-455`), keyed by `expr as *const ast::Expr as usize`. Since `ClassCall` and payload-carrying `FnCall` nodes are distinct AST nodes with distinct addresses, no key collisions are possible — verified by reading both call sites and the shared `Collector` struct fields.
- `compile_variant_construction` (`:1198-1256`) validates `call.args.len() != info.field_types.len()` before doing anything else, returning a clear error before compiling any argument or touching the bump pointer.
- `TypeName` arm (`:1075-1081`) correctly special-cases `Some(info) if info.field_types.is_empty()``i32.const tag`, which transparently covers `True`/`False` since those are pre-registered in `enum_variants` with `field_types: vec![]` (`plum-checker/src/lib.rs:78-79`) — no hardcoded `True`/`False` branch needed, matching brief intent. Bare use of a payload-carrying variant name correctly produces a distinct, clear error (`Some(_)` arm) rather than silently doing the wrong thing.
- Confirmed via `plum-checker/src/lib.rs:368` (`ast::CasePattern::Class` handling in the checker) that the 2 still-failing tests' failure mode is genuinely inside match/type-checking or match-lowering territory, not variant construction — the report's isolate-verification claim (swap match body for a literal return) is consistent with the code structure; no evidence in the diff contradicts it.
- No existing test asserted on the old `TypeName` error text (grep for "is not yet supported as a value" / "carries a payload" in tests turned up nothing), so the rewritten arm carries no regression risk to existing assertions.

### Strengths
- Exact adherence to the prescribed layout invariant Task 4 depends on; field/tag offset arithmetic double-checked against the sibling `ClassCall` code path and found consistent in convention (8-byte stride) while correctly differing in the tag-slot offset.
- New function's error messages follow the file's existing `format!("codegen: ...")` convention consistently (arg-count mismatch, missing scratch slot, bare-payload-variant use).
- Argument re-extraction pattern (`ast::Arg::Positional/Keyword/Pair`) duplicated faithfully from the existing ordinary-call path rather than diverging.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- The `ast::Arg::Positional/Keyword/Pair``arg_expr` extraction closure is now duplicated three times in this file (ordinary `FnCall`, `compile_variant_construction`, and `Attribute::Method`). Pre-existing duplication pattern, not introduced by this task, and out of scope to refactor here — flagging only for awareness, not requesting a fix.
- No test exercises the arg-count-mismatch error path or the bare-payload-variant-name error path added in this task; the brief didn't require it and the isolate-verification technique substitutes for confidence on the happy path, but these two new error branches are currently only reachable in the report's manual testing, not in committed test code.

### Assessment
**Task quality:** Approved
**Reasoning:** The diff matches the brief's prescribed edits exactly, the heap layout Task 4 depends on is implemented precisely as specified (tag at offset 0, fields at `(i+1)*8`, 8-byte stride), the scratch-pool reuse is collision-free by construction (distinct pointer keys), and the two still-failing tests fail for the correct, in-scope reason (unimplemented match lowering, not a construction bug).</result>
<usage><subagent_tokens>55279</subagent_tokens><tool_uses>8</tool_uses><duration_ms>59053</duration_ms></usage>
</task-notification>

**Claude:** Task 3 approved. Updating the ledger and moving to Task 4.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-19-g`
  - `Agent: Implement Task 4: match lowering for enums`

**Claude:** Task 4 implementer running in the background.

**You:** <task-notification>
<task-id>adc11508de0e9d355</task-id>
<tool-use-id>toolu_01XVyMGpn3WusRNfd1ARpWt6</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/adc11508de0e9d355.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 4: match lowering for enums" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE_WITH_CONCERNS

Commit: `8ecbf56` — feat(plum-wasm-codegen): compile general enum match patterns (tags + constructors)

Test summary: `plum-wasm-codegen` codegen_tests 30/30 green (includes the 2 carried-over tests, now GREEN, and 2 new tests). Workspace-wide: everything green except 1 pre-existing failure in `examples_test.rs` (`match_example_reports_clear_unsupported_pattern_errors`), which the brief explicitly documents as expected fallout, deferred to Task 5.

Concerns:
- The example-test failure noted above is scope for Task 5, not this task — flagging so it isn't mistaken for a regression.
- `.superpowers/sdd/task-4-report.md` previously contained an unrelated report from a differently-numbered older plan iteration; I overwrote it with this task's report per instructions.

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-4-report.md`</result>
<usage><subagent_tokens>62425</subagent_tokens><tool_uses>26</tool_uses><duration_ms>173732</duration_ms></usage>
</task-notification>
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 380a51c`
  - `Agent: Review Task 4 (spec + quality)`

**Claude:** Task 4 reviewer running in the background.

**You:** <task-notification>
<task-id>ac87a492bd7b68db9</task-id>
<tool-use-id>toolu_01VxdjK6YNbPgCXsvCmdDdcP</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/ac87a492bd7b68db9.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 4 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Confirmed: `EnumVariantInfo` has `tag: i32` and `field_types: Vec&lt;PlumType&gt;`, consistent with usage. `True`/`False` are pre-registered with tags 1/0, so `Bool` matching stays backward compatible. This confirms the constructor arm's zip and `Collector`'s binding logic type-check correctly, and locals get the right `ValType` via `plum_type_to_valtype`.

### Spec Compliance
- ✅ Spec compliant

All four prescribed edits match the brief's exact prescribed code verbatim:
- `Collector::walk_stmt`'s `Match` arm (lib.rs, Class-pattern branch added, binding field names via `self.bind(n, fty.clone())`) — confirmed real field types are threaded through `collector.named` → local slot `ValType` assignment (lib.rs:573-578), so constructor-pattern bindings get correctly-typed wasm locals, not `TVar`/placeholder.
- `compile_variant_eq_arm` generalized to `ctx.enum_variants.get(name)` lookup using `info.tag`, replacing the hardcoded `True`/`False` match. `Bool`'s variants remain registered with `tag: 1`/`tag: 0` in `plum-checker/src/lib.rs:78-79`, so existing Bool-matching behavior is preserved under the generalized path.
- New `compile_variant_constructor_arm`: tag test loads from offset 0 (`I32Load` at `offset: 0`), field loads at `(i+1)*8` via `I64Load`/`F64Load`/`I32Load` selected by `plum_type_to_valtype(field_ty)` — matches Task 3's `[tag][field0][field1]...]` 8-byte-stride layout exactly.
- Non-`Name`/`Wildcard` sub-patterns produce a clear error: `"codegen: only bare bindings or '_' are supported inside a constructor pattern"` — no silent misbehavior.
- Wired into `compile_match_arms`'s `CasePattern::Class` branch, replacing the old blanket `Err`.
- The two new tests (`non_bool_bare_tag_pattern_runs_correctly`, `constructor_pattern_wildcard_field_runs_correctly`) match the brief's prescribed test code verbatim.
- Task 3's two carried-over tests require no test-file changes — confirmed the diff touches only lib.rs logic plus two newly appended tests; no modification to the existing `payload_variant_construction_compiles_and_runs`/`multi_field_variant_construction_compiles_and_runs` tests.
- Out-of-scope items respected: no generics monomorphization, no multi-subject match handling added, no nested constructor-pattern support (explicitly rejected with an error).
- Noted per instructions, not flagged as an issue: `examples_test.rs`'s `match_example_reports_clear_unsupported_pattern_errors` now fails as an expected, deferred consequence (Task 5's job).

### Strengths
- Field-load `ValType` dispatch (`I64Load`/`F64Load`/`I32Load` by `plum_type_to_valtype`) correctly generalizes beyond the Int-only case Task 3 exercised, so a hypothetical `Float`-payload variant would load correctly too.
- Error messages are consistent in style and wording with the rest of the file (`"codegen: ... not yet supported"`, `"codegen: ... expects N field(s), got M"`), matching sibling error strings elsewhere in `compile_match_arms`.
- Reuses `scratch_local`/`subject_vt`/`rest`-recursion plumbing exactly as established by neighboring arm-compiler functions — no new control-flow pattern introduced.
- `ctx.type_env.borrow_mut().insert(...)` for each bound field keeps downstream `compile_expr` type-checking consistent with the local's real `ValType`, avoiding a mismatch between the `LocalCtx.locals` local index and the type used to interpret it in later expression codegen.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- `compile_variant_constructor_arm`'s tag-mismatch error path: if the `if` (tag equality) test fails, execution falls into `Else``compile_match_arms(rest, ...)`, consistent with other arms — fine, no issue, just noting no dedicated test exercises a tag mismatch cascading past multiple constructor arms (e.g., matching `None` when the first case is `Some(x)`); indirectly covered by `constructor_pattern_wildcard_field_runs_correctly`'s `None` case and Task 3's multi-variant tests, so coverage is adequate but not maximally explicit.
- No test exercises the new "only bare bindings or `_`" error path for a rejected sub-pattern kind (e.g., `Some(Some(x))`); brief did not require it, so this is optional polish, not a gap against spec.

### Assessment
**Task quality:** Approved
**Reasoning:** The diff is a verbatim, correctly-wired implementation of the brief's prescribed edits; offset arithmetic matches Task 3's runtime layout exactly, error handling for unsupported sub-patterns is clear and non-silent, and the Collector-side type binding fix is verified (via `plum-checker`'s `EnumVariantInfo`/`plum_type_to_valtype`) to produce correctly-typed wasm locals rather than placeholders.</result>
<usage><subagent_tokens>49558</subagent_tokens><tool_uses>8</tool_uses><duration_ms>53507</duration_ms></usage>
</task-notification>

**Claude:** Task 4 approved. Updating the ledger and moving to the final task, Task 5.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-19-g`
  - `Agent: Implement Task 5: examples and docs`

**Claude:** Task 5 implementer running in the background.

**You:** <task-notification>
<task-id>ae7026553adbccb20</task-id>
<tool-use-id>toolu_0154Dt9ezrDBvz62oxTEbcea</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/ae7026553adbccb20.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 5: examples and docs" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>---

## Final Status Report

**Status:** BLOCKED

**Issue:** Critical codegen bug prevents completion of Task 5 as specified in the brief. The codegen generates invalid wasm (validation error: "type mismatch: expected i64 but nothing on stack") whenever any match statement is added to match.plum, even simple patterns like `match n x =&gt; x`.

**What was completed:**
- ✅ Test code updated correctly (examples_test.rs)
- ✅ README documentation updated (Known Gaps section)
- ⚠️ match.plum has minimal working version but cannot include required match statements

**What's blocked:**
- Step 2/3: Cannot add `main() -&gt; Int = describeOption(Some(5))` due to codegen bug
- Step 6: Cannot commit working solution since test suite has the failing test

**Test summary:**
- Workspace: 133/133 tests pass (including all codegen tests showing match IS implemented)
- Tree-sitter: 18/18 grammar tests pass
- Examples test: 6/7 pass; 1 BLOCKED by codegen (not a test failure but a code generation issue)

**Concerns:**
The codegen tests (codegen_tests.rs) all pass including match and constructor pattern tests. The bug only manifests when match statements appear in actual exported functions in match.plum. This suggests the issue is context-dependent, likely in how Tasks 1-4's implementation handles multiple exported functions containing match statements.

**Report file:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-5-report.md`</result>
<usage><subagent_tokens>73515</subagent_tokens><tool_uses>50</tool_uses><duration_ms>418882</duration_ms></usage>
</task-notification>
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/task-5-report.md`

**Claude:** Let me verify this diagnosis myself before deciding how to proceed — the implementer's finding could indicate a real bug or just a convention mismatch (every existing match codegen test uses explicit `return` in each arm; match.plum's example functions don't).
  - `Bash: cat > /tmp/repro.plum << 'EOF' bindExample(n: Int) -> Int = match n x => return x main() -> Int = bindExample(5) EOF cat > /tmp/repro_novar.plum << 'EOF' bindEx`
  - `Bash: cat >> /Users/pyrossh/Code/plum/plum-wasm-codegen/tests/codegen_tests.rs << 'EOF' #[test] fn debug_temp_tail_match_with_return() { let src = "\ bindExample(n: I`
  - `Bash: git checkout -- plum-wasm-codegen/tests/codegen_tests.rs cat >> /Users/pyrossh/Code/plum/plum-wasm-codegen/tests/codegen_tests.rs << 'EOF' #[test] fn debug_temp`
  - `Bash: cat /Users/pyrossh/Code/plum/examples/match.plum; echo ---; git status --short; echo ---diff---; git diff --stat`
  - `Bash: git diff examples/match.plum`

**Claude:** I'll revert match.plum and dispatch a fix subagent with the precise diagnosis so it doesn't re-debug from scratch.
  - `Bash: git checkout -- examples/match.plum && cat examples/match.plum`
  - `Agent: Fix and complete Task 5`

**Claude:** I found the root cause myself (a pre-existing, orthogonal codegen limitation: `match`/`if` as a function's final statement drops its value unless every arm uses explicit `return` — unrelated to the enum work) and re-dispatched Task 5 with that diagnosis plus a fix, so it can proceed without re-debugging.

**You:** <task-notification>
<task-id>abd1d11071520b842</task-id>
<tool-use-id>toolu_01EAA4WjZ9986du9fG7LBPEC</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/abd1d11071520b842.output</output-file>
<status>completed</status>
<summary>Agent "Fix and complete Task 5" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

- Commit created: `fe36a46` — "docs+test: general enum support is complete; update known gaps and example"
- Test summary: `cargo test --workspace` fully green (37 tests across all crates, 0 failed); `tree-sitter-cli test` 18/18 corpus tests passed (100%)
- `spawn_task` confirmed succeeded (task_id `task_c2430dde`)
- Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-5-report.md`

Key deviation: rewrote all match arms in `examples/match.plum`'s five existing functions to use explicit `return &lt;expr&gt;` (codebase convention) since a bare tail expression through a match statement drops the value in current codegen — a pre-existing, orthogonal gap now tracked via the spawned follow-up task rather than fixed here.</result>
<usage><subagent_tokens>52601</subagent_tokens><tool_uses>18</tool_uses><duration_ms>117583</duration_ms></usage>
</task-notification>

**Claude:** Task 5 done. Dispatching the reviewer now.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 8ecbf56`
  - `Agent: Review Task 5 (spec + quality)`

**Claude:** Task 5 reviewer running in the background.

**You:** <task-notification>
<task-id>aa5e8263476ed7665</task-id>
<tool-use-id>toolu_01TUUruo43JLyQjpp6N8C6GJ</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/aa5e8263476ed7665.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 5 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Good — other examples already use explicit `return` consistently, confirming this was pre-existing codebase convention, not something invented for this task. Nothing else in the diff's scope is affected.

Nothing further to verify; I have enough to write the review.

### Spec Compliance
- ✅ Spec compliant — all four brief steps executed exactly: `examples_test.rs` matches the brief's prescribed replacement verbatim (`plum-wasm-codegen/tests/examples_test.rs:72-90`, confirmed via `git show fe36a46`), exactly one `strings_example_reports_clear_interpolation_error` remains, `examples/match.plum` has the brief's exact `main` appended, and `README.md`'s Known Gaps bullets match the brief's prescribed replacement text exactly (`README.md` diff hunk).
- The documented deviation (rewriting all 5 pre-existing functions' match arms to add explicit `return`) is a reasonable, minimal, in-scope fix: verified against `git show 8ecbf56:examples/match.plum` that every arm in all 5 functions was a bare tail expression, so every rewritten arm actually needed the change (no unnecessary edits) — each is a 1:1 syntactic substitution (`&lt;expr&gt;``return &lt;expr&gt;`) with no semantic change, since match is the sole/final statement in each function body. This also matches the codebase's existing convention (`plum-wasm-codegen/tests/codegen_tests.rs` and other examples like `control_flow.plum`, which already write every arm/branch with explicit `return`), so it doesn't introduce a new style.
- Checked the README's `match` section prose above Known Gaps (brief's Step 4 second half, "around line 304"): the pre-existing prose already only mentions multi-subject match as unsupported ("Multiple comma-separated subjects/patterns are accepted by the grammar but not yet lowered by codegen") — no stale enum-tag caveat existed there to trim. Report's claim of "no changes needed" is correct.

### Strengths
- The deviation is well-reasoned and well-scoped: it fixes only the example file (data), not the compiler, staying inside the stated boundary against fixing the tail-position gap in this task.
- `match_example_compiles_and_runs_correctly` goes further than the brief's literal text by using `assert_compiles` (which runs `wasmparser::validate` on the bytes), consistent with the pattern used by `methods_compiles_and_runs_correctly` right above it — genuine end-to-end validation, not just "doesn't error."
- The follow-up gap was properly externalized via `spawn_task` rather than silently left undocumented anywhere, and the report's root-cause description (`compile_block_as_fn_body` dropping match/if tail values when arms lack explicit `return`) is specific and actionable.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
- README's Known Gaps section makes an explicit, workspace-wide promise: "plum-wasm-codegen reports a clear error rather than silently producing wrong code" for unsupported syntax. The newly-discovered tail-position gap directly violates that promise for a very ordinary, non-enum-specific pattern (any function ending in `match`/`if` without explicit `return`) — as evidenced by the fact that all 5 pre-existing functions in this very file happened to hit it. Tracking it solely via an internal `spawn_task` follow-up (invisible to a reader of the public README) means the docs now overstate the codegen's actual safety guarantees. This isn't enum-specific and the brief didn't anticipate it, but given this is explicitly the plan's final documentation task and the gap was discovered during it, a one-line addition to Known Gaps (or a caveat on the "reports a clear error" sentence) would have been more complete. Not a blocker since it's genuinely orthogonal to enum support and the brief's scope, but worth a follow-up doc patch.

#### Minor (Nice to Have)
None.

### Assessment
**Task quality:** Approved
**Reasoning:** All brief-mandated edits are present and verified byte-for-byte against the diff; the out-of-brief deviation (adding `return` to pre-existing match arms) is necessary, minimal, semantically inert, and consistent with established codebase convention rather than a disguised compiler fix. The one gap — not mentioning the tail-position bug in the public README — is a documentation-completeness nit, not a spec or correctness defect, and doesn't block this task or the overall enum-support feature, which otherwise reads as complete and coherent from this final diff.</result>
<usage><subagent_tokens>44534</subagent_tokens><tool_uses>7</tool_uses><duration_ms>61972</duration_ms></usage>
</task-notification>
  - `Agent: Add README doc note for tail-position gap`

**Claude:** Small doc fix running in the background for Task 5's Important finding.

**You:** <task-notification>
<task-id>a434fcdc54c20da87</task-id>
<tool-use-id>toolu_01AP4Aj1tRXiue6M7TJW5ip7</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a434fcdc54c20da87.output</output-file>
<status>completed</status>
<summary>Agent "Add README doc note for tail-position gap" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

Commit created: `3fc279f docs: note tail-position match/if return-value gap in Known Gaps`

Test summary: All 47 tests passed (plum-checker, plum-wasm-codegen, examples, tree-sitter bindings — no failures)</result>
<usage><subagent_tokens>27199</subagent_tokens><tool_uses>5</tool_uses><duration_ms>28204</duration_ms></usage>
</task-notification>

**Claude:** Let me quickly verify the doc addition directly before closing out Task 5 (trivial enough not to need a full reviewer round).
  - `Bash: git show 3fc279f -- README.md`

**Claude:** Clean, one-line addition, matches existing style. Task 5 is complete. Updating the ledger and moving to the final whole-branch review.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** All 5 tasks complete. Now dispatching the final whole-branch review on the most capable model.
  - `Bash: git merge-base master 2cee306 2>/dev/null; echo "---"; git log --oneline 2cee306~1..3fc279f`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 2cee306 3fc279f`
  - `Skill: superpowers:requesting-code-review`

**You:** Base directory for this skill: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/requesting-code-review

# Requesting Code Review

Dispatch a code reviewer subagent to catch issues before they cascade. The reviewer gets precisely crafted context for evaluation — never your session's history. This keeps the reviewer focused on the work product, not your thought process, and preserves your own context for continued work.

**Core principle:** Review early, review often.

## When to Request Review

**Mandatory:**
- After each task in subagent-driven development
- After completing major feature
- Before merge to main

**Optional but valuable:**
- When stuck (fresh perspective)
- Before refactoring (baseline check)
- After fixing complex bug

## How to Request

**1. Get git SHAs:**
```bash
BASE_SHA=$(git rev-parse HEAD~1)  # or origin/main
HEAD_SHA=$(git rev-parse HEAD)
```

**2. Dispatch code reviewer subagent:**

Dispatch a `general-purpose` subagent, filling the template at [code-reviewer.md](code-reviewer.md)

**Placeholders:**
- `{DESCRIPTION}` - Brief summary of what you built
- `{PLAN_OR_REQUIREMENTS}` - What it should do
- `{BASE_SHA}` - Starting commit
- `{HEAD_SHA}` - Ending commit

**3. Act on feedback:**
- Fix Critical issues immediately
- Fix Important issues before proceeding
- Note Minor issues for later
- Push back if reviewer is wrong (with reasoning)

## Example

```
[Just completed Task 2: Add verification function]

You: Let me request code review before proceeding.

BASE_SHA=$(git log --oneline | grep "Task 1" | head -1 | awk '{print $1}')
HEAD_SHA=$(git rev-parse HEAD)

[Dispatch code reviewer subagent]
  DESCRIPTION: Added verifyIndex() and repairIndex() with 4 issue types
  PLAN_OR_REQUIREMENTS: Task 2 from docs/superpowers/plans/deployment-plan.md
  BASE_SHA: a7981ec
  HEAD_SHA: 3df7661

[Subagent returns]:
  Strengths: Clean architecture, real tests
  Issues:
    Important: Missing progress indicators
    Minor: Magic number (100) for reporting interval
  Assessment: Ready to proceed

You: [Fix progress indicators]
[Continue to Task 3]
```

## Integration with Workflows

**Subagent-Driven Development:**
- Review after EACH task
- Catch issues before they compound
- Fix before moving to next task

**Executing Plans:**
- Review after each task or at natural checkpoints
- Get feedback, apply, continue

**Ad-Hoc Development:**
- Review before merge
- Review when stuck

## Red Flags

**Never:**
- Skip review because "it's simple"
- Ignore Critical issues
- Proceed with unfixed Important issues
- Argue with valid technical feedback

**If reviewer wrong:**
- Push back with technical reasoning
- Show code/tests that prove it works
- Request clarification

See template at: [code-reviewer.md](code-reviewer.md)
  - `Bash: cat /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/requesting-code-review/code-reviewer.md`

**Claude:** Dispatching the final whole-branch review on the most capable model, per this skill's guidance.
  - `Agent: Final whole-branch review of general enum support`

**Claude:** All 5 tasks are done and individually approved. The final whole-branch review is running in the background now on Opus.

**You:** <task-notification>
<task-id>aeb5850a8aaa7e6ab</task-id>
<tool-use-id>toolu_01A551GRfK3PxTriCy3hcDvi</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/aeb5850a8aaa7e6ab.output</output-file>
<status>completed</status>
<summary>Agent "Final whole-branch review of general enum support" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Definitively confirmed. The canonical `Some(Int) | None` ordering silently miscompiles when the value is `None`. Now I have everything I need.

---

## Code Review: General Enum Support (2cee306..3fc279f)

### Strengths

- **Clean, additive checker change.** Extending `EnumVariants` from `BTreeMap&lt;String,String&gt;` to carry `EnumVariantInfo { enum_name, tag, field_types }` is well-scoped. The genuine pre-existing bug (a bare non-Bool tag like `None` typing as `TNamed("None")` instead of its owning enum) is correctly fixed at `plum-checker/src/lib.rs` `infer_expr`'s `TypeName` arm, and it's covered by a clear regression test with an explanatory comment.
- **Real validation added.** Arg-count/type checks for variant construction (`infer_expr` `FnCall` arm) and field-count/type checks for constructor patterns (`check_pattern`) replace the previous permissive fallback, and each new rule has both a positive and a negative test.
- **Layout convention is genuinely shared.** Construction (`compile_variant_construction`) and destructuring (`compile_variant_constructor_arm`) agree on tag at offset 0 (`I32Store`/`I32Load`, align 2) and field `i` at offset `(i+1)*8` with width-matched i32/i64/f64 access — the same 8-byte stride as class instances. HEAP_BASE (65536) is 8-aligned so the tag word stays aligned.
- **Grammar deviation was the right call.** The plan proposed `conflicts: [[$.fn_call, $.class_call]]`; the implementation instead declares the conflict at the argument-list level (`[[$.fn_argument_list, $.class_argument_list]]`) with a comment explaining that the ambiguity is really at the empty-`()` state. That is a more accurate fix than the plan's guess, and the corpus test proves `Some(v)` parses as `fn_call` with a `type_identifier` callee and no ERROR node.
- **Good error hygiene for out-of-scope constructs in code paths that are reached** — non-enum subject, wrong field count, nested sub-patterns, payload-carrying name used bare all return clear `codegen:` errors.
- Tests execute real wasm via wasmtime (`run_main`), not just "compiles".

### Issues

#### Critical (Must Fix)

**1. Constructor-pattern match arm dereferences payload-free sibling values → silent miscompile.**
`plum-wasm-codegen/src/lib.rs`, `compile_variant_constructor_arm` (the arm added at `compile_match_arms`'s `CasePattern::Class` case, ~lines 950–990).

The arm unconditionally emits `LocalGet(scratch); I32Load(offset 0); I32Const(tag); I32Eq` to read the subject's stored tag. But a payload-free sibling variant (e.g. `None`) is represented as a small integer tag, **not** a pointer. When such a value flows into a constructor arm ordered before its own tag arm, the code loads i32 from a tiny address (e.g. address 1). Low linear memory is zero-initialized, so the load returns 0 — which equals the tag of a payload variant declared at index 0 (`Some` in the canonical `Some(Int) | None`). The match wrongly enters the `Some` branch and destructures garbage.

The design spec's collision-safety argument ("heap pointers are always ≥ 65536, tags are small ints") only covers **bare-tag** patterns (direct i32 compare, which is safe). It never addressed the reverse direction — a **constructor** pattern tested against a payload-free value — and the implementation inherits that gap.

I confirmed this empirically (standalone wasmtime harness, repo untouched):

```
Some(Int)|None ; match None (Some first)  =  0  expect 5  &gt;&gt;&gt; WRONG &lt;&lt;&lt;
None|Some(Int) ; match None (None first)  =  5  expect 5  ok
Some(Int)|None ; match Some(7)            =  7  expect 7  ok
```

`unwrapOr(None, 5)` returns `0`. This is the single most common enum operation (`Option`/`Result` matching), it's silent (valid wasm, wrong result), and it's **tag-layout/arm-order dependent**`None | Some(Int)` happens to work, `Some(Int) | None` does not. The shipped `examples/match.plum` `describeOption` uses the broken ordering; it only escapes because `main` calls it exclusively with `Some(5)`.

Why untested: every codegen/example test constructs and matches only the payload variant (`Some(7)`, `Some(99)`, `Some(5)`); `None` is never the runtime subject of a match whose constructor arm precedes it. So the whole suite is green while the feature is broken.

Fix: before dereferencing in a constructor arm, guard that the subject is actually a heap pointer, e.g. emit `subject &gt;= I32Const(HEAP_BASE)` and AND it with the tag-equality test (payload-free values are always `&lt; HEAP_BASE`, so they can never satisfy a constructor arm). Add a regression test that matches a `None`-valued subject through a `Some(v)`-first match and asserts the `default` is returned.

#### Important (Should Fix)

**2. No clear-error test for nested constructor patterns (a stated out-of-scope guard).**
The code correctly returns `"codegen: only bare bindings or '_' are supported inside a constructor pattern"` for a nested sub-pattern, and the README lists it as a gap — but nothing tests it. The prompt specifically asks whether out-of-scope constructs are covered by a clear-error test or "just asserted in prose"; this one is prose + code only. Add a test compiling `Some(Some(v))` and asserting the error string, so a future refactor can't silently turn the miscompile back on. (Multi-subject `match` does have a codegen guard at `compile_match`, but likewise no test in this diff.)

**3. Enum-variant name silently shadows a class of the same name.**
Both `infer_expr`'s `FnCall` arm and `compile_expr`'s `FnCall`/`TypeName` arms consult `ctx.enum_variants` *before* classes/functions. `build_global_tables` inserts into `classes` and `enum_variants` independently with no collision check. So declaring `enum E = | Cat` alongside `type Cat = ...` makes `Cat(...)` resolve as variant construction, silently shadowing the class constructor with no diagnostic. Low likelihood, but it's exactly the collision the review brief calls out. At minimum, detect and reject a name registered as both; ideally document the precedence.

#### Minor (Nice to Have)

**4. Missing enum/class interop test the plan called for.** Task 3's testing plan specified "a variant used as a class field/method to confirm interop with existing class codegen" (and the brief asks about "deeply nested classes containing enum fields"). No such test was added. Given both use the same bump heap and the shared `classcall_scratch` pool, a small end-to-end test (class field holding a payload enum, constructed then destructured) would be worthwhile insurance — though I found no actual defect in the pool sharing itself (distinct `expr` pointer keys, verified by reasoning + the passing multi-field test).

**5. Binding type-env insert in the constructor arm is never scoped back.** `compile_variant_constructor_arm` does `ctx.type_env.borrow_mut().insert(n, ...)` but nothing removes it after the arm, unlike the `Collector`'s `saved`/restore dance. Harmless today (each field binding has a distinct local; the `Collector` already established bindings), but if two arms bind the same name at different types the stale hint could mislead `infer_local_type`. Cosmetic/robustness.

### Recommendations

- Block merge on issue #1 and land it with a `None`-through-constructor-arm regression test; the fix is a few instructions (pointer lower-bound guard) and is localized to `compile_variant_constructor_arm`.
- Fold in issues #2 and #3 while touching this code — both are small and both protect the invariants the design relies on.
- The two documented deviations are the right calls and should **not** block merge: (a) rewriting `bare_enum_tag_...`'s body to a single inline `o == None` genuinely sidesteps an unrelated pre-existing grammar bug without weakening the test's intent; (b) using explicit `return` in every `match.plum` arm matches the codebase's existing convention throughout its suite and correctly avoids masking the pre-existing tail-position gap, which is properly deferred and now documented in Known Gaps. Both discovered gaps are legitimately out of scope.
- README Known Gaps is otherwise accurate for what shipped — but note it currently *claims* enum-tag/constructor patterns work, while issue #1 means they don't for a common case. Fix the code, not the doc.

### Assessment

**Ready to merge?** No — with fixes.

**Reasoning:** The checker work, layout convention, grammar fix, and deviations are all sound, but there is a confirmed silent miscompile (issue #1): matching a payload-free variant (`None`) through a constructor arm dereferences a small-int tag as a heap pointer, returning wrong results for the canonical `Some(Int) | None` ordering — the entire test suite misses it because `None` is never the runtime match subject. That must be fixed and regression-tested before this whole-branch feature can be considered correct.</result>
<usage><subagent_tokens>88823</subagent_tokens><tool_uses>9</tool_uses><duration_ms>308597</duration_ms></usage>
</task-notification>

**Claude:** Critical finding — a real silent miscompile. Let me verify independently before dispatching a fix.
  - `Bash: cat >> /Users/pyrossh/Code/plum/plum-wasm-codegen/tests/codegen_tests.rs << 'EOF' #[test] fn debug_temp_none_through_some_first_match() { let src = "\ enum Opti`

**Claude:** Confirmed independently: `unwrapOr(None, 5)` returns `0` instead of `5` — a real silent miscompile. Dispatching one fix subagent for the complete findings list.
  - `Agent: Fix final review findings`

**Claude:** Fix subagent for the final review's Critical/Important/Minor findings is running in the background.

**You:** <task-notification>
<task-id>a3081e556aafe5ebc</task-id>
<tool-use-id>toolu_01B8hVpbJt8FZXPCJQbn47U3</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a3081e556aafe5ebc.output</output-file>
<status>completed</status>
<summary>Agent "Fix final review findings" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

Commits:
- 5e131a3 fix(plum-wasm-codegen): guard constructor-pattern match against payload-free sibling values
- 42d88a3 fix(plum-checker): reject a name declared as both a class and an enum variant

Test summary: `cargo test --workspace` fully green; `tree-sitter-cli test` 18/18 parses succeed (100%).

Findings fixed: #1 (critical, TDD RED→GREEN confirmed), #2 (test added, behavior already correct), #3 (checker fix + test), #4 (test added, no bug found). Deferred: #5 (stale `type_env` binding after constructor-pattern arm) — left as noted per the review's own guidance, since it's cosmetic/theoretical risk and a proper fix would need scoped-environment restructuring for no demonstrated bug.

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/final-review-fix-report.md`</result>
<usage><subagent_tokens>70671</subagent_tokens><tool_uses>38</tool_uses><duration_ms>225357</duration_ms></usage>
</task-notification>

**Claude:** Let me independently re-verify the critical fix before re-reviewing.
  - `Bash: git log --oneline -5`
  - `Bash: cat >> /Users/pyrossh/Code/plum/plum-wasm-codegen/tests/codegen_tests.rs << 'EOF' #[test] fn debug_temp_verify_none_fix() { let src = "\ enum Option = | Some(In`

**Claude:** Confirmed independently: the fix works, returns 5 correctly now. Dispatching a re-review of the whole branch.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 2cee306 42d88a3`
  - `Agent: Re-review final branch after fixes`

**Claude:** Re-review of the whole branch running in the background.

**You:** <task-notification>
<task-id>a4778467b2e5b6ab3</task-id>
<tool-use-id>toolu_01FhiUM5xZDf31EMtJMFwfUD</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a4778467b2e5b6ab3.output</output-file>
<status>completed</status>
<summary>Agent "Re-review final branch after fixes" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>I have everything I need. My analysis is complete.

## Review: general enum support (2cee306..42d88a3)

### Strengths

- **The Critical fix is correct.** I walked the emitted wasm in `compile_variant_constructor_arm` (`plum-wasm-codegen/src/lib.rs:986-1024`) instruction-by-instruction against all three edge conditions:
  - **(a) Payload-free subject (small int &lt; HEAP_BASE):** `LocalGet; I32Const(65536); I32GeU` yields 0, so the outer `If` takes the `Else` (`I32Const(0)`) branch and the `I32Load` is *never executed*. No matter what garbage sits at that low address, the arm cannot match and control falls to `rest`. The guard uses **unsigned** `I32GeU`, which is the right choice (pointers can exceed i32 signed range). Tags are assigned by `enumerate()` position, so they are always small and always `&lt; HEAP_BASE` — the invariant the guard relies on is guaranteed, and it's exactly the representation the design spec documents ("heap pointers are always ≥ 65536 and per-enum tags are small ints... can never collide").
  - **(b) Genuine payload pointer (≥ HEAP_BASE):** takes the `If` branch, loads the tag word at offset 0, compares to the expected tag; matches → binds fields + body, mismatches → yields 0 → `rest`. Correct.
  - **(c) Stack discipline is well-formed.** Inner `If(BlockType::Result(ValType::I32))` leaves exactly one i32 on both branches; outer `If(BlockType::Empty)` consumes it and both branches are net-zero statement blocks. `If`/`Else`/`End` nesting balances. `cargo test --workspace` is fully green (all crates), and the wasm-validator-backed tests (`assert_valid`, `output_validates`) pass, confirming no type-stack imbalance. This mirrors the existing discipline in `compile_variant_eq_arm` and the `Int` arm.
- The regression test `constructor_pattern_does_not_misfire_on_payload_free_sibling` (`unwrapOr(None, 5) == 5`) directly locks the fix, and `payload_variant_construction_compiles_and_runs` proves the pointer path is unaffected. TDD RED→GREEN is credible and I re-verified GREEN independently.
- Construction (`compile_variant_construction`) and destructuring use a **consistent** memory layout: tag at offset 0, fields at `(i+1)*8`, 8-byte stride matching class fields, width-correct i32/i64/f64 loads/stores. Bump allocations are 8-aligned (base 65536, sizes are multiples of 8), so field offsets 8/16/... are naturally aligned. Store side and load side agree.
- The class/enum-variant collision check (`plum-checker/src/lib.rs`, `check_source`) is a clean, targeted fix placed to avoid changing `build_global_tables`' signature, with a clear message and matching test. Both Important findings and the Minor interop finding are now covered by tests.
- README "Known gaps" is honest and now correctly documents the tail-position match/if return-value gap and the nested-constructor-pattern restriction.

### Issues

#### Critical (Must Fix)
None. The prior Critical finding's fix is genuinely correct, not merely guarded.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- **Collision diagnostic lives only in `check_source`, but `compile_source` calls `build_global_tables` directly** (`plum-wasm-codegen/src/lib.rs:293`), not `check_source`. Codegen documents the assumption that input was already accepted by the checker (line 284), so in the normal pipeline this is fine — but a caller that runs codegen on unchecked source would not get the class/enum collision diagnostic and would silently resolve to the enum variant. This is consistent with the checker being the single gate for *all* semantic checks, so it's an observation rather than a defect. No action strictly required.
- **Deferred `type_env` scoping (finding #5) remains, and the deferral reasoning is sound.** I confirmed the insert at `lib.rs:1018` happens immediately before each arm's `compile_block`, and every binding name maps to a distinct pre-allocated wasm local, so a stale `type_env` entry cannot cause a wrong local to be read/written — it only affects compile-time type-hint bookkeeping. Correctly assessed as cosmetic; fine to leave.

### Recommendations
- Consider a follow-up test with an enum that *interleaves* payload and payload-free variants (e.g. `enum E = A(Int) | B | C(Int)`), matching a `B` value against a `C(x)` constructor arm, to lock the tag-numbering-across-mixed-variants invariant explicitly. Current tests cover `Some/None` and `Rect/Circle` but not interleaved tags. Low priority — the logic already handles it correctly.

### Assessment

**Ready to merge?** Yes

**Reasoning:** The previously-Critical miscompile is fixed correctly at the instruction level — the `&gt;= HEAP_BASE` guard provably prevents any payload-free subject from satisfying a constructor-pattern tag test regardless of low-memory contents, the pointer path is preserved, and the `If`/`Else`/`End` nesting is stack-balanced and validator-clean. The cumulative implementation matches the design spec's runtime representation and scope, all workspace tests and the tree-sitter corpus pass, and the two remaining items are genuinely cosmetic/observational.</result>
<usage><subagent_tokens>64626</subagent_tokens><tool_uses>8</tool_uses><duration_ms>138292</duration_ms></usage>
</task-notification>

**Claude:** Ready to merge. Updating the ledger and moving to branch completion.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** All tasks complete and the final review is clean. Now finishing the branch.
  - `Skill: superpowers:finishing-a-development-branch`

**You:** Base directory for this skill: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/finishing-a-development-branch

# Finishing a Development Branch

## Overview

Guide completion of development work by presenting clear options and handling chosen workflow.

**Core principle:** Verify tests → Detect environment → Present options → Execute choice → Clean up.

**Announce at start:** "I'm using the finishing-a-development-branch skill to complete this work."

## The Process

### Step 1: Verify Tests

**Before presenting options, verify tests pass:**

```bash
# Run project's test suite
npm test / cargo test / pytest / go test ./...
```

**If tests fail:**
```
Tests failing (<N> failures). Must fix before completing:

[Show failures]

Cannot proceed with merge/PR until tests pass.
```

Stop. Don't proceed to Step 2.

**If tests pass:** Continue to Step 2.

### Step 2: Detect Environment

**Determine workspace state before presenting options:**

```bash
GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P)
GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P)
```

This determines which menu to show and how cleanup works:

| State | Menu | Cleanup |
|-------|------|---------|
| `GIT_DIR == GIT_COMMON` (normal repo) | Standard 4 options | No worktree to clean up |
| `GIT_DIR != GIT_COMMON`, named branch | Standard 4 options | Provenance-based (see Step 6) |
| `GIT_DIR != GIT_COMMON`, detached HEAD | Reduced 3 options (no merge) | No cleanup (externally managed) |

### Step 3: Determine Base Branch

```bash
# Try common base branches
git merge-base HEAD main 2>/dev/null || git merge-base HEAD master 2>/dev/null
```

Or ask: "This branch split from main - is that correct?"

### Step 4: Present Options

**Normal repo and named-branch worktree — present exactly these 4 options:**

```
Implementation complete. What would you like to do?

1. Merge back to <base-branch> locally
2. Push and create a Pull Request
3. Keep the branch as-is (I'll handle it later)
4. Discard this work

Which option?
```

**Detached HEAD — present exactly these 3 options:**

```
Implementation complete. You're on a detached HEAD (externally managed workspace).

1. Push as new branch and create a Pull Request
2. Keep as-is (I'll handle it later)
3. Discard this work

Which option?
```

**Don't add explanation** - keep options concise.

### Step 5: Execute Choice

#### Option 1: Merge Locally

```bash
# Get main repo root for CWD safety
MAIN_ROOT=$(git -C "$(git rev-parse --git-common-dir)/.." rev-parse --show-toplevel)
cd "$MAIN_ROOT"

# Merge first — verify success before removing anything
git checkout <base-branch>
git pull
git merge <feature-branch>

# Verify tests on merged result
<test command>

# Only after merge succeeds: cleanup worktree (Step 6), then delete branch
```

Then: Cleanup worktree (Step 6), then delete branch:

```bash
git branch -d <feature-branch>
```

#### Option 2: Push and Create PR

```bash
# Push branch
git push -u origin <feature-branch>
```

**Do NOT clean up worktree** — user needs it alive to iterate on PR feedback.

#### Option 3: Keep As-Is

Report: "Keeping branch <name>. Worktree preserved at <path>."

**Don't cleanup worktree.**

#### Option 4: Discard

**Confirm first:**
```
This will permanently delete:
- Branch <name>
- All commits: <commit-list>
- Worktree at <path>

Type 'discard' to confirm.
```

Wait for exact confirmation.

If confirmed:
```bash
MAIN_ROOT=$(git -C "$(git rev-parse --git-common-dir)/.." rev-parse --show-toplevel)
cd "$MAIN_ROOT"
```

Then: Cleanup worktree (Step 6), then force-delete branch:
```bash
git branch -D <feature-branch>
```

### Step 6: Cleanup Workspace

**Only runs for Options 1 and 4.** Options 2 and 3 always preserve the worktree.

```bash
GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P)
GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P)
WORKTREE_PATH=$(git rev-parse --show-toplevel)
```

**If `GIT_DIR == GIT_COMMON`:** Normal repo, no worktree to clean up. Done.

**If worktree path is under `.worktrees/` or `worktrees/`:** Superpowers created this worktree — we own cleanup.

```bash
MAIN_ROOT=$(git -C "$(git rev-parse --git-common-dir)/.." rev-parse --show-toplevel)
cd "$MAIN_ROOT"
git worktree remove "$WORKTREE_PATH"
git worktree prune  # Self-healing: clean up any stale registrations
```

**Otherwise:** The host environment (harness) owns this workspace. Do NOT remove it. If your platform provides a workspace-exit tool, use it. Otherwise, leave the workspace in place.

## Quick Reference

| Option | Merge | Push | Keep Worktree | Cleanup Branch |
|--------|-------|------|---------------|----------------|
| 1. Merge locally | yes | - | - | yes |
| 2. Create PR | - | yes | yes | - |
| 3. Keep as-is | - | - | yes | - |
| 4. Discard | - | - | - | yes (force) |

## Common Mistakes

**Skipping test verification**
- **Problem:** Merge broken code, create failing PR
- **Fix:** Always verify tests before offering options

**Open-ended questions**
- **Problem:** "What should I do next?" is ambiguous
- **Fix:** Present exactly 4 structured options (or 3 for detached HEAD)

**Cleaning up worktree for Option 2**
- **Problem:** Remove worktree user needs for PR iteration
- **Fix:** Only cleanup for Options 1 and 4

**Deleting branch before removing worktree**
- **Problem:** `git branch -d` fails because worktree still references the branch
- **Fix:** Merge first, remove worktree, then delete branch

**Running git worktree remove from inside the worktree**
- **Problem:** Command fails silently when CWD is inside the worktree being removed
- **Fix:** Always `cd` to main repo root before `git worktree remove`

**Cleaning up harness-owned worktrees**
- **Problem:** Removing a worktree the harness created causes phantom state
- **Fix:** Only clean up worktrees under `.worktrees/` or `worktrees/`

**No confirmation for discard**
- **Problem:** Accidentally delete work
- **Fix:** Require typed "discard" confirmation

## Red Flags

**Never:**
- Proceed with failing tests
- Merge without verifying tests on result
- Delete work without confirmation
- Force-push without explicit request
- Remove a worktree before confirming merge success
- Clean up worktrees you didn't create (provenance check)
- Run `git worktree remove` from inside the worktree

**Always:**
- Verify tests before offering options
- Detect environment before presenting menu
- Present exactly 4 options (or 3 for detached HEAD)
- Get typed confirmation for Option 4
- Clean up worktree for Options 1 & 4 only
- `cd` to main repo root before worktree removal
- Run `git worktree prune` after removal

**Claude:** I'm using the finishing-a-development-branch skill to complete this work.
  - `Bash: cargo test --workspace 2>&1 | tail -20 && echo "---tree-sitter---" && cd tooling/tree-sitter-plum && npx --yes tree-sitter-cli test 2>&1 | tail -10`

**Claude:** Tests pass. This is a normal repo checkout (no worktree — the user declined earlier), and we're already directly on `master` with all commits landed there.
  - `Bash: GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P); GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/dev/null && pwd -P); echo "$GIT_DIR / $`
  - `Bash: git remote -v; git log --oneline -1; git status --short`

**Claude:** Everything is already committed directly to `master` (no separate branch was created, since you opted to work in place). There's nothing to merge locally — the remaining choice is just whether to push.
  - `AskUserQuestion`

**Claude:** General enum support is complete — all 10 commits are on `master` locally, tests fully green (workspace + tree-sitter corpus), left unpushed per your choice.

**Summary of what shipped:**
- Grammar: positional variant-construction calls (`Some(v)`) now parse as expressions
- Checker: per-variant tags/field types, fixed a bare-tag typing bug, real construction/pattern validation
- Codegen: variant construction (bump-heap `[tag][fields...]` layout) and match lowering (tag + constructor patterns) for any concrete enum
- A real silent miscompile was caught by the final review and fixed: matching a payload-free variant like `None` through a `Some(v)`-first arm used to return wrong results — now guarded correctly
- Two unrelated pre-existing gaps were discovered along the way and spawned as separate follow-up tasks rather than fixed here: a grammar limitation on multi-line body trailing statements, and a codegen gap where tail-position `match`/`if` without explicit `return` silently drops its value

Generics monomorphization (needed for the actual `libs/std` `Option`/`Result`/`List`/`Map`) remains the next big piece whenever you want to tackle it.

**You:** Yes please tackle that and fix the gaps, Two unrelated pre-existing gaps were discovered along the way and spawned as separate follow-up tasks rather than fixed here: a grammar limitation on multi-line body trailing statements, and a codegen gap where tail-position `match`/`if` without explicit `return` silently drops its value
  - `Skill: superpowers:brainstorming`

**You:** Skill /superpowers:brainstorming is already loaded above; instructions unchanged.

**Claude:** This is really three separate pieces of work: two small, already-diagnosed bug fixes, and one large feature (generics monomorphization) that needs its own design discussion. Let me confirm the split before diving in.
  - `AskUserQuestion`

**Claude:** Let me look at the current code for both gaps before designing the fix.
  - `Bash: grep -n "_statement\|fn:\|body:" tooling/tree-sitter-plum/grammar.js | head -40`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`
  - `Bash: grep -n "expression:\|primary_expression:" tooling/tree-sitter-plum/grammar.js`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`

**Claude:** Both root causes are clear now. Let me present a compact design for both fixes (Gap A is a one-line grammar change; Gap B is more involved).

**Gap A — grammar trailing-statement limitation**

`tooling/tree-sitter-plum/grammar.js`'s `_statement` rule (line 166-179) lists `$.primary_expression` as its last alternative instead of the full `$.expression` — so a bare comparison/boolean/ternary expression used as a statement (most commonly the final line of an indented body) doesn't parse. `$.expression` is already a strict superset of `$.primary_expression` (grammar.js:261-269) and is already used elsewhere (assign, assert, return, if-condition) without issue, so this is a one-line swap: `$.primary_expression``$.expression` in `_statement`, regenerate, add corpus tests for the previously-broken forms (bare comparison, boolean-op, ternary as a tail statement).

**Gap B — tail-position match/if drops its value**

`compile_block_as_fn_body` only preserves a function's return value when the *literal* last statement is a bare `Stmt::Expr`. A `Stmt::If`/`Stmt::Match` in that position compiles its arms as ordinary statement blocks (`BlockType::Empty`), so any bare tail expression inside an arm gets `Drop`ped — producing wasm that fails `wasmparser::validate` (a hard, loud failure, not a silent wrong answer — but `plum-cli`'s `compile` command writes the bytes to disk without validating, so today a user hitting this gets a corrupt `.wasm` file with no error message).

Fix: add a recursive "value position" compile path for `Stmt::If`/`Stmt::Match`, used when they're the function body's tail statement (mirroring the existing `Stmt::Expr` special case) — each arm's own last statement is compiled the same way, recursively (so `return` and a bare tail expression can coexist across arms/branches: `return` is stack-polymorphic in wasm, so mixing it with a branch that leaves a value on the stack validates fine). If any arm's tail statement still isn't return/expr/nested-if/nested-match (some genuinely malformed shape), emit a clear compile-time `Err` instead of ever emitting bytes that could fail validation silently.
  - `AskUserQuestion`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-tail-position-and-grammar-gaps-design.md`

**Claude:** Self-review: no placeholders, sections are internally consistent, scope is bounded with an explicit out-of-scope list. Committing.
  - `Bash: git add docs/superpowers/specs/2026-07-20-tail-position-and-grammar-gaps-design.md && git commit -m "$(cat <<'EOF' docs: add design spec for grammar trailing-st`

**Claude:** Spec written and committed to [2026-07-20-tail-position-and-grammar-gaps-design.md](docs/superpowers/specs/2026-07-20-tail-position-and-grammar-gaps-design.md). Please review it before I write the implementation plan.

**You:** Yes
  - `Skill: superpowers:writing-plans`

**You:** Skill /superpowers:writing-plans is already loaded above; instructions unchanged.
  - `Bash: grep -n "fn stmt_always_diverges\|fn block_always_diverges\|fn compile_block_as_fn_body\|fn compile_fn_body\|fn compile_stmt\b\|fn compile_block\b\|fn compile_m`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Bash: grep -n "fn ret_type_to_wasm\|fn ast_type_to_wasm\|fn compile_variant_eq_arm\|fn compile_variant_constructor_arm" plum-wasm-codegen/src/lib.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Bash: grep -n "stmt_always_diverges\|block_always_diverges\|has_return_value\|compile_block_as_fn_body\|compile_match(\|compile_match_arms(\|compile_variant_eq_arm(\|`
  - `Bash: cat examples/match.plum; echo ---README---; grep -n "tail-position\|Known gaps" -A10 README.md | head -20`
  - `Read: /Users/pyrossh/Code/plum/README.md`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-tail-position-and-grammar-gaps.md`

**Claude:** Self-review complete: every spec requirement maps to a task, function signatures are consistent across all steps, no placeholders. Plan saved to [docs/superpowers/plans/2026-07-20-tail-position-and-grammar-gaps.md](docs/superpowers/plans/2026-07-20-tail-position-and-grammar-gaps.md).

Two execution options:

1. **Subagent-Driven (recommended)** — fresh subagent per task, review between tasks, fast iteration
2. **Inline Execution** — execute tasks in this session with checkpoints

Which approach?

**You:** 1
  - `Skill: superpowers:subagent-driven-development`

**You:** Skill /superpowers:subagent-driven-development is already loaded above; instructions unchanged.

**Claude:** Continuing to work directly on `master` (consistent with your earlier choice), tests already confirmed green. Setting up the ledger for this new plan.
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** That plan is fully complete — starting a fresh ledger for the new one.
  - `Write: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Pre-flight scan: I authored this plan and already cross-checked consistency during writing-plans self-review — no conflicts found. Dispatching Task 1.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-t`
  - `Agent: Implement Task 1: grammar trailing-statement fix`

**Claude:** Task 1 implementer running in the background.

**You:** <task-notification>
<task-id>a7b9f9b486b63ffc9</task-id>
<tool-use-id>toolu_01XsPzrrFjMmV81FMo9anpae</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a7b9f9b486b63ffc9.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 1: grammar trailing-statement fix" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Perfect! Task 1 is complete. Here's the summary:

---

**Status:** DONE

**Commit created:** `aeaba6f` - fix(tree-sitter-plum): allow a full expression as a body's trailing statement

**Test summary:** All 21 corpus tests pass (18 original + 3 new, 100% success rate); all Rust workspace tests pass (no failures).

**No concerns.** The grammar change correctly allows comparison operators, boolean operators, and ternary expressions as trailing statements in function bodies. All new test cases generate valid parse trees with no ERROR or MISSING nodes.

**Report file:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md`</result>
<usage><subagent_tokens>45228</subagent_tokens><tool_uses>30</tool_uses><duration_ms>177255</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 1 reviewer.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 1ccf9bd`
  - `Agent: Review Task 1 (spec + quality)`

**Claude:** Task 1 reviewer running in the background.

**You:** <task-notification>
<task-id>a45144bc3431fd3a6</task-id>
<tool-use-id>toolu_01929urKJKZw7c793j2Xbufb</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a45144bc3431fd3a6.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 1 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>### Spec Compliance
- ✅ Spec compliant. `tooling/tree-sitter-plum/grammar.js:178` changes exactly `$.primary_expression``$.expression` in the `_statement` rule, matching the brief's prescribed diff verbatim (confirmed in diff lines 36-38). No other grammar rule was touched.
- The 3 new corpus cases in `test/corpus/function.txt` (diff lines 153-248) match the brief's Step 3 input verbatim (`isNone`/`o == None`, `bothTrue`/`a &amp;&amp; b`, `pick`/`cond ? a : b`) and each produces a clean tree with no `ERROR`/`MISSING` node: `(expression (comparison_operator ...))`, `(expression (boolean_operator ...))`, `(expression (ternary_expression ...))`, each a single complete node directly inside `body`.
- Verified structural consistency against `grammar.js`'s existing rule definitions (lines 293-311, 351-359, 434-438): `comparison_operator` takes bare `$.primary_expression` operands (no `expression` wrapper) while `boolean_operator` and `ternary_expression` take `$.expression`-wrapped operands — the generated trees correctly reflect this pre-existing asymmetry, it is not a bug introduced by this task.
- Report claims "no Rust code changes needed." Verified directly: `git log 1ccf9bd..aeaba6f` shows one commit touching only `tooling/tree-sitter-plum/*`; `plum-core/src/parser.rs` already contains a generic `unwrap_expr_node` (peels `"expression"` and `"primary_expression"` wrapper kinds recursively, lines 18-25) applied at the top of `parse_stmt` (line 299) before dispatch, plus a catch-all arm `kind if is_expression_kind(kind) =&gt; Stmt::Expr(...)` (line 322) where `is_expression_kind` (line 758) already lists `comparison_operator`, `boolean_operator`, `ternary_expression`. So both the pre-existing tree-shape change (extra `expression` wrapper around all former bare-`primary_expression` statements) and the newly-reachable node kinds are handled without any Rust changes — confirms the report's claim and the brief's stated interface contract.

### Strengths
- Genuinely minimal grammar diff: one line changed, nothing else in `grammar.js`.
- Binary-diffed generated artifacts (`grammar.json`, `node-types.json`, `parser.c`) are consistent with a `tree-sitter generate` byproduct, not hand-edited — no textual diff to review, and the report explicitly documents running `generate`.
- The 8 pre-existing corpus files that changed (`enum.txt`, `for.txt`, `if.txt`, `literals.txt`, `match.txt`, `type.txt`, `while.txt`) all show a uniform, mechanical transformation: every previously-bare `(primary_expression ...)` statement is now wrapped one level deeper as `(expression (primary_expression ...))`, with no other structural change — consistent with the single grammar-rule edit and not evidence of scope creep.
- Report is transparent about the wider snapshot churn (8 files, not just the 3 new cases) and explains the root cause correctly (statements dispatched via `$.expression` now materialize an explicit wrapper node).

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- The brief's Step 2 only anticipated "all pre-existing corpus cases still pass" without flagging that 8 snapshot files would need `-u` regeneration due to the new wrapper node; the implementer diverged from the literal step sequence (ran `-u` broadly instead of only for the 3 new cases) but this was the only viable path given the grammar's actual behavior, and the report calls it out honestly rather than hiding it.

### Assessment
**Task quality:** Approved
**Reasoning:** The change is exactly the one-line grammar edit specified, the 3 new corpus cases parse cleanly and match the grammar's existing (asymmetric) operand-wrapping conventions, and independent inspection of `plum-core/src/parser.rs` confirms the zero-Rust-change claim is correct — the generic `unwrap_expr_node`/`is_expression_kind` machinery already absorbs both the new node kinds and the incidental wrapper-node tree-shape change across the other 8 corpus files.</result>
<usage><subagent_tokens>56133</subagent_tokens><tool_uses>16</tool_uses><duration_ms>106376</duration_ms></usage>
</task-notification>

**Claude:** Task 1 approved. Updating the ledger and moving to Task 2 — the larger codegen fix.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-t`
  - `Agent: Implement Task 2: tail-position value propagation`

**Claude:** Task 2 (the larger codegen fix) implementer running in the background.

**You:** <task-notification>
<task-id>a53bf4b3aff0c2168</task-id>
<tool-use-id>toolu_013wWcr89DTUQ7cU8n94Wg7G</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a53bf4b3aff0c2168.output</output-file>
<status>completed</status>
<summary>Agent "Check plum-core statement dispatch for expression node kind" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>This directly answers the question — I have everything needed.

Found in `/Users/pyrossh/Code/plum/plum-core/src/parser.rs`.

**1. The generic unwrap helper** (lines 17–25):
```rust
/// Peel transparent `expression` / `primary_expression` wrapper nodes.
fn unwrap_expr_node&lt;'b&gt;(&amp;self, node: Node&lt;'b&gt;) -&gt; Node&lt;'b&gt; {
    match node.kind() {
        "expression" | "primary_expression" =&gt; {
            node.named_child(0).map(|c| self.unwrap_expr_node(c)).unwrap_or(node)
        }
        _ =&gt; node,
    }
}
```
This already unifies "expression" and "primary_expression" in one arm, recursing to peel nested wrapper layers (handles the case where "expression" wraps another "expression" or a "primary_expression").

**2. The statement dispatcher** `parse_stmt` (lines 298–325):
```rust
fn parse_stmt(&amp;self, node: Node) -&gt; Option&lt;Stmt&gt; {
    let node = self.unwrap_expr_node(node);
    Some(match node.kind() {
        "assign" =&gt; Stmt::Assign(self.parse_assign(node)),
        "break" =&gt; Stmt::Break,
        "continue" =&gt; Stmt::Continue,
        "return" =&gt; { ... }
        "todo" =&gt; Stmt::Todo,
        "assert" =&gt; { ... }
        "for" =&gt; Stmt::For(self.parse_for(node)),
        "while" =&gt; Stmt::While(self.parse_while(node)),
        "if" =&gt; Stmt::If(self.parse_if(node)),
        "match" =&gt; Stmt::Match(self.parse_match(node)),
        kind if is_expression_kind(kind) =&gt; Stmt::Expr(self.parse_expression(node)),
        _ =&gt; return None,
    })
}
```

Note `parse_stmt` calls `unwrap_expr_node(node)` on entry (line 299), before the match — so an incoming child of kind `"expression"` is peeled down to its inner node (e.g. `comparison_operator`, `boolean_operator`, `ternary_expression`, or `primary_expression`) before dispatch ever happens. The match itself has no explicit `"expression"` arm, but it doesn't need one because of this pre-unwrap step, and the fallback `is_expression_kind(kind)` catch-all (line 322) handles whatever concrete expression kind remains after unwrapping.

**3. Verdict:** Yes, the code is already correct for the grammar change. Since `parse_stmt` unconditionally unwraps the top-level node (including new `"expression"` wrapper nodes) before matching on `.kind()`, a trailing bare `a == b`, `a &amp;&amp; b`, or `cond ? a : b` statement — now wrapped in an `"expression"` node — will be unwrapped to its real inner kind (`comparison_operator`/`boolean_operator`/`ternary_expression`) and then hit the `is_expression_kind(kind)` fallback arm, producing `Stmt::Expr(...)` exactly as before. Old bare `primary_expression` statements are likewise still handled: `unwrap_expr_node` treats `"primary_expression"` the same way as `"expression"`, recursively peeling until a non-wrapper kind is reached. No panic or unhandled-kind path is hit; both old and new grammar shapes converge to the same behavior via the shared `unwrap_expr_node` helper used at the top of `parse_stmt`.</result>
<usage><subagent_tokens>27190</subagent_tokens><tool_uses>5</tool_uses><duration_ms>156279</duration_ms></usage>
</task-notification>

**Claude:** That notification was from an earlier, already-resolved sub-check (Task 1's review). Continuing to wait on Task 2's implementer.

**You:** <task-notification>
<task-id>ab983613718859f48</task-id>
<tool-use-id>toolu_0155p7X8CNbuJLv9Cf9tD85y</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/ab983613718859f48.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 2: tail-position value propagation" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE_WITH_CONCERNS

Commit: `3254688` — fix(plum-wasm-codegen): propagate a value through tail-position match/if without explicit return

Test summary: `cargo test -p plum-wasm-codegen --test codegen_tests` 41/41 passed (8 new + 33 pre-existing); `cargo test --workspace` all green; `tree-sitter-cli test` 21/21.

Concern: transcribing brief Step 3i verbatim broke 9 pre-existing tests (enum matches without a trailing wildcard, e.g. `Some`/`None`, `True`/`False`, each arm using explicit `return`) — the brief's own code treats "ran out of patterns" as an unconditional error in value position, which regresses those. Fixed with a small addition inside `compile_match` only (not touching any of the specified function signatures): detect when the listed patterns already cover every variant of an enum, and in that case append a synthetic `_ =&gt; todo` case before recursing, so the fallback traps (valid, unreachable) instead of erroring. This is a deviation from "verbatim transcription," documented in detail in the report.

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md`</result>
<usage><subagent_tokens>125432</subagent_tokens><tool_uses>47</tool_uses><duration_ms>665503</duration_ms></usage>
</task-notification>
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package aeaba6f`
  - `Bash: grep -n "synthetic\|todo\|already cover\|covers_every\|is_exhaustive\|fn compile_match\b" -A5 -B2 /Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md | h`
  - `Agent: Review Task 2 (spec + quality)`

**Claude:** Task 2's implementer found a real regression risk in my plan's exact prescribed code (non-exhaustive enum matches without a trailing wildcard would have wrongly errored) and worked around it with a targeted heuristic. Dispatched the reviewer with heavy scrutiny on that specific fix's correctness before accepting it.

**You:** <task-notification>
<task-id>a293f1c47614c7a1e</task-id>
<tool-use-id>toolu_01BpQxEiciA9C5B12cxY5mJg</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a293f1c47614c7a1e.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 2 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>All confirmed. Case/Block/Stmt derive Clone as used.

### Spec Compliance
- ✅ Spec compliant, with one documented deviation from brief Step 3i's exact code, evaluated below on technical merit rather than accepted on faith.
- All other steps (3a–3h, 3j–3k) match the brief's prescribed code essentially verbatim: `block_type_for` (lib.rs:257-259 in new file), removal of `stmt_always_diverges`/`block_always_diverges`, the four new helpers (`compile_case_body`, `compile_block_in_value_position`, `compile_stmt_in_value_position`, `compile_if`), simplified `compile_block_as_fn_body`, `compile_fn_body`'s `ret_type_to_wasm` call site, `compile_stmt`'s `If`/`Match` arms calling `compile_if`/`compile_match(..., None)`, and `result_vt: Option&lt;ValType&gt;` threaded through `compile_match`/`compile_match_arms`/`compile_variant_eq_arm`/`compile_variant_constructor_arm` in the exact parameter positions/order the brief specifies.
- `compile_match_arms`'s "no more cases" fallback deviates from the brief's unconditional `Err`: `compile_match` (lib.rs, `match_covers_every_enum_variant` + synthetic-case append) intercepts before calling `compile_match_arms` when `result_vt.is_some()` and the cases provably cover every enum variant. Verified the claim that no signatures changed beyond what the brief specifies — the new logic lives entirely inside `compile_match`, calling the unmodified-signature `compile_match_arms` with one extra synthetic case appended to a cloned `Vec`.

### Scrutiny of the deviation (all five points independently verified against the diff, not the report's narrative)

1. **Distinct-tag counting, not pattern count**: `tags_seen: BTreeSet&lt;i32&gt;` (dedup by `.insert`), compared via `tags_seen.len() == total_variants`. Traced a concrete adversarial case: enum `A(0)/B(1)/C(2)`, cases `[A, A, B]` (redundant duplicate arm for `A`, `C` genuinely missing) → `tags_seen = {0,1}`, len 2 ≠ 3 → correctly returns `false`. No false positive from duplicate arms.
2. **Rejects enum-tag patterns mixed with a binding/wildcard**: confirmed via the pattern-matching guard — `Wildcard` and lowercase `Name` both fall through to the `_ =&gt; return false` arm. Also confirmed via `compile_match_arms` that a `Wildcard`/binding arm never recurses into `rest` (comment: "any following cases are unreachable" / "always matches"), so the "ran out of patterns" branch is genuinely unreachable in that case regardless — the heuristic not firing there is correct and matches existing (already-working) behavior.
3. **Gated on `result_vt.is_some()`**: confirmed at the single call site — `if result_vt.is_some() &amp;&amp; match_covers_every_enum_variant(...)`. Ordinary statement position (`None`) is untouched, still hits `Ok(())` on empty cases.
4. **Rejects cross-enum / non-variant-pattern mixes**: `enum_name` mismatch check (`Some(en) if en != info.enum_name =&gt; return false`) rejects multi-enum case lists; `Int`/`String`/`Float` patterns fall to the same catch-all `_ =&gt; return false` since only `Name`(variant)/`Class` patterns are matched positively.
5. **Synthetic case safety**: `ast::Case`/`ast::Block`/`ast::CasePattern`/`ast::Stmt::Todo` all derive `Clone`/are freshly constructed (`cases.push(ast::Case { patterns: vec![CasePattern::Wildcard], body: Block { stmts: vec![Stmt::Todo] } })`) on a locally cloned `Vec` — not parsed from source, not shared/aliased, cannot collide with a user's own `Todo`. `Stmt::Todo` compiles to `Instruction::Unreachable` (pre-existing, unchanged, lib.rs:852-855) and is already one of the "compiles normally" arms in `compile_stmt_in_value_position` — verified both facts directly rather than trusting the report.
6. **Tag uniqueness precondition**: verified in `plum-checker/src/lib.rs:51-53,78-79` that `EnumVariantInfo.tag` is a unique per-variant `i32` within an `enum_name` (e.g. `True`/`False` = 1/0) — this is the invariant the distinct-tag-count comparison depends on, and it holds.
7. **Necessity check**: traced why the literal brief code breaks 9 tests — an exhaustive `Some(v) =&gt; return v; None =&gt; return default` match with no wildcard still recurses `compile_match_arms` down to an empty `rest` inside the final `Else` branch (unreachable at runtime, but still code-generated and thus subject to the literal Step 3i unconditional `Err`). Confirms the report's justification is real, not a rationalization.
8. **Scope**: this is a narrow, single-purpose heuristic (bare variant/constructor patterns only, single enum, exact distinct-tag-count match) — it does not attempt reachability analysis, guards, nested patterns, or multi-subject matches, and errs toward `Err` (the safe/conservative direction) on anything it can't prove. Consistent with "no full exhaustiveness analysis" — it decides one narrow, decidable question (do these top-level tag patterns partition the whole enum), not general match completeness.

No false positive or false negative found despite deliberately trying to construct one (redundant-arm-with-missing-variant, cross-enum, mixed wildcard/binding, non-I32 subject).

### Strengths
- Error messages consistently use the `"codegen: ..."` prefix (verified across all new error sites: if-without-else, non-exhaustive match, empty branch, and the generic value-position fallback).
- Doc comments on every new function are accurate and match actual behavior (checked `compile_stmt_in_value_position`'s comment on stack polymorphism against the real Return/Todo handling).
- The deviation is flagged transparently in the report rather than silently introduced, and is genuinely necessitated by a real conflict between the brief's literal Step 3i code and its own Step 4 acceptance criteria (verified independently, not taken on faith).
- Test additions (codegen_tests.rs) match the brief's Step 1 verbatim; assertions correctly target the `"doesn't produce a return value"` substring.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- No direct unit test targets `match_covers_every_enum_variant`'s edge cases (duplicate-arm-with-missing-variant, cross-enum mix) — current coverage is indirect via the 9 pre-existing tests plus the 3 new "clear error" tests. A small dedicated test (e.g. `match A =&gt; 1, A =&gt; 2` on a 3-variant enum, expecting the non-exhaustive error) would make this guarantee explicit rather than implicit, given how load-bearing the correctness of this helper is.
- `match_covers_every_enum_variant`'s `total_variants` computation (`ctx.enum_variants.values().filter(|v| v.enum_name == en).count()`) does an O(n) scan of all enum variants in the program per match compiled; harmless at current scale, not worth changing.

### Assessment
**Task quality:** Approved
**Reasoning:** All prescribed edits match the brief; the one deviation from Step 3i's literal code is well-motivated (demonstrated by tracing the actual recursion that breaks 9 pre-existing tests), and its replacement logic (`match_covers_every_enum_variant` + synthetic `_ =&gt; todo` case) was independently verified against every failure mode in the scrutiny list — distinct-tag counting, gating on `result_vt`, cross-enum/non-variant rejection, and synthetic-case AST safety — with no false positive or false negative found.</result>
<usage><subagent_tokens>68728</subagent_tokens><tool_uses>8</tool_uses><duration_ms>118573</duration_ms></usage>
</task-notification>

**Claude:** Task 2 approved — the deviation was rigorously verified correct. Updating the ledger and moving to Task 3.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-t`

**Claude:** Task 3 is small and mechanical (transcription + testing) — using the cheap tier.
  - `Agent: Implement Task 3: restore idiomatic examples/docs`

**Claude:** Task 3 implementer running in the background.

**You:** <task-notification>
<task-id>aa26cf0e42b2c4dd5</task-id>
<tool-use-id>toolu_01FEj2YTYrL9Wf4tMaLZC5Fz</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/aa26cf0e42b2c4dd5.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 3: restore idiomatic examples/docs" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Perfect! Task 3 is complete. Here's the summary:

**Status:** DONE

**Commit created:**
- `87b54a9` — docs+test: restore examples/match.plum to idiomatic bare-tail style; tail-position gap fixed

**Test results:**
-`cargo test --workspace`: 84 tests passed (0 failed)
- ✅ Tree-sitter corpus: 21 parses successful (0 failed)
-`match_example_compiles_and_runs_correctly` passes with bare-tail-expression arms

**Files changed:**
- `/Users/pyrossh/Code/plum/examples/match.plum` — reverted all match arms to bare-expression style (removed `return` keywords)
- `/Users/pyrossh/Code/plum/README.md` — removed the now-fixed tail-position gap bullet from Known Gaps

**Self-review findings:** None — Task 2's fix works correctly with bare-tail expressions, the revert is clean, and all tests pass. This completes the full plan (Tasks 1, 2, and 3).

**Report:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-3-report.md`</result>
<usage><subagent_tokens>41440</subagent_tokens><tool_uses>14</tool_uses><duration_ms>100509</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 3 reviewer.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 3254688`
  - `Agent: Review Task 3 (spec + quality)`

**Claude:** Task 3 reviewer running in the background.

**You:** <task-notification>
<task-id>a964036782244b4bf</task-id>
<tool-use-id>toolu_013ZNqYnwcnGUZwrUgnnbcFY</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a964036782244b4bf.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 3 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>### Spec Compliance
✅ Spec compliant — verified line-by-line against the brief's prescribed content.

**`examples/match.plum`**: diff shows exactly the five `return X``X` reversions (describeNumber's three arms, describeBool's two, bindExample's one, describeColor's three, describeOption's two) with no other line changes. Resulting file content matches the brief's prescribed block verbatim, including `main() -&gt; Int = describeOption(Some(5))` (unchanged, diff context confirms it was already bare and untouched).

**`README.md`**: diff removes exactly one line — the `match`/`if` tail-position workaround bullet — from the Known Gaps list. The three remaining bullets (string interpolation, multi-subject match, user-defined generics) plus the fourth (nested constructor patterns) are left intact as context lines, with the trailing `closure` paragraph and blank-line spacing preserved. No stray double-blank-line or list-formatting artifact introduced.

### Strengths
- Diff is a pure, minimal revert — zero scope creep, matches "straightforward revert, not new design" instruction from the plan.
- Report's claimed test results (84 workspace tests, 21/21 tree-sitter corpus) are consistent with a docs+example-only change; nothing in the diff touches grammar, codegen, or checker code that could plausibly regress those suites.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
None.

### Assessment
**Task quality:** Approved
**Reasoning:** The diff matches the brief's prescribed file contents exactly (both `examples/match.plum` and the README Known Gaps edit), the example now genuinely demonstrates the plan's two fixes (bare-expression trailing statement in grammar; tail match/if value propagation without explicit `return`), and the Known Gaps section post-diff accurately reflects only the gaps still known to be true. This closes out the plan coherently.</result>
<usage><subagent_tokens>35692</subagent_tokens><tool_uses>3</tool_uses><duration_ms>15654</duration_ms></usage>
</task-notification>

**Claude:** Task 3 approved. All 3 tasks complete — dispatching the final whole-branch review.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: git log --oneline 1ccf9bd~1..87b54a9 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 42d88a3 87b54a9`
  - `Agent: Final whole-branch review of tail-position and grammar fixes`

**Claude:** Final whole-branch review running in the background on Opus, with heavy scrutiny on the `match_covers_every_enum_variant` deviation.

**You:** <task-notification>
<task-id>afce8b9b38333c6be</task-id>
<tool-use-id>toolu_01JmvVqW6yXkpmoxFbGyhv1d</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/afce8b9b38333c6be.output</output-file>
<status>completed</status>
<summary>Agent "Final whole-branch review of tail-position and grammar fixes" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>I have everything I need. All 8 new tests pass, the full workspace suite is green (from the fresh run), the tree-sitter corpus is 21/21, the removed helpers have zero remaining references, tags are provably unique-per-enum (sequential `enumerate()`), the checker performs no exhaustiveness analysis (so the codegen defensive errors are genuinely reachable), and plum-core's `unwrap_expr_node` peels both wrapper kinds. Here is my review.

---

### Strengths

- **The `match_covers_every_enum_variant` deviation is genuinely correct** — I verified all four attack vectors from your brief independently (details below). This is the riskiest part of the change and it holds up.
- **Clean, minimal threading of `Option&lt;ValType&gt;`.** The `result_vt: None` path is byte-for-byte the old statement-position behavior (each `If`/arm uses `BlockType::Empty`, bodies via `compile_block`), so non-tail and Unit-returning `if`/`match` are provably unchanged. The `Some(vt)` path is purely additive.
- **Good decomposition.** `block_type_for`, `compile_case_body`, `compile_block_in_value_position`, `compile_stmt_in_value_position`, and `compile_if` each have one clear job; `compile_if` correctly de-duplicates the old inline `Stmt::If` codegen (single source of truth now).
- **Dead-code removal is safe.** `stmt_always_diverges`/`block_always_diverges` have zero remaining references anywhere in the workspace (grep confirmed). Their `Unreachable`-at-`end` trick is fully subsumed by wasm stack-polymorphism in the new design.
- **Tests exercise real behavior.** The five positive tests assert concrete `run_main` results via wasmtime execution (which itself validates), not just `compile` success. The three error tests are properly distinct: if-without-else, non-exhaustive-match, and non-value tail statement each hit a different code path and a differently-parenthesized error message.
- **Grammar fix + corpus regeneration is legitimate.** The wider-than-planned corpus churn (enum/for/if/literals/match/type/while `.txt`) is a pure mechanical snapshot update: every trailing statement now carries one extra `(expression …)` wrapper node. plum-core's `unwrap_expr_node` (parser.rs:17-20) peels both `expression` and `primary_expression`, so no Rust change was needed — empirically confirmed by the full green suite.

### Issues

#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)

- **`plum-cli compile` still writes unvalidated bytes** (`plum-cli/src/main.rs`) — explicitly out of scope per the spec, and moot for *this* class of bug since the root cause is gone. But the README's blanket promise that codegen "reports a clear error rather than silently producing wrong code" is only fully true because these specific paths now error; a future codegen bug would still reach disk silently. Worth the separate follow-up the spec already names. Not a blocker.
- **Latent (pre-existing, not introduced here):** a tail bare `Stmt::Expr` whose expression is `Unit`-typed in a non-`Unit` function would leave nothing on the stack and produce invalid wasm that `compile_source` returns `Ok` on. This is guarded upstream by the checker's return-type check and predates this change (the old simple-tail-expr case had the identical shape), so it is not a regression — noting only for completeness.

### Recommendations

- Accept both deviations. Task 2's `match_covers_every_enum_variant` is the right call, and Task 1's corpus regeneration is unavoidable mechanical churn.
- File the `plum-cli` output-validation follow-up so the README guarantee becomes structurally true rather than incidentally true.

### `match_covers_every_enum_variant` — verification of the four risks

- **(a) duplicate arms masking a missing variant — SAFE.** Tags are assigned `tag as i32` via `enumerate()` per enum (checker lib.rs:92-98), so within one enum distinct variants have distinct tags, and the function rejects any case whose variant belongs to a different enum. `tags_seen` is a `BTreeSet&lt;i32&gt;`, so duplicate arms for one variant collapse to a single tag and cannot inflate the count. `A, B, B` over a 3-variant enum → `{0,1}`, `2 != 3``false` → correct non-exhaustive error. Counting distinct tags is exactly counting distinct covered variants.
- **(b) variant patterns mixed with wildcard/binding — SAFE.** A `CasePattern::Wildcard`, or a lowercase/`Name` binding, falls to the `_ =&gt; return false` arm, so the heuristic declines to fire. That's correct: such an arm is itself a catch-all, so `compile_match_arms` never runs out of patterns and never emits the spurious error the heuristic exists to prevent.
- **(c) synthetic `_ =&gt; todo` colliding with a real arm — SAFE.** It fires only when *every* user arm is a bare variant/constructor pattern (no wildcard), and it is appended strictly after them. It is reached only if no real tag matched — impossible at runtime for an exhaustive match — and `Stmt::Todo` lowers to `Unreachable`, which is stack-polymorphic and satisfies the enclosing `BlockType::Result(vt)` without a real value. No shadowing possible.
- **(d) firing on a genuinely non-exhaustive match — SAFE.** It returns `true` only when `tags_seen.len() == total_variants` for the single enum involved (`total_variants` counted by `enum_name` match). A missing variant strictly lowers the distinct-tag count, so the heuristic declines and the empty-`cases` branch in `compile_match_arms` produces the intended clear error. Confirmed the checker does no exhaustiveness analysis (`check_match`, lib.rs:326), so this codegen-level gate is the real and only backstop — and it holds.

### Wasm semantics — nested `if`-inside-`match`-arm (deepest path)

Traced `classify(-5)`: outer match compiles to `If(Result I64)` [then: `i64.const 1`] `Else` [ binding `x`; inner `If(Result I64)` [then: `-1`] `Else` [`2`] `End` ] `End`. Every reachable path of both blocks leaves exactly one `I64`; the mixed-arm test's `return`-ending arm is stack-polymorphic after `return`. All branch/`Result` typings are satisfied. Validation and execution both pass (verified by the green wasmtime-backed tests).

### Assessment

**Ready to merge?** Yes

**Reasoning:** The high-risk `match_covers_every_enum_variant` heuristic is airtight under adversarial analysis (unique per-enum tags + distinct-tag counting make it impossible to mask a missing variant or mis-fire), the value-position threading is a minimal additive change that leaves all existing paths byte-identical, dead code removal is confirmed safe, and the full workspace + tree-sitter suites are green with execution-based (not just validation-based) test coverage. Both deviations are justified and minimal; the only open items are explicitly out-of-scope follow-ups.</result>
<usage><subagent_tokens>91210</subagent_tokens><tool_uses>15</tool_uses><duration_ms>229512</duration_ms></usage>
</task-notification>

**Claude:** Ready to merge — no Critical or Important findings. Updating the ledger.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Both bugs are fixed and reviewed clean. Wrapping up this branch.
  - `Bash: cargo test --workspace 2>&1 | tail -5 && cd tooling/tree-sitter-plum && npx --yes tree-sitter-cli test 2>&1 | tail -3 && git status --short && git log --oneline`
  - `Bash: cd /Users/pyrossh/Code/plum && git log --oneline --all -- docs/superpowers/plans/2026-07-19-general-enum-support.md | head -3`

**Claude:** I missed committing the plan document itself earlier — fixing that now.
  - `Bash: git add docs/superpowers/plans/2026-07-20-tail-position-and-grammar-gaps.md && git commit -m "$(cat <<'EOF' docs: add implementation plan for tail-position and`

**Claude:** Now let's finish the branch properly.
  - `Skill: superpowers:finishing-a-development-branch`

**You:** Skill /superpowers:finishing-a-development-branch is already loaded above; instructions unchanged.
  - `Bash: cargo test --workspace 2>&1 | tail -3; GIT_DIR=$(cd "$(git rev-parse --git-dir)" 2>/dev/null && pwd -P); GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" 2>/`

**Claude:** Tests pass, and we're already on `master` directly (no separate branch, per your earlier choice) — everything is already committed there.
  - `AskUserQuestion`

**Claude:** Both gaps are fixed and left local per your choice. Now moving to the big one — generics monomorphization, as planned.
  - `mcp__ccd_session__mark_chapter: Generics monomorphization design`
  - `Bash: grep -rn "generic\|Generic" plum-checker/src/*.rs plum-wasm-codegen/src/lib.rs plum-core/src/ast.rs | grep -v "^Binary"`
  - `Bash: cat libs/std/list.plum; echo ---MAP---; cat libs/std/map.plum`
  - `Bash: cat examples/types.plum; echo ---FUNCTIONS---; cat examples/functions.plum`

**Claude:** I've explored the current state. Before designing, one important finding to share: `libs/std/list.plum` and `map.plum` use a lot of syntax that doesn't exist in the current grammar at all — closures (`|v| ...`), `Nil` (removed earlier this session), `?.` optional chaining, `:=`, decorators (`@HasTrait`), colon-arrow return types. Generics monomorphization alone won't make those files compile; that's a much bigger, separate gap. The realistic, achievable target is the currently-documented generic syntax (`type Box(a) = value: a`, `trait Comparable(a: Ord)`, `wrap(value: a) -> Bool`, multi-param `Map(a, b)`-style types) — making **user-declared generics in the existing language subset** actually monomorphize and compile.

Given the existing codegen already assumes every value has a statically-known concrete wasm type (I32/I64/F64) at compile time, the natural fit is **full monomorphization by specialization** (like Rust/C++ templates): collect every concrete instantiation site in the program (`Box(Int)`, `wrap(5)` inferring `a=Int`, etc.), generate one specialized copy of the type/function per unique concrete type-argument combination with a mangled name, then run each specialized copy through the existing non-generic codegen path unchanged. The alternative — type erasure / boxed uniform representation — would be a much bigger architectural rewrite (every generic-typed value becomes a runtime-tagged pointer) that doesn't fit how the rest of the compiler works today.
  - `AskUserQuestion`

**Claude:** That's exactly the right steer for this codebase specifically. Since every reference-shaped value here (class instances, `Str`, enum tags/payload pointers, `Bool`) is **already** represented as a uniform `i32` regardless of concrete type, a C#-style hybrid becomes much simpler than either of my original two options:

**Refined design: erasure by default, boundary-boxing only for primitives**

- Every generic type parameter (`a`, `b`, ...) compiles as `ValType::I32`, uniformly, with **zero specialization** — a generic class/function's body is compiled exactly **once**, not once per instantiation. Its fields/locals of generic type just get the same 8-byte-stride slot every field already gets, stored/loaded via `I32Store`/`I32Load`.
- When a concrete **reference-shaped** value (class instance, `Str`, `Bool`, enum value) flows into or out of a generic-typed position, nothing special happens — it's already an `i32`, passed straight through.
- When a concrete **`Int`/`Float`** value flows *into* a generic-typed position (e.g. `Box(value: 5)`), it gets **boxed**: bump-allocate an 8-byte cell, store the `i64`/`f64` there, pass the pointer through the generic slot.
- When a generic-typed value flows *out* into a position that statically needs a concrete `Int`/`Float` (arithmetic, a non-generic `Int`-returning context), it gets **unboxed**: load the `i64`/`f64` back through the pointer.
- This needs the checker to resolve, at each usage site, what concrete type a generic parameter is bound to (via existing unification, extended to substitute type arguments at call/construction sites) — just enough bookkeeping to tell codegen "box here" or "pass through" — not a full specialization/name-mangling pass, and no combinatorial blowup from nested generic calls.

This avoids monomorphization's code-bloat and specialization-pass complexity almost entirely, at the cost of a small, local box/unbox adapter only where primitives cross a generic boundary.
  - `AskUserQuestion`

**Claude:** Locking in full monomorphization by specialization. Here's the concrete design, in sections.

**1. Instantiation-site collection**

Walk the whole program to find every concrete generic instantiation actually needed: explicit type annotations with concrete args (`Box[Int]`, `Map(Str, Int)`), constructor calls where field types resolve generics via inference (`Box(value: 5)``a=Int`), and generic function/method calls where argument types resolve the params (`wrap(5)``a=Int`, `pair(1, "x")``a=Int, b=Str`). This reuses `plum-checker`'s existing unification — extended to bind generic parameter names to concrete types per call site instead of ignoring them.

**2. Specialization mechanism**

For each unique `(generic item, concrete type args)` pair, produce a substituted copy of the `Class`/`Fn`/`Enum` AST — every generic parameter name in a declared position (field types, param types, return types) replaced by its concrete type — under an internal mangled name (e.g. `Box$Int`, `wrap$Int$Str`), then run that copy through the **existing, unchanged** non-generic checker/codegen pipeline. User-facing syntax doesn't change at all; specialization is invisible plumbing.

**3. Transitive/recursive generics**

A specialized body can itself call another still-generic function — that's a new instantiation site, now with concrete types known. Handled as a worklist: start from concrete call sites in non-generic code, specialize, scan the freshly-specialized bodies for further instantiation sites, repeat to a fixed point. A recursion-depth guard catches pathological unbounded nesting (e.g. a generic type embedding itself in a growing chain) with a clear compile error rather than hanging.

**4. Out of scope**

Trait-bound enforcement stays permissive (unchanged from today — monomorphization needs substitution, not bound-checking). Generic code with zero call sites anywhere just never gets specialized/compiled (natural tree-shaking falls out of the worklist). And — as established earlier — this does **not** make `libs/std/list.plum`/`map.plum` compile; those need closures, `Nil`/optional-chaining, decorators, and other unimplemented syntax first.
  - `AskUserQuestion`

**Claude:** **5. Testing plan**

- **Checker tests**: a generic class instantiated at two different concrete types in the same program (`Box(Int)` and `Box(Str)`) resolves each usage's fields to the right concrete type; a generic function called with different concrete argument types infers correctly per call site; a multi-param generic (`Pair(a, b)`-shaped) resolves both params independently.
- **Codegen tests** (compiled and run via `wasmtime`, matching existing style): a generic class specialized at `Int` and at `Str` in the same program — confirms the two specializations don't alias or collide (distinct mangled identities, distinct compiled bodies); a generic function called at multiple concrete types; a generic method on a generic class; a transitively-generic call (a specialized generic function calling another generic function) to prove the worklist reaches fixed point correctly; a deliberately pathological unbounded-recursive-generic case asserting a clear compile error rather than a hang.
- **Examples**: a new example exercising `Box(a)`/`wrap(value: a)`/`pair(first: a, second: b)` end-to-end (these already exist in `examples/types.plum`/`functions.plum` as declarations only — extend with actual instantiations and a `main` that runs them), and update README's Known Gaps to drop the now-solved "user-defined generics aren't monomorphized" bullet (while keeping the note that this alone doesn't unblock `libs/std`'s actual `List`/`Map`, which need other unimplemented syntax too).
  - `AskUserQuestion`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generics-monomorphization-design.md`

**Claude:** Self-review: no placeholders, sections are internally consistent (mangling scheme, worklist, and out-of-scope boundary all agree), and the scope boundary about `libs/std` is stated clearly up front and echoed in Out of Scope so it can't be mistaken for a promise this solves the std lib. Committing.
  - `Bash: git add docs/superpowers/specs/2026-07-20-generics-monomorphization-design.md && git commit -m "$(cat <<'EOF' docs: add design spec for generics monomorphizatio`

**Claude:** Spec written and committed to [2026-07-20-generics-monomorphization-design.md](docs/superpowers/specs/2026-07-20-generics-monomorphization-design.md). Please review it before I write the implementation plan.

**You:** Yes looks good
  - `Skill: superpowers:writing-plans`

**You:** Skill /superpowers:writing-plans is already loaded above; instructions unchanged.
  - `Read: /Users/pyrossh/Code/plum/plum-checker/src/types.rs`

**Claude:** Before finalizing the plan, one gap I found while digging into the exact mechanics: the design spec covers generic classes/functions/traits, but doesn't address generic **enums** (`enum Option = | Some(a) | None`) — which is exactly the shape `libs/std/option.plum`'s real `Option` uses, and matters for the stated goal of eventually unblocking std. Enums don't have an explicit generics-declaration in the grammar today, so they'd need the same "detect a single-lowercase-letter field type as a generic placeholder" mechanism already needed for functions (which also lack an explicit generics list).
  - `AskUserQuestion`

**Claude:** Updating the spec to reflect this before writing the plan.
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generics-monomorphization-design.md`
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generics-monomorphization-design.md`
  - `Bash: git add docs/superpowers/specs/2026-07-20-generics-monomorphization-design.md && git commit -m "docs: extend generics monomorphization spec to cover generic enu`
  - `Bash: grep -n "fn parse_type\|fn parse_class\|fn parse_trait\|fn parse_enum\|fn parse_source\|fn parse_field\|GenericParam" plum-core/src/parser.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-core/src/parser.rs`
  - `Bash: grep -rln "#\[cfg(test)\]" plum-checker/src plum-core/src plum-wasm-codegen/src 2>/dev/null; grep -n "pub mod\|mod " plum-checker/src/lib.rs | head`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-generics-monomorphization.md`

**Claude:** Self-review: every spec requirement (instantiation collection, specialization, worklist/transitive, testing plan, generic enums) maps to a task; type signatures are consistent across Task 1's primitives and Task 2's driver; no placeholders. Given this is genuinely new algorithmic code (not a modification of existing tested logic like the smaller fixes), I've flagged in Global Constraints that correctness here rests primarily on the test suite, not on line-exact transcription — this is the appropriate posture for greenfield compiler-pass work.

Plan saved to [docs/superpowers/plans/2026-07-20-generics-monomorphization.md](docs/superpowers/plans/2026-07-20-generics-monomorphization.md).
  - `Bash: git add docs/superpowers/plans/2026-07-20-generics-monomorphization.md && git commit -m "$(cat <<'EOF' docs: add implementation plan for generics monomorphizati`

**Claude:** Two execution options:

1. **Subagent-Driven (recommended)** — fresh subagent per task, review between tasks, fast iteration
2. **Inline Execution** — execute tasks in this session with checkpoints

Which approach?

**You:** 1
  - `Bash: cargo test --workspace 2>&1 | tail -5`

**Claude:** Baseline clean. Setting up a fresh ledger for this plan.
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Write: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Pre-flight scan complete (self-authored, already cross-checked in writing-plans self-review). Dispatching Task 1.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-g`
  - `Agent: Implement Task 1: generics specialization primitives`

**Claude:** Task 1 implementer running in the background.

**You:** <task-notification>
<task-id>a3a9df0eb061889e5</task-id>
<tool-use-id>toolu_014s81ZPmsnuM3t2bMnXpmg7</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a3a9df0eb061889e5.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 1: generics specialization primitives" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>## Summary

**Status:** DONE

**Commit:** `1b220d2 feat(plum-checker): core generics specialization primitives`

**Test Summary:** All 9 monomorphize tests passing; full plum-checker suite green (40 tests total).

**Discrepancies:** None — brief code compiled exactly as written.

**Report:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md`

---

Task 1 complete. All 9 pure specialization primitives (generic detection, mangling, class/fn/enum specialization) implemented, tested, and committed. Zero integration wiring yet—ready for Task 2.</result>
<usage><subagent_tokens>46829</subagent_tokens><tool_uses>14</tool_uses><duration_ms>100642</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 1 reviewer.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 34f159a`
  - `Agent: Review Task 1 (spec + quality)`

**Claude:** Task 1 reviewer running in the background.

**You:** <task-notification>
<task-id>aa93a0b92186e335d</task-id>
<tool-use-id>toolu_01Vp4dZk8PLEiH7bZMTn52ou</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/aa93a0b92186e335d.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 1 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>### Spec Compliance
- ✅ Spec compliant

Diff is a verbatim transcription of the brief's prescribed code (167 lines in `monomorphize.rs`, matching the diff stat exactly) plus the exact prescribed test file and the one-line `lib.rs` module registration. No discrepancies claimed by the report, and none found on inspection.

Verified against real definitions (`plum-core/src/ast.rs`, `plum-checker/src/types.rs`):
- `ast::Class { name, implements, generics: Vec&lt;GenericParam&gt;, fields }`, `GenericParam { name, bounds }` — matches `class_generic_params`'s use of `c.generics.iter().map(|g| g.name.clone())`.
- `ast::Fn { name, type_param: Option&lt;String&gt;, params, returns, body }`, `ParamType::{Type,Variadic}`, `Type { name, generics: Vec&lt;Type&gt; }` — matches `fn_generic_params`/`specialize_fn`.
- `ast::Enum { name, variants }`, `EnumVariant { name, fields: Vec&lt;String&gt; }` (field types are bare strings, not `Type`) — matches `enum_generic_params`/`specialize_enum`'s direct string comparison/substitution.
- `ast::ReturnType { name, generics: Vec&lt;GenericParam&gt; }` — note this is `Vec&lt;GenericParam&gt;`, not `Vec&lt;Type&gt;` (asymmetric with `Type.generics: Vec&lt;Type&gt;`). `specialize_fn` correctly works around this by building a throwaway `ast::Type{name: r.name.clone(), generics: vec![]}`, substituting only the top-level name, and discarding the substituted `generics` (reconstructing `ReturnType{generics: vec![]}`). This means nested return-type generics (e.g. hypothetical `List&lt;a&gt;` as a return type) are not substituted — but this is a pre-existing AST asymmetry, not a bug the implementer introduced, and it's not exercised by this task's test cases.
- `plum_checker::types::PlumType` variants (`TInt/TFloat/TBool/TStr/TUnit/TVar/TFun/TNamed`) and its `Display` impl — matches `mangle`'s use of `t.to_string()` and `plum_type_to_ast_type`'s exhaustive match.

Detection-order/uniqueness logic (`fn_generic_params`, `enum_generic_params`) uses a linear `Vec` + `!names.iter().any(...)` scan to preserve first-appearance order while deduping — correct for the required semantics. `class_generic_params` trusts the declared list directly (no dedup), which matches the brief (`Class`/`Trait` generics come from the explicit list, not implicit detection).

No error-producing code paths exist in this task (pure data transforms, no `Result`), so the `"monomorphize: ..."` error-prefix constraint doesn't apply here — correctly deferred to a later task.

No wiring: `lib.rs` diff is exactly `pub mod monomorphize;` added; nothing in the diff touches `check_source`/`compile_source` or any call site.

### Strengths
- Doc comments accurately explain the *why* (e.g., why `Fn`/`Enum` need implicit detection vs `Class`'s explicit list; why body rewriting is deferred to Task 2).
- Test suite (`plum-checker/tests/monomorphize_tests.rs`) exercises all 9 functions/struct with both happy-path and edge cases (empty generics, method-receiver override, multi-arg mangle).

### Issues
#(none found)

#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- `specialize_fn`'s return-type handling silently drops any (currently impossible, given `ReturnType.generics: Vec&lt;GenericParam&gt;`) nested generic substitution — worth a one-line comment noting this is a known AST-shape limitation, so a future task doesn't rediscover it as a "bug." Not a defect in this task's scope.

### Assessment
**Task quality:** Approved
**Reasoning:** The implementation is a literal, unmodified copy of the brief's prescribed code; all types/shapes were independently verified against `plum-core/src/ast.rs` and `plum-checker/src/types.rs` and match exactly; no wiring was added; all `pub` functions are genuinely public API with no dead-code workarounds needed.</result>
<usage><subagent_tokens>52914</subagent_tokens><tool_uses>5</tool_uses><duration_ms>43201</duration_ms></usage>
</task-notification>

**Claude:** Task 1 approved. Updating the ledger and moving to Task 2 — the core worklist driver, the hardest part of this feature.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-g`

**Claude:** This is the architecturally hardest task — dispatching on the most capable model.
  - `Agent: Implement Task 2: monomorphization worklist driver`

**Claude:** Task 2 (worklist driver) implementer running in the background — this is the biggest, most judgment-heavy piece of the whole feature, so it may take a while.

**You:** <task-notification>
<task-id>a7c05a66e453e9fc7</task-id>
<tool-use-id>toolu_01XybSXdqZV4m2jftVNacnAy</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a7c05a66e453e9fc7.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 2: monomorphization worklist driver" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

**Commit:** `22140cf` feat(plum-checker): monomorphize generic classes/functions/methods/enums before checking and codegen

**Tests:** All 4 new checker tests green with correct resolution; full `cargo test --workspace` green.

**Divergence from brief:** Moderate — the worklist/fixed-point driver, mangling, and traversal are the brief's as-is, but I had to fix real blockers the brief's design assumed away: (1) a plum-core parser bug where `c.generics` was never populated (grammar inlines `generic_type`/`generic`, so the driver's `!c.generics.is_empty()` check always saw no generics); (2) inference-based return-type resolution in the monomorphizer, since `-&gt; Box` (bare generic class) and `-&gt; a` (which the grammar can't even parse → `returns: None`) both left the specialized return type wrong; (3) registering specialized classes/methods into the inference tables so method-body field access resolves. Also recurse into string interpolation. All documented in the report.

**Concerns:** `-&gt; a` generic-param returns are unparseable by the current grammar and handled via body-tail inference rather than a grammar fix; a proper grammar change + `parser.c` regen is a worthwhile follow-up but was out of scope. No current case is affected.

**Report:** /Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md</result>
<usage><subagent_tokens>124271</subagent_tokens><tool_uses>56</tool_uses><duration_ms>995896</duration_ms></usage>
</task-notification>
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md`

**Claude:** This has significant, well-justified deviations touching a different crate (a real parser bug: `Class.generics` was never populated at all) plus a new return-type-inference mechanism. Given the risk, dispatching the reviewer on the most capable model with heavy scrutiny.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 1b220d2`
  - `Agent: Review Task 2 (spec + quality)`

**Claude:** Task 2 reviewer running in the background with heavy scrutiny on the return-type-inference mechanism, given it's the riskiest part of this diff.

**You:** <task-notification>
<task-id>a50ac39524b3a83d6</task-id>
<tool-use-id>toolu_01AvDdGe53V4iX3qQovuVyZ1</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a50ac39524b3a83d6.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 2 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>I have verified all four deviations against the actual grammar, parser, and checker code. Confirmed: `generic_type`/`generic` are inlined (grammar.js:59), `generic` is only `a|b|c|d` (line 505), `return_type` requires a capitalized `type_identifier` (lines 132-133) so `-&gt; a` cannot parse, and the parser rewrite correctly reconstructs `GenericParam`s. No existing test depends on the old broken `generics: []` behavior (the only `generics.is_empty()` assertion is in `monomorphize_tests.rs:72`, checking a *specialized* class, which is correct). The one real soundness issue is in deviation #2's specialization branch. Report follows.

### Spec Compliance
- ✅ Spec compliant on the four target behaviors: all 4 checker tests are wired end-to-end, `monomorphize_source` is the single entry point, and it is invoked as the first line of both `check_source` (lib.rs:30-33) and `compile_source` (wasm-codegen lib.rs:640). Error prefix `"monomorphize: ..."` is used consistently.
- ✅ Deviation #1 (parser fix) is verified correct and necessary — see Strengths.
- ✅ Deviation #3 (registering specializations) is correct: entries are keyed by mangled name (`Box$Int`, `(mangled_receiver, method)`, mangled fn name), never clobber the original generic templates (which live under their bare names in `self.classes`/`methods`/`global_env`), and cannot leak between specializations since all keys are mangled-unique within one fresh `Monomorphizer`.
- ✅ Deviation #4 (string-interp recursion) is a harmless strictly-more-correct superset (see Minor).
- ⚠️ Deviation #2 (return-type inference) is functionally correct for the intended cases but its **specialization branch is unsound** — it can mask real type errors (see Critical/Important).

### Strengths
- Deviation #1 is the right call. The inlining claim is real (grammar.js:59 `inline: ($) =&gt; [$.generic_type, $.generic]`; `generic: choice($.a,$.b,$.c,$.d)` line 505). Under the old code `children_of_kind(g, "generic_type")` returned nothing, so **every** parsed `Class`/`Trait` had `generics: []`, making Task 1's `class_generic_params` dead against real source. The rewrite (parser.rs:114-124) matching `"a"|"b"|"c"|"d"` exactly mirrors the grammar's only legal generic spellings; tracing `type Cat(a: Ord)``[{a, [Ord]}]` and `type Pair(a, b)``[{a,[]},{b,[]}]` both reconstruct correctly. It is a one-function, demonstrably-necessary root-cause fix with no in-crate alternative (re-deriving class generic-ness inside `plum-checker` would require re-walking tree-sitter nodes, i.e. duplicating the parser). Removing the now-dead `parse_generic_type` is appropriate.
- Deviation #2's *ordinary-function* branch is precise: `classes_generic.contains_key(&amp;rt.name)` only fires when the declared return names a *known generic class* (an inherently unresolved bare-generic return), and never on a concrete class like `Int`/non-generic `Foo`. A function `-&gt; Int` with a buggy body is left untouched and correctly rejected by the checker (verified against check_fn at lib.rs:186-218).
- `maybe_rewrite_return` guards against `TVar`/`TFun` tails, so it never fabricates a return from an un-inferrable body.
- RED evidence is consistent with pre-change code: the two function tests (using `fn_generic_params`, which reads `f.params`/`f.returns` directly, not `c.generics`) would pass without the parser fix, while the two class/method tests would fail — exactly as the report's "intermediate" state describes. The reported RED messages (`field 'value': expected a, found Int`, `call 'wrap' arg 0: expected a, found Int`) match what the checker produces on un-monomorphized source.

### Issues

#### Critical (Must Fix)
None that block the task's stated goal (all target tests and the existing suite behave correctly).

#### Important (Should Fix)
- **Specialization return-type overwrite masks real type errors**`monomorphize.rs`, `rewrite_fn_body` called with `resolve_return: true` (worklist arms at diff lines 428 and 447) → `maybe_rewrite_return` (diff lines 135-150). For every specialization, `needs` is unconditionally `true` (`None =&gt; resolve_return` and `Some(rt) =&gt; resolve_return || ...`), so the declared return type is **always** overwritten with the inferred tail type — even when it was a genuine, concrete, correct annotation. Consequence: a generic function with a concrete return and an inconsistent body, e.g.

  ```
  wrong(x: a) -&gt; Int =
    "hello"
  ```

  when instantiated (`wrong("s")``wrong$Str`) has `returns` rewritten `Int``Str`, after which `check_fn` (lib.rs:186-218) validates the body against the rewritten `Str` and passes. The `expected Int, found Str` error is silently swallowed. This is the exact "fixing the annotation instead of letting the checker reject it" failure the review flagged. It is narrow (only generic functions/methods whose body is internally inconsistent with a *concrete* declared return) and no current test exercises it, but it is a genuine soundness hole in a pass that sits ahead of the type checker. Fix: for specializations, restrict the overwrite to the cases that actually need it — `f.returns` is `None` (the unparseable `-&gt; a` case) or names a generic param / still-generic class — rather than forcing it unconditionally. That keeps the `getValue&lt;Box&gt;() -&gt; a` and `-&gt; Box` cases working while letting the checker reject concrete-return mismatches. (For multi-path bodies the risk is lower — early `Return(Some)` statements are still checked at lib.rs:243-251 — but against the tail-derived type, which can produce a misleading diagnostic.)

#### Minor (Nice to Have)
- Deviation #4 (string-interp recursion, diff 329-335): technically outside the plan's stated scope, but it only rewrites call sites that would otherwise be left generic, is a no-op when there are none, and string interpolation doesn't compile yet anyway — so practical risk is nil and it is strictly more correct. Acceptable; worth a one-line note in the plan's scope ledger rather than a change.
- `PendingSpecialization::Enum` is `#[allow(dead_code)]` (diff 67-68) and never enqueued — there is no `resolve_enum_instantiation`, so generic enums are dropped from output but never specialized at instantiation sites. This matches the brief's own driver (which also never pushed the Enum arm), so it is a pre-existing scope gap, not a regression — but generic-enum instantiation is silently unsupported and worth tracking as a follow-up.
- The signature-registration blocks in the two worklist arms (diff 432-439 and 450-457) are near-identical; a small helper would reduce duplication.

### Assessment
**Task quality:** Needs fixes
**Reasoning:** The four target behaviors work, the tests are genuine, and deviations #1/#3/#4 are correct and well-justified; but deviation #2's specialization branch unconditionally overwrites concrete return annotations, which can mask real type errors in generic function bodies — a soundness regression in a pre-checker pass that should be tightened to only touch genuinely-unresolved returns.</result>
<usage><subagent_tokens>66799</subagent_tokens><tool_uses>11</tool_uses><duration_ms>174916</duration_ms></usage>
</task-notification>

**Claude:** Good review — caught a real soundness gap (return overwrite masks type errors) and also surfaced that generic enum instantiation is never actually wired up (only dropped, never specialized) — which Task 3's tests will need. There's also a deeper architectural issue worth resolving now: the existing enum representation keys variants by bare name globally, so specializing the same generic enum at two different concrete types in one program would collide. Dispatching one fix for all of this, scoping enums pragmatically to single-instantiation-per-program for this plan.
  - `Agent: Fix Task 2 review findings`

**Claude:** Fix subagent running in the background — this covers both the soundness gap and wiring up generic enum instantiation (which Task 3 needs).

**You:** <task-notification>
<task-id>aef4c98c9947c02ba</task-id>
<tool-use-id>toolu_01AD6sR8EAWmWNJeXPiP6sB8</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/aef4c98c9947c02ba.output</output-file>
<status>completed</status>
<summary>Agent "Fix Task 2 review findings" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Both findings fixed, committed, all green.

**Status:** DONE

**Commit created:**
- `c273ea5` fix(plum-checker): stop masking type errors in specializations; resolve generic enum instantiation

**Test summary:** `cargo test --workspace` fully green (checker 37 tests incl. 3 new); tree-sitter corpus 22/22 parses green (incl. new generic-variant case).

**Both findings fixed:**
- Finding 1: `maybe_rewrite_return`'s `Some(rt)` arm no longer overwrites a concrete, correct declared return type for specializations — the internally-inconsistent generic function (`wrong(x: a) -&gt; Int = "hello"`) is now correctly rejected. RED→GREEN proven.
- Finding 2: `resolve_enum_instantiation` implemented; single instantiation resolves end-to-end (verified by inspecting monomorphized `Option$Int` output), multi-instantiation of one generic enum produces a clear `monomorphize:` error. RED→GREEN proven.

**One flagged deviation:** the task said I "shouldn't need to touch the grammar," but Finding 2 was impossible without it — `enum Option = | Some(a)` produced an ERROR node (variant params rejected lowercase generics), so generic enums were never even recognized as generic. I extended `enum_field` to accept generics, regenerated `parser.c`, updated `parse_enum_variant`, and added a corpus test. This mirrors round 1's class-generics parser fix.

**Report:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md` (new "## Fix round 2" section; file is gitignored so left uncommitted on disk as required).

Known scope limitation (documented in the collision error message and the `enum_variant_owner` field comment): one generic enum may be instantiated at only a single concrete type per program; a second is detected and rejected, never silently mishandled.</result>
<usage><subagent_tokens>108062</subagent_tokens><tool_uses>66</tool_uses><duration_ms>768905</duration_ms></usage>
</task-notification>
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 1b220d2`
  - `Agent: Re-review Task 2 after fix round 2`

**Claude:** Re-review of Task 2's cumulative diff running in the background.

**You:** <task-notification>
<task-id>a527ef038e07d98b5</task-id>
<tool-use-id>toolu_01FpAM8QNaCz1W3hih3QGdzP</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a527ef038e07d98b5.output</output-file>
<status>completed</status>
<summary>Agent "Re-review Task 2 after fix round 2" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Concrete `Some(Int)` parses clean (no ERROR node) — grammar change is purely additive. My independent verification is complete.

## Review Report

### Spec Compliance
- ✅ Spec compliant. All three high-risk areas (return-overwrite fix, enum collision detection, grammar change) verified sound against the actual diff and source, not the report.

### Strengths
- The return-overwrite fix is genuinely correct, not a patch over the symptom. The `Some(rt)` arm now gates purely on `is_generic_param_name(&amp;rt.name) || classes_generic.contains_key(&amp;rt.name)` with no `resolve_return ||`, so a concrete declared return is never touched.
- The design cleanly separates concerns: `specialize_fn` (monomorphize.rs:146-149) already substitutes generic-param-letter returns to concrete via `subst` *before* `maybe_rewrite_return` runs. This means the `is_generic_param_name` branch only ever fires on returns that `subst` couldn't resolve — making the "concrete-but-wrong body" adversarial case fall through to the checker correctly.
- Enum collision detection is placed in the worklist arm (guarded by `specialized.insert`), so it runs exactly once per mangled enum and cannot false-positive on repeated construction sites.
- Grammar change is strictly additive (`choice($.type_identifier, $.generic)`); verified concrete `Some(Int)` still parses with no ERROR/MISSING node.

### Detailed verification of the priority checks

**1. Return-overwrite soundness — CONFIRMED SOUND.**
Traced the adversarial case `wrong(x: a) -&gt; Int = "hello"` called via `useIt() -&gt; Str = wrong("s")`:
- `wrong` enqueues as `wrong$Str`; `specialize_fn` keeps `returns: Some("Int")` (`"Int"` ∉ subst).
- `rewrite_fn_body(resolve_return=true)`: tail `"hello"` infers `TStr`. `maybe_rewrite_return`: `needs = is_generic_param_name("Int")(false) || classes_generic.contains_key("Int")(false) = false`. Return type **left as `Int`**.
- Checker then sees `wrong$Str() -&gt; Int = "hello"` → surfaces `expected Int, found Str`. Correctly rejected, not masked. Also traced `id(x:a) -&gt; a = "hello"` (param-letter return): `specialize_fn` substitutes return to concrete `Int`, `needs=false`, still rejected. The `None`-arm fill-in stays gated on `resolve_return`, so ordinary void functions aren't given fabricated returns.

**2. Enum collision detection — CONFIRMED AIRTIGHT.**
- (a) No false positive: second `Some(5)``mangled == "Option$Int"` already in `enqueued`, not re-pushed; worklist processes `Option$Int` once, claims `Some`/`None` once.
- (b) Genuine collision: `Some(5)``Option$Int` and `Some("x")``Option$Str` both enqueue; second worklist arm finds `enum_variant_owner["Some"] == "Option$Int" != "Option$Str"` → returns `Err` before any output is checked (the error aborts `monomorphize_source`, so no corrupt `EnumVariants` registration occurs). Error string contains "monomorphize"/"Option"/"more than one concrete type" as the test asserts.
- (c) Variant names left unmangled: `specialize_enum` uses `v.name.clone()`; `resolve_enum_instantiation` explicitly does not rewrite `call.name`. Post-monomorphization only `Option$Int` remains with variants `Some`/`None`, so the flat bare-name `EnumVariants` lookup resolves match patterns correctly with no match-lowering changes.

**3. Grammar change — REAL, NECESSARY, NON-REGRESSING.** `enum_field` previously accepted only `$.type_identifier` for payloads; a lowercase `generic` letter would not match. New corpus test parses `Some(a)` to `(field (type_identifier) (a))` with no ERROR; concrete `Some(Int)` independently re-verified as still clean.

**4. Tests — REAL and PROVE their claims.** The inconsistent-body test asserts `is_err` (would pass Ok under the old buggy overwrite). The single-instantiation test goes beyond a permissive `check_source` Ok — it inspects `monomorphize_source` output for a concrete `Option$Int` with `Some` field `["Int"]` and asserts the generic `Option` template is dropped. Multi-instantiation test asserts the specific collision error.

**5. Whole-task coherence — GOOD.** The parser fix (deviation #1) and the new enum-field parser change are exact analogues (both handle grammar-inlined nodes: `a`/`b`/`c`/`d`), and interact cleanly — `enum_generic_params` now sees populated variant fields just as `class_generic_params` sees populated `c.generics`. Both pipeline entry points route through `monomorphize_source`.

### Issues

#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- **Bare generic-class return (`-&gt; Box`) can still mask a wrong body.** When a return names a generic class used bare, `maybe_rewrite_return` overwrites it with the body's tail type unconditionally (`classes_generic.contains_key(&amp;rt.name)` branch). This is the documented "body-determined" design and `-&gt; Box` bare is itself unresolvable at the annotation, so it's defensible — but it is a narrower instance of the exact masking class the round-2 fix targeted. Worth a follow-up note that a bare-generic-class return does not get the same "reject inconsistent body" protection that a concrete return now does.
- **Double monomorphization in the CLI path.** `plum-cli/src/main.rs` calls `check_source` (line 110) then `compile_source` (line 118); each runs `monomorphize_source`. Harmless (the second pass sees no generics) but redundant work. Not a correctness issue.
- **No corpus test locks in the concrete-payload enum case** (`Some(Int)`) — it's only exercised indirectly via examples. Adding it beside the new generic-variant corpus case would guard against future grammar regressions. (I verified it parses correctly today.)

### Assessment
**Task quality:** Approved
**Reasoning:** Both round-2 findings are fixed at the root: the return-overwrite condition is genuinely sound (traced adversarially — inconsistent bodies are now rejected, not masked), and the enum collision detection is airtight with correct no-op re-use, correct error on genuine collision, and correctly unmangled variant names. The grammar change is necessary, minimal, additive, and non-regressing. Remaining items are minor/documented tradeoffs, not defects.</result>
<usage><subagent_tokens>65510</subagent_tokens><tool_uses>8</tool_uses><duration_ms>168374</duration_ms></usage>
</task-notification>

**Claude:** Task 2 approved — both critical findings genuinely resolved. Updating the ledger and moving to Task 3.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: git rev-parse HEAD`
  - `Agent: Implement Task 3: generics codegen tests and docs`

**Claude:** Task 3 (final task) implementer running in the background.

**You:** <task-notification>
<task-id>abd0d65f1164b733b</task-id>
<tool-use-id>toolu_01N2bivUDjFQM9N6NdNDcJzh</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/abd0d65f1164b733b.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 3: generics codegen tests and docs" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>I'll pause here and wait for the monitor to notify me that the test process has completed before proceeding further.</result>
<usage><subagent_tokens>77871</subagent_tokens><tool_uses>37</tool_uses><duration_ms>283597</duration_ms></usage>
</task-notification>

**Claude:** That doesn't look like a proper completion report — let me check the actual state of things directly.
  - `Bash: git log --oneline -3; git status --short; cat /Users/pyrossh/Code/plum/.superpowers/sdd/task-3-report.md 2>/dev/null | tail -60`
  - `Bash: git diff --stat; echo ---; cargo test --workspace 2>&1 | tail -15`
  - `Bash: cargo test --workspace 2>&1 | grep -E "test result|FAILED|error\["`

**Claude:** The implementer got interrupted before finishing (examples/README updates and commit are still pending), but its test additions are solid — full workspace green. Resuming it to finish rather than starting over.
  - `ToolSearch: SendMessage`
  - `SendMessage: abd0d65f1164b733b`

**Claude:** Resumed and running in the background.

**You:** <task-notification>
<task-id>abd0d65f1164b733b</task-id>
<tool-use-id>toolu_01Cb7v21B3GSVnyVED5a2mwP</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/abd0d65f1164b733b.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 3: generics codegen tests and docs" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>All done. Report written and committed.

**Status:** DONE

Commit: `2e28ecc` — "test+docs: generics monomorphization complete; extend examples, update known gaps" (branch `fix/monomorphize-review-round2`)

Test summary: `cargo test --workspace` fully green (17 test-result blocks, 0 failures); `npx tree-sitter-cli test` 22/22 corpus tests pass (100%).

Concerns:
- The plan's suggested recursion-guard repro (`Wrap(a)`/`inner: Wrap(Box(a))`) doesn't actually reach the worklist (a class's own field-type declaration is never a call site), matching the plan's own caveat — I substituted a genuine repro (`recurse` re-wrapping its arg in `Box` each recursive call) that does grow the worklist unboundedly and correctly trips the 10,000 guard.
- That test takes ~100-112s real time (O(n²) from growing mangled-name string bookkeeping up to the cap) — bounded and deterministic, but slow; out of scope to optimize here.
- `.superpowers/sdd/task-3-brief.md` contains stale content from an unrelated, already-merged plan — flagged in the report for repo hygiene, didn't block the work since the actual plan doc was authoritative.

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-3-report.md`</result>
<usage><subagent_tokens>102292</subagent_tokens><tool_uses>22</tool_uses><duration_ms>524520</duration_ms></usage>
</task-notification>

**Claude:** Task 3 is done — but it committed on a separate branch (`fix/monomorphize-review-round2`) rather than `master`. Let me check and reconcile that before reviewing.
  - `Bash: git branch -a; echo ---; git log --oneline -5; echo ---; git status --short`
  - `Bash: git log --oneline master -3; echo ---; git merge-base master fix/monomorphize-review-round2`

**Claude:** The Task 2 fix-round-2 subagent apparently created this branch on its own rather than committing to `master` directly (deviating from the established convention this session). Since history is linear, I can fast-forward `master` to include it cleanly.
  - `Bash: git checkout master && git merge --ff-only fix/monomorphize-review-round2 && git log --oneline -5 && git branch -d fix/monomorphize-review-round2`

**Claude:** Back on `master` with everything merged. Now dispatching the Task 3 reviewer.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package c273ea5 2e28ecc`
  - `Agent: Review Task 3 (spec + quality)`

**Claude:** Task 3 reviewer running in the background.

**You:** <task-notification>
<task-id>a6bc6040d8cfaf1dc</task-id>
<tool-use-id>toolu_01464hnV6ctd48HuUhKWfSH2</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a6bc6040d8cfaf1dc.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 3 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>All spot checks confirm the report's claims exactly. Now finalizing the review.

### Spec Compliance
- ✅ Spec compliant. The `.superpowers/sdd/task-3-brief.md` handed to both the implementer and this review is stale (it describes an unrelated, already-merged `examples/match.plum` bare-tail-style task, commit `87b54a9`). The implementer correctly identified this and worked from the real Task 3 section in `docs/superpowers/plans/2026-07-20-generics-monomorphization.md:874-1074`, which I verified directly. Against that authoritative source:
  - All 5 codegen tests (`plum-wasm-codegen/tests/codegen_tests.rs`) match the plan verbatim, character-for-character, and all pass (confirmed by running them).
  - The recursion-guard test repro substitution is a legitimate, plan-anticipated adjustment, not scope drift — the plan (line 1014) explicitly predicts the `Wrap(a)`/`inner: Wrap(Box(a))` repro may not reach the worklist and instructs the implementer to "investigate and adjust the repro source until it genuinely exercises an unbounded worklist growth." The substituted `recurse`/`Box` repro genuinely does: traced through `plum-checker/src/monomorphize.rs`, `recurse(5)``recurse$Int` → body infers `Box$Int``recurse(b)` demands `recurse$Box$Int` → ... a strictly-growing, never-repeating chain that never hits the `specialized`/`enqueued` dedup. Ran it directly: passes in ~101s, error message contains `"monomorphize"`, confirming the guard trips rather than hanging.
  - Examples extensions (`examples/functions.plum`'s `useWrap`/`usePair`, `examples/types.plum`'s `makeIntBox`/`makeStrBox`) match the plan verbatim; all 8 example tests across both crates pass (confirmed).
  - README changes match the plan's prescribed Known-Gaps edit and Generics-section rewrite, plus a good-faith addition documenting the single-instantiation-per-enum limitation (carried over from Task 2, correctly not "fixed" here).
  - `cargo test --workspace` and tree-sitter corpus: not re-run in full per instructions; re-ran the tree-sitter suite (22/22, 100%) and all touched Rust test files individually — all green, consistent with the report.

### Strengths
- The class-vs-enum aliasing distinction is verified correct at the source level: classes are mangled by full type-argument name (`Box$Int`, `Box$Str` are distinct structs, no collision possible), while enums additionally go through a bare-variant-name ownership table (`enum_variant_owner` in `plum-checker/src/monomorphize.rs:188-193`) — this is the actual, source-confirmed reason the enum limitation exists and the class case doesn't need it. `examples/types.plum`'s two `Box` instantiations at `Int` and `Str` correctly exercise the unrestricted class case.
- The recursion-guard repro substitution shows real engineering: the implementer read the actual pass logic rather than guessing, and the substitution genuinely exercises the guard's intended failure mode (unbounded, non-repeating specialization growth) rather than an unrelated early error.
- README's new Generics prose and Known Gaps entry are accurate and specific — they name the exact unrelated blockers (closures, `Nil`/optional chaining, decorators, colon-arrow syntax) still preventing `libs/std` from compiling, rather than vaguely blaming generics.
- Process transparency: the report explicitly flags the stale-brief mismatch, the runtime cost of the new test, and the process deviation (branch name), rather than glossing over them.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
- The new recursion-guard test adds ~100-112s to every `cargo test --workspace` run, permanently, for a test whose only job is confirming a hardcoded 10,000-iteration cap trips. This is a real, ongoing CI/dev-loop tax. It's directly inherited from the plan's own design (plan line 1013-1014 expects exactly this test and this guard value, and fixing it would mean touching `monomorphize.rs`'s guard/bookkeeping, which is out of scope for this task) — so it's not an implementer defect, but it's worth flagging to whoever owns the plan/repo going forward: either lower the guard cap, make the bookkeeping sub-quadratic, or mark this specific test `#[ignore]`/slow-tagged so default `cargo test` runs stay fast. Not a blocker for this task's approval since the plan mandated this exact behavior.

#### Minor (Nice to Have)
- Process deviation noted per the task instructions: the implementer committed to a self-created branch (`fix/monomorphize-review-round2`) instead of `master`, later fast-forwarded by the controller. Commit content is correct; this is purely a process note, not a code issue.
- README documents the generic-enum single-instantiation limitation but still doesn't mention the two other plan-wide out-of-scope generics limitations (a method adding its own extra generic parameter beyond its receiver's; trait-bound enforcement for `Comparable(a: Ord)`-style bounds). Neither was ever documented before this plan either, and documenting them wasn't requested by Task 3's brief, so this is a nice-to-have for future polish, not a gap introduced by this diff.

### Assessment
**Task quality:** Approved
**Reasoning:** All prescribed tests, examples, and docs from the plan's real Task 3 section are present, verbatim-matched, and independently verified to pass; the one deliberate deviation (recursion-guard repro) is exactly the adjustment the plan itself anticipated and is proven, by tracing the actual monomorphizer source, to genuinely exercise the guard. The cumulative 3-task feature reads as coherent: classes/functions/methods/enums are monomorphized end-to-end through real wasm execution, the one genuine limitation (generic-enum single-instantiation) is correctly preserved and now documented in README, and out-of-scope items (extra method generics, trait-bound enforcement, `libs/std` compiling) are correctly left alone and accurately described as still-open gaps.</result>
<usage><subagent_tokens>56640</subagent_tokens><tool_uses>16</tool_uses><duration_ms>248011</duration_ms></usage>
</task-notification>

**Claude:** One real, ongoing cost flagged: the recursion-guard test permanently adds ~100s to every test run. Worth a quick fix before finishing — dispatching a small change to mark it `#[ignore]` (with a comment explaining how to run it) so the default suite stays fast.
  - `Agent: Mark slow recursion-guard test as ignored`

**Claude:** Fix subagent running in the background.

**You:** <task-notification>
<task-id>a61ac371a07e0a4ac</task-id>
<tool-use-id>toolu_01L4hgGa2PY4V5PiiRYCayc6</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a61ac371a07e0a4ac.output</output-file>
<status>completed</status>
<summary>Agent "Mark slow recursion-guard test as ignored" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

**Commit:** `6613e6d test(plum-checker): mark slow recursion-guard test as ignored by default`

**Test verification:**
- Default run: Test skipped (shown as "1 ignored" in output)
- With `--ignored`: Test passes in 104.98s as expected
- Full workspace: All suites green (~1.2s total, 106 tests passed, 1 ignored)</result>
<usage><subagent_tokens>30808</subagent_tokens><tool_uses>9</tool_uses><duration_ms>165821</duration_ms></usage>
</task-notification>

**Claude:** Task 3 approved and the slow-test finding addressed. Updating the ledger and moving to the final whole-branch review.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: git log --oneline 34f159a~1..6613e6d && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review`
  - `Agent: Final whole-branch review of generics monomorphization`

**Claude:** Final whole-branch review running in the background on Opus, independently re-verifying the two highest-risk fixes from scratch rather than trusting prior approvals.

**You:** <task-notification>
<task-id>a3121982d2d930f6c</task-id>
<tool-use-id>toolu_01GR37yPYZ8LTsLU2Nx36uRG</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a3121982d2d930f6c.output</output-file>
<status>completed</status>
<summary>Agent "Final whole-branch review of generics monomorphization" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>I have completed my independent verification. All checks pass; both high-risk fixes are sound under my own re-derivation.

---

## Review: Generics Monomorphization (34f159a..6613e6d)

### Strengths

- **Architecturally clean.** `monomorphize_source` runs as the first line of both `check_source` and `compile_source` and emits a fully-concrete `ast::Source`. I verified both consuming pipelines are otherwise untouched (`plum-checker/src/lib.rs:142`, `plum-wasm-codegen/src/lib.rs:1310` are one-line additions each). The "no downstream changes" claim holds.
- **Both previously-approved high-risk fixes independently re-verified as sound** (details below).
- **Grammar fixes are minimal and correct**, disambiguated cleanly by identifier case.
- **Tests genuinely execute** via `wasmtime``run_main` (`codegen_tests.rs:195`) instantiates the module and calls `main`, asserting real return values (7, 42, 9, 42, 13), not just validating bytes.
- **Full suite green**: `cargo test --workspace` all pass (1 intentionally ignored), `tree-sitter-cli test` 22/22, corpus + formatter tests unaffected.
- **Git state clean**: HEAD at 6613e6d, only `master` local, 5 expected commits, no orphaned `fix/monomorphize-review-round2` branch, no residue. Only untracked items are unrelated `.recall/` dirs.

### Issues

#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)

1. **`plum-checker/src/monomorphize.rs:380` — O(n²) `global_env` cloning.** `rewrite_fn_body` does `let mut env = self.global_env.clone()` on every specialization, while `global_env` grows by one entry per specialized free function (`:811`). This is the real source of the ~100s runaway-guard-test runtime: reaching the 10,000 cap costs ~50M map-entry clones. It's benign for normal programs (few specializations) and only bites the pathological error path, but it means a user who *does* write unbounded generic recursion waits ~100s before the clear error fires. If cheap, thread a shared reference or a persistent map. Not a merge blocker.

2. **`plum-checker/tests/checker_tests.rs:1028` — the guard test is `#[ignore]`, so CI never exercises the recursion guard.** The `guard &gt; 10_000` logic can now regress silently. Consider a fast variant (e.g. a lower, test-injectable cap, or a repro that trips a smaller bound) so at least one guard test runs by default. The chosen repro itself is sound — I traced it: `recurse` re-wraps its arg in `Box` each call, so each specialization demands `recurse$Box$…$Int` at a strictly larger type, growing the worklist without bound until the cap. Confirmed `infer_expr` returns `TNamed(mangled_name)` for a rewritten `ClassCall` even before the class is registered (`lib.rs:510/513`), which is what makes the chain grow.

3. **`plum-checker/src/monomorphize.rs:434` — a void generic function with an expression body could get a fabricated return type.** For `f(x: a) = x` (no `-&gt; ` annotation, `returns: None`, body `FnBody::Expr`), the specialization path (`resolve_return: true`) infers a tail and fills in a return. This matches the documented intent (the grammar can't spell `-&gt; a`) and is harmless since the body genuinely produces that value, but it's a slightly surprising edge of the `None → fabricate` rule. Worth a one-line note only.

### Re-verification of the two high-risk claims (from scratch)

**Return-type-overwrite fix — SOUND.** Traced the adversarial case `wrong(x: a) -&gt; Int = "hello"` instantiated via `useIt() -&gt; Str = wrong("s")`:
- `specialize_fn` produces `wrong$Str` with `returns: Some("Int")`, body `"hello"`.
- `maybe_rewrite_return` (`:434`) computes `needs` from `Some(rt)` where `rt.name == "Int"`: `is_generic_param_name("Int") == false` and `classes_generic.contains_key("Int") == false``needs == false`. The concrete `-&gt; Int` is **not** overwritten.
- `check_fn` then rejects the Str-body-vs-Int-return mismatch (`lib.rs:34` / `:212`). The real error is surfaced, not masked. The dedicated test at `:909` asserts exactly this (`is_err`), and it passes.

**Enum single-instantiation + collision detection — SOUND on all three sub-claims:**
- (a) *No false positive on re-use*: two `Some(5)` sites both mangle to `Option$Int`; the second is deduped by `enqueued`/`specialized` (`:612`) and the arm's `if !m.specialized.insert(...) { continue; }` (`:815`), so `"Some"` is claimed once. Verified.
- (b) *Genuine collision detected before corruption*: `Some(5)` + `Some("x")` enqueue `Option$Int` and `Option$Str`. When the second drains, `enum_variant_owner["Some"]` already holds the other mangled name, and the `Err(...)` returns at `:829` **before** `m.produced.push` at `:837`. No corrupted registration. Test at `:972` confirms the clear `monomorphize:` error.
- (c) *Variant names left unmangled*: `specialize_enum` (`:318`) copies `v.name` verbatim, mangling only the enum name; `resolve_enum_instantiation` deliberately does **not** rewrite `call.name` (`:616`). So the unmodified, bare-variant-keyed match/`EnumVariants` lookup still resolves `Some`. Test at `:962` asserts `some.fields == ["Int"]` and that the generic `Option` template is dropped.

### Grammar &amp; non-generics-safety

- **Both grammar fixes real, necessary, minimal.** `inline: [$.generic_type, $.generic]` (`grammar.js:59`) confirms both are inlined — so `generics`' children *are* the flattened letter/bound nodes, making the old `parse_generic_type`-child search dead, exactly as claimed. The `enum_field` fix adds `choice($.type_identifier, $.generic)` (`:149`). Neither regresses: `implements` is a separate `commaSep1($.type_identifier)` field (`:97`) disambiguated from `generics` by case (`Stringable` vs `a`), and concrete `Some(Int)` still matches `type_identifier`. Corpus + formatter suites stay green.
- **Non-generic passthrough is a true no-op.** For a source with zero generic items, all classification maps are empty, every `resolve_*_instantiation` early-returns, `maybe_rewrite_return`'s `needs` is always false (concrete return names are capitalized, never generic-param or generic-class), and items are copied unchanged. No misclassification risk in either direction — the receiver/generic detection keys strictly off the `generics` list, `fn_generic_params`, and `enum_generic_params`.

### Documentation
README accurately now states generics are monomorphized and compiled through the unmodified pipeline, correctly attributes the `libs/std` blocker to *unrelated* unimplemented features (closures/`Nil`/decorators/colon-arrow), and documents the enum single-instantiation limitation with the `monomorphize:`-error behavior.

### Recommendations
1. Add a fast, default-running guard test (lower injectable cap) so the recursion bound isn't CI-invisible.
2. Consider avoiding the per-specialization `global_env.clone()` if inexpensive.

### Assessment

**Ready to merge?** Yes

**Reasoning:** Both high-risk fixes are sound under independent re-derivation, the grammar changes are minimal and non-regressing, the pass degenerates to a true no-op for non-generic code, and the entire workspace + tree-sitter suites are green. The only findings are minor (a quadratic clone on an error path, an ignored guard test, and a benign void-fn edge) — none block merge.</result>
<usage><subagent_tokens>93289</subagent_tokens><tool_uses>11</tool_uses><duration_ms>186410</duration_ms></usage>
</task-notification>

**Claude:** Ready to merge — both high-risk fixes independently re-verified sound, only Minor findings remain (none blocking). Updating the ledger and finishing the branch.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Skill: superpowers:finishing-a-development-branch`

**You:** Skill /superpowers:finishing-a-development-branch is already loaded above; instructions unchanged.
  - `Bash: cargo test --workspace 2>&1 | tail -5; echo ---; git branch --show-current; git status --short`

**Claude:** Tests pass, we're on `master` directly, everything already committed.
  - `AskUserQuestion`

**Claude:** Left local per your choice. That closes out this session's full arc:

1. **General enum support** — payload-free and payload-carrying enum variants now construct and match-destructure to working wasm (a real silent miscompile was caught and fixed by the final review before merge).
2. **Two pre-existing bugs fixed** — a grammar limitation on bare expressions as trailing statements, and a codegen gap where tail-position `match`/`if` without explicit `return` silently dropped its value.
3. **Generics monomorphization** — generic classes, functions, methods, and enums (single-lowercase-letter syntax) now fully specialize and compile, via a new pass that runs ahead of both the checker and codegen with zero changes needed to either. Along the way this surfaced and fixed a genuine pre-existing parser bug (`Class.generics` was never populated at all) and a second grammar gap (generic enum variant payloads didn't parse), plus caught and closed a soundness hole where specialization could have silently masked real type errors.

Known limitations, documented in README: a generic enum can only be instantiated at one concrete type per program (a second instantiation is rejected with a clear error, not silently broken); `libs/std`'s actual `list.plum`/`map.plum` still don't compile — they need closures, `Nil`/optional chaining, decorators, and other unimplemented syntax unrelated to generics.

**You:** Can you remove the extra syntax in list/map and keep it simple and standard. And also can you fix the generic enum one concrete type, does this mean we can't have more than one Option<Int> like Option<Str>?

**Claude:** Good question — yes, as it currently stands, within one compiled file you can't have both `Option$Int` and `Option$Str` simultaneously; the second instantiation hits the collision guard and fails with a clear error rather than corrupting anything. That's a real limitation for a general-purpose `Option`, since basically any realistic program wants it at multiple types.

The root cause: the existing enum runtime (from earlier work) keys `Some`/`None` by **bare variant name globally** across the whole program, not per-instantiation. To truly support `Option<Int>` and `Option<Str>` coexisting, I'd need to mangle variant names too (`Some$Int`, `Some$Str`) and rewrite every `match` pattern that references them — which Task 2 explicitly scoped out as "much bigger than this task." I can do that now as a proper follow-up.

On simplifying `list.plum`/`map.plum`: worth flagging before I start — `List`'s current draft leans on **closures** (`|v| ...`) for basically every higher-order method (`map`, `each`, `filter`, `reduce`, `sort`, `any`, `every`, `retain`, `reject`), and closures don't exist in the language at all yet (parsed by the grammar but "not wired into any reachable rule"). "Simple and standard, using only current syntax" would mean dropping essentially all of those and leaving `List` as a basic container (add/get/set/remove/length/first/last) — a big cut, not just a syntax cleanup.
  - `AskUserQuestion`

**Claude:** This is now three related follow-ups: (1) the generic-enum multi-instantiation fix, (2) implementing closures, (3) rewriting `list.plum`/`map.plum` to use only supported syntax (closures once they exist, `Option` instead of `Nil`). Each is its own real chunk of compiler work.
  - `AskUserQuestion`
  - `mcp__ccd_session__mark_chapter: Generic enum multi-instantiation fix`
  - `Skill: superpowers:brainstorming`

**You:** Skill /superpowers:brainstorming is already loaded above; instructions unchanged.

**Claude:** I already have deep context on the exact mechanism from building this earlier. Here's the design.

**Root cause and fix**

`EnumVariants` is keyed by bare variant name (`"Some"`, `"None"`) globally. To let `Option$Int` and `Option$Str` coexist, variant names need to be mangled too, consistently with their enum's own suffix: `"Some"``"Some$Int"` / `"Some$Str"`, `"None"``"None$Int"` / `"None$Str"` (yes, even payload-free variants — their runtime tag is identical either way, but their *static type* differs per instantiation, so leaving them unmangled would still create inference ambiguity).

This touches three places, all inside the already-built `monomorphize.rs`:

1. **`specialize_enum`** (already exists from the last plan) — extend it to rename each variant using the same suffix as the enum's own mangled name, not just the enum itself.
2. **Construction call-site rewriting** — currently `resolve_enum_instantiation` deliberately does *not* rewrite `call.name` for a variant constructor like `Some(5)`. That has to change: once the enum's instantiation is resolved, rewrite `"Some"``"Some$Int"` at the call site.
3. **Match-pattern rewriting** — this is the new piece. `rewrite_stmt`'s existing `Match` handling infers the subject's type but never touches the patterns themselves. It needs to: infer the subject's concrete type, check whether it's a specialized generic enum, and if so rewrite every `Some`/`None`-shaped pattern name in that match's cases to the corresponding mangled variant name, using a mangling table recorded when that enum was specialized.

With that in place, the single-instantiation collision guard becomes unnecessary and can be removed — there's no more collision since every specialization gets its own uniquely-named variants.
  - `AskUserQuestion`

**Claude:** What would you like to change about the design?

**You:** Yes loos good

**Claude:** Great — proceeding to write the spec.
  - `Bash: ls docs/superpowers/specs/`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md`

**Claude:** Self-review: no placeholders, sections consistent, out-of-scope boundary matches the prior plan. Committing.
  - `Bash: git add docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md && git commit -m "$(cat <<'EOF' docs: add design spec for generic-enum mult`

**Claude:** Spec written and committed to [2026-07-20-generic-enum-multi-instantiation-design.md](docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md). Please review before I write the implementation plan.

**You:** Yes
  - `Skill: superpowers:writing-plans`

**You:** Skill /superpowers:writing-plans is already loaded above; instructions unchanged.
  - `Read: /Users/pyrossh/Code/plum/plum-checker/src/monomorphize.rs`
  - `Bash: grep -n "instantiated at more than one concrete type\|enum_variant_owner\|resolve_enum_instantiation" plum-checker/tests/checker_tests.rs`
  - `Bash: grep -n "monomorphize\|more than one concrete" plum-checker/tests/checker_tests.rs | head -30`
  - `Read: /Users/pyrossh/Code/plum/plum-checker/tests/checker_tests.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-checker/tests/checker_tests.rs`
  - `Bash: grep -n "generic_enum\|fn generic_" plum-wasm-codegen/tests/codegen_tests.rs`
  - `Bash: grep -n "single concrete\|generic enum\|instantiation" README.md`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md`

**Claude:** Self-review caught a real issue: Step 3g left in my own scratch reasoning (a "wait, don't do this" aside) instead of a clean instruction. Fixing that now.
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md`
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md`
  - `Bash: git add docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md && git commit -m "$(cat <<'EOF' docs: add implementation plan for generic-enum mul`

**Claude:** Plan saved to [docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md](docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md).

Two execution options:

1. **Subagent-Driven (recommended)** — fresh subagent per task, review between tasks, fast iteration
2. **Inline Execution** — execute tasks in this session with checkpoints

Which approach?

**You:** 1
  - `Bash: cargo test --workspace 2>&1 | tail -5`

**Claude:** Baseline clean. Setting up a fresh ledger.
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Write: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Pre-flight scan clean (self-authored, self-reviewed). Dispatching Task 1.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-g`
  - `Agent: Implement generic enum multi-instantiation`

**Claude:** Task 1 implementer running in the background.

**You:** <task-notification>
<task-id>a9ab25fe3cb73fe68</task-id>
<tool-use-id>toolu_018xEeCG4wnNS7MmPshCdwRz</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a9ab25fe3cb73fe68.output</output-file>
<status>completed</status>
<summary>Agent "Implement generic enum multi-instantiation" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Clean build, no debug leftovers, no commit made. Everything is documented in the report.

**Status:** BLOCKED

- Commits created: none (blocker prevents a clean, non-regressing commit)
- Tests: checker-side new/updated tests pass (`generic_enum_single_instantiation_type_checks`, `generic_enum_multiple_instantiations_coexist_and_type_check` both `ok`); codegen-side fails — including the **pre-existing** `generic_enum_specialized_and_matched_runs_correctly`, now broken as a regression, plus the new coexistence test, both with `"codegen: unknown enum variant 'Some'"`.
- Concern: the brief's design (subject-type-driven match-pattern rewriting) breaks whenever a generic enum flows through an ordinary function's bare-typed parameter (e.g. `unwrapOr(o: Option, ...)`), because such params' declared type is never mangled — only locally-inferred types (from `x = Some(5)`) get resolved correctly. This is confirmed root-caused (not order-dependent — verified by reordering source), and fixing it properly requires extending the specialization machinery to treat bare-generic-enum param/return types as their own specialization axis, well beyond Task 1's stated 7-edit scope.
- Full diagnosis, evidence, and two concrete recommended paths forward: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md`</result>
<usage><subagent_tokens>172529</subagent_tokens><tool_uses>64</tool_uses><duration_ms>1082618</duration_ms></usage>
</task-notification>

**Claude:** Real design gap found — let me read the full diagnosis before deciding how to proceed.
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md`

**Claude:** The implementer found something important and correctly stopped rather than paper over it: my plan's design has a real gap that also, on reflection, likely affects generic **classes** too, not just enums — an *ordinary* function taking a bare generic-typed parameter (e.g. `unwrapOr(o: Option, default: Int) -> Int`, the completely normal way you'd write this) never gets its own declared param type resolved to a concrete specialization. The existing single-instantiation support only worked before by accident, because with only one enum instantiation globally, unmangled `Some`/`None` happened to match regardless. Enforcing mangling now breaks that — the implementer correctly caught this would be a regression and stopped instead of committing it.

The good news: the fix reuses almost everything already built. Ordinary functions whose param/return bare-names a generic class or enum need to become a specialization target themselves (like real generic functions already are) — and since `specialize_fn`'s substitution mechanism already works on any bare type name, no new substitution logic is needed, just a new way to *detect* this shape and resolve it from the call site's argument type.
  - `AskUserQuestion`

**Claude:** Given the scope, let me check what's currently sitting uncommitted in the working tree before deciding how to restructure the plan.
  - `Bash: git status --short; echo ---; git diff --stat`

**Claude:** Here's the design for the real fix (covering both classes and enums uniformly), building on what's already correctly implemented and sitting uncommitted:

**Mechanism**: An "ordinary" function/method (no lowercase-letter generic params) whose param or return type bare-names a *generic* class or enum (e.g. `unwrapOr(o: Option, default: Int) -> Int`, `sumBox(b: Box) -> Int`) gets treated as its own specialization target — mirroring exactly how a true generic function already works:

- **Return position**: already half-solved — `maybe_rewrite_return` already resolves a bare generic-*class* return (`-> Box`) from the body's inferred tail type; it just never checked the enum registry. One-line fix: add the enum check alongside the existing class check.
- **Param position** (the actual blocker): new detection (`fn_bare_generic_refs`) finds params whose declared type bare-names a generic class/enum. Such a function gets deferred and specialized per call site — resolving the substitution from the *calling* argument's already-rewritten concrete type (e.g. `unwrapOr(Some(13), ...)``o`'s binding becomes `Option → Option$Int`), mangling the function itself (`unwrapOr$Option$Int`), and reusing `specialize_fn` **unchanged** — its substitution mechanism already works on any bare type name, not just single letters.

This is scoped to free functions for now (a method introducing this shape is a further edge case, consistent with the existing "extra method-level generics" out-of-scope note).
  - `AskUserQuestion`

**Claude:** What would you like to change?

**You:** Do it write the plan
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md`

**Claude:** I'll add a section documenting the newly-discovered gap and its fix.
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md`
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md`
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md`
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md`

**Claude:** Committing the spec amendment.
  - `Bash: git add docs/superpowers/specs/2026-07-20-generic-enum-multi-instantiation-design.md && git commit -m "$(cat <<'EOF' docs: amend generic-enum spec with bare-gen`

**Claude:** Let me read the current (already-modified) state of `monomorphize.rs` to write precise instructions for the new plan.
  - `Read: /Users/pyrossh/Code/plum/plum-checker/src/monomorphize.rs`
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md`

**Claude:** Self-review: no placeholders, new fields/methods are consistent with usage across all edit points, arg-count handling mirrors existing precedent (`resolve_fn_instantiation`'s zip pattern) rather than introducing a new inconsistency. Committing.
  - `Bash: git add docs/superpowers/plans/2026-07-20-generic-enum-multi-instantiation.md && git commit -m "$(cat <<'EOF' docs: revise generic-enum plan with bare-generic-t`
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Regenerating the task brief and dispatching.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-g`
  - `Agent: Implement bare-generic-typed-param specialization`

**Claude:** Task 1 implementer running in the background on Opus, given the stakes of getting this architecturally correct.

**You:** <task-notification>
<task-id>a22dd252d3746f9e7</task-id>
<tool-use-id>toolu_01UdHHsX3B3yLEXN3e22Scxg</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a22dd252d3746f9e7.output</output-file>
<status>completed</status>
<summary>Agent "Implement bare-generic-typed-param specialization" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Committed successfully as `2216237`.

**Status:** DONE

- **Commit:** `2216237` feat(plum-checker): support multiple concrete instantiations of the same generic enum (includes both the already-uncommitted variant-mangling work and this task's new bare-generic-ref specialization mechanism, per Step 7)
- **Tests:** all green — checker 39 passed/1 ignored, codegen 48 passed, full `cargo test --workspace` clean, tree-sitter 22/22 parses
- **Pre-existing regression fixed:** `generic_enum_specialized_and_matched_runs_correctly` went RED→GREEN with no change to the test itself — the actual proof the blocker is resolved
- **Concerns:** none. All 7 edits (3a-3g) applied verbatim; no changes to `specialize_fn` needed (its substitution already handles bare generic keys); per-call-site mangling correctly disambiguates the same bare-typed function called at two concrete types
- **Report:** /Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md

One note: the commit landed with a locally-configured git identity (Peter John / pyrossh@...) rather than an explicit one — flagging since git emitted the "configured automatically" warning, though the Co-Authored-By trailer is present.</result>
<usage><subagent_tokens>74597</subagent_tokens><tool_uses>32</tool_uses><duration_ms>260908</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the reviewer with heavy scrutiny given the stakes of this fix.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package e698d53`
  - `Agent: Review Task 1 (spec + quality)`

**Claude:** Task 1 reviewer running in the background.

**You:** <task-notification>
<task-id>af9b305d2dd833dc5</task-id>
<tool-use-id>toolu_01JZy6UuJBnn5HxK6EtUFJBP</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/af9b305d2dd833dc5.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 1 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>I have everything I need. All key claims verified against the actual diff and source.

## Spec Compliance
- ✅ Spec compliant. All 7 edits (3a–3g) applied as specified; both new checker tests and the new codegen test present; README updated; the pre-existing codegen regression test is genuinely unmodified.

## Strengths
- The pre-existing regression test `generic_enum_specialized_and_matched_runs_correctly` (codegen_tests.rs:759–777) is confirmed **byte-for-byte unchanged** — the diff hunk `@@ -769,10 +769,66 @@` contains only additions, no deletions, in that region. It is exactly the `unwrapOr(o: Option, default: Int)` bare-generic-enum shape and executes via wasmtime `run_main` returning 13. This is real proof the blocker is fixed by `monomorphize.rs` changes alone.
- The fix genuinely reuses `PendingSpecialization::Fn`/`specialize_fn`/worklist unchanged; only detection (`fn_bare_generic_refs`) and resolution (`resolve_bare_generic_fn_instantiation`) were added, as the brief intended.
- The class case (`ordinary_function_with_bare_generic_class_param_runs_correctly`, line 817) is a real end-to-end wasm execution test (`run_main == 11`), not a compile-only check.

## Verification of the critical risks

**Risk 1 (same function, two instantiations — aliasing):** Traced correct. `resolve_bare_generic_fn_instantiation` infers arg types per call site (args are already rewritten to `Some$Int`/`Some$Str` before infer runs, and `enum_variants` was taught the mangled variants, so infer yields `TNamed("Option$Int")` vs `TNamed("Option$Str")`). Each site produces a distinct `mangled` name, a fresh local `bindings` map, and a separate worklist entry deduped by mangled name. `specialize_fn` emits independent copies. No shared mutable state, no aliasing. *Caveat:* the codegen test exercises two **differently-named** functions (`unwrapIntOr`/`unwrapStrOr`), not literally the same function at two instantiations — so risk #1's exact scenario is reasoned-correct but not directly test-covered (minor).

**Risk 2 (classification mutual exclusivity):** The second-loop `else if` chain (`receiver_is_generic``fns_generic``fns_bare_generic`) is strictly mutually exclusive; a truly-generic-AND-bare-referencing function lands only in `fns_generic`, matching the out-of-scope constraint. The third-loop pass-through guard `!receiver_is_generic &amp;&amp; !is_generic_fn &amp;&amp; !is_bare_generic_fn` is consistent with the classification, so bare-generic fns are excluded from pass-through and only emitted via specialization. No double-classification, no gap.

**Risk 3 (class + enum symmetry):** `fn_bare_generic_refs` checks both `classes_generic` and `enums_generic_by_name`. Class-side test genuinely executes. Symmetric.

**Risk 4 (regression actually fixed):** Verified above — test unchanged, passes via the monomorphize change. The checker-side tests that *were* modified (`..._matched_type_checks` assertion, and the rename `multi_instantiation_is_a_clear_error``multiple_instantiations_coexist_and_type_check`) belong to the bundled variant-mangling work, which the brief explicitly authorizes; the constraint "must not change that test" targets the codegen test, which is intact.

**Risk 5 (borrow/logic of the edits):** `resolve_enum_instantiation` correctly changed to `&amp;mut ast::FnCall`; sequential `&amp;mut` borrows in the `FnCall` arm are fine. The eager `enum_variant_mangling`/`enum_variants` population borrows the `&amp;'a` enum and mutates disjoint `self` fields — sound. The final `call.name = self.enum_variant_mangling[&amp;mangled][&amp;variant.name]` is safe because the table was just populated with every variant. No borrow-checker or indexing hazard.

## Issues

#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- The literal "same free function called at two different concrete instantiations" case (e.g. `unwrapOr(Some(5))` and `unwrapOr(Some("x"))` in one program) is reasoned-correct but not directly covered by a test — the multi-instantiation codegen test uses two distinct function names instead. A one-function-two-call-sites test would close the last gap on risk #1. Not blocking; the mangling logic is demonstrably per-call-site.
- An uncalled bare-generic function is silently dropped (excluded from pass-through, never enqueued). This is consistent with how truly-generic functions already behave and is acceptable, but worth noting as intentional.

## Assessment
**Task quality:** Approved
**Reasoning:** All 7 edits match the brief, the classification chain is provably exclusive, class/enum are handled symmetrically, and the pre-existing codegen regression test is verified unmodified and re-passing via the monomorphize change alone — the actual blocker is genuinely fixed, not just reported fixed. Only minor test-coverage gaps remain.</result>
<usage><subagent_tokens>60315</subagent_tokens><tool_uses>5</tool_uses><duration_ms>119009</duration_ms></usage>
</task-notification>

**Claude:** Approved, but given this is the highest-risk part of the whole fix, let me close that one test-coverage gap (the literal "same function, two instantiations" case) for extra confidence before finishing.
  - `Agent: Add same-function-two-instantiations test`
  - `mcp__ccd_session__dismiss_task: placeholder`

**Claude:** Test-coverage fix running in the background.

**You:** <task-notification>
<task-id>aa775880e02d1eab8</task-id>
<tool-use-id>toolu_01Y4d8ikHcs3HkbyqD3ZthaV</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/aa775880e02d1eab8.output</output-file>
<status>completed</status>
<summary>Agent "Add same-function-two-instantiations test" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>---

**Status:** DONE

**Commit created:** b6a9042 - test(plum-wasm-codegen): cover same bare-generic-enum-param function at multiple instantiations

**What the test proves:** The same generic function `unwrapOr` taking a bare generic-enum-typed param, called twice at the same concrete instantiation (`Option$Int`), correctly deduplicates and specializes without aliasing. Cross-type disambiguation (calling at both `Option$Int` and `Option$Str`) cannot be type-safely tested here since `unwrapOr`'s return type is fixed to `Int`; the test validates same-type reuse/dedup per the review's reasoning.

**Test summary:** `unwrapOr(Some(5), 0) + unwrapOr(Some(37), 0)` expects 42; all 49 codegen + 57 workspace tests pass.</result>
<usage><subagent_tokens>39039</subagent_tokens><tool_uses>8</tool_uses><duration_ms>90067</duration_ms></usage>
</task-notification>

**Claude:** Good — the implementer correctly recognized the cross-type disambiguation couldn't be tested cleanly given `unwrapOr`'s fixed `Int` return, and fell back to the honest same-type-reuse proof (which the review's own reasoning already covered analytically). Updating the ledger and moving to the final review.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: git log --oneline aedebc1~1..b6a9042 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review`
  - `Agent: Final whole-branch review of generic-enum multi-instantiation fix`

**Claude:** Final whole-branch review running in the background on Opus.

**You:** <task-notification>
<task-id>aa7860c5c0709764f</task-id>
<tool-use-id>toolu_01VfPrFMiE4N1HTre8G4qsia</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/aa7860c5c0709764f.output</output-file>
<status>completed</status>
<summary>Agent "Final whole-branch review of generic-enum multi-instantiation fix" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>I have completed my independent verification. Every claim traces through correctly, all tests pass, and the git history is clean.

---

## Code Review: Generic Enum Multi-Instantiation (aedebc1..b6a9042)

### Strengths

- **The pre-existing regression is genuinely fixed — independently confirmed.** `generic_enum_specialized_and_matched_runs_correctly` (codegen_tests.rs:759) is byte-for-byte unmodified (the diff's `@@ -769,10 +769,92 @@` hunk is pure append; the test body still ends at `assert_eq!(run_main(&amp;bytes), 13)`) and passes. I traced the full mechanism rather than trusting the report:
  1. `unwrapOr(o: Option, ...)` is classified into `fns_bare_generic` because `fn_bare_generic_refs` finds `Option` in `enums_generic_by_name` (monomorphize.rs:719).
  2. It is therefore excluded from direct pass-through (the `is_bare_generic_fn` guard, monomorphize.rs:738).
  3. At the call `unwrapOr(Some(13), 0)`, args are rewritten **first** (`Some(13)``Some$Int`, monomorphize.rs:609-616), so when `resolve_bare_generic_fn_instantiation` runs it infers the argument's *already-rewritten* type `TNamed("Option$Int")` — not the stale `Option` — binds `{Option: Option$Int}`, mangles to `unwrapOr$Option$Int`, and reuses `PendingSpecialization::Fn`/`specialize_fn` unchanged (monomorphize.rs:516-522).
  4. `specialize_fn` substitutes the param to `o: Option$Int`; the body's `match o` then infers subject `Option$Int`, finds the mangling table, and rewrites `Some`/`None` patterns to `Some$Int`/`None$Int` — matching the specialized enum's mangled variants. The key that "could never match" now matches. This is a real fix, not a report artifact.

- **Classification chain is sound and mutually exclusive.** The three-way `else if` (methods_generic_on / fns_generic / fns_bare_generic, monomorphize.rs:714-721) cannot double-classify. A truly-generic function that *also* bare-references (out-of-scope) lands in `fns_generic` (branch 2 wins); a method with a bare-ref param (out-of-scope) falls through to normal pass-through. Both out-of-scope shapes degrade to a downstream checker error (a bare, dropped `Option`/`Box` referenced by a non-existent template), **not** silent corruption — consistent with the documented scope.

- **Dedup is correct.** `same_bare_generic_enum_param_function_called_multiple_times_runs_correctly` passes (5+37=42); the second call site finds `unwrapOr$Option$Int` already in `enqueued` and does not re-enqueue, but still rewrites `call.name` to the shared mangled name. One specialization, two references.

- **Symmetry is real, not claimed.** `fn_bare_generic_refs` checks both `classes_generic` and `enums_generic_by_name` (monomorphize.rs:476), and `ordinary_function_with_bare_generic_class_param_runs_correctly` is a genuine wasmtime execution test (`compile_source` + `run_main`, expects 11), not a compile-only check. Traced the class path: `sumBox$Box$Int` specializes correctly and dispatches `getBoxValue` on the mangled receiver.

- **Part 1 (variant mangling) is correct and the collision-guard removal is safe.** Construction sites (monomorphize.rs:596) and both `CasePattern::Name` and `CasePattern::Class` match arms (monomorphize.rs:373-392) are rewritten. The removed `enum_variant_owner` guard is genuinely unnecessary: distinct instantiations yield distinct mangled variants (`Some$Int` vs `Some$Str`) and same-instantiation is deduped, so no flat-table collision can recur. A subtle-but-important detail present in the committed code (and *not* shown in the plan snippets) — the eager population of `self.enum_variants` with mangled `EnumVariantInfo` including correct `tag`/`field_types` (monomorphize.rs:573-590) — is what keeps inference on rewritten construction sites working; its `enumerate()` tag assignment correctly mirrors declaration order, matching `specialize_enum`.

- **Process discipline.** The prior BLOCKED self-report (rather than committing a regression) and the plan revision in response are exactly right. Git history is clean: linear on `master`, no orphaned branches, no half-applied state, working tree clean (only untracked `.recall/`).

- **Full `cargo test --workspace` is green** (checker 39 pass/1 ignored, codegen 49 pass, all other crates pass). Only warnings are pre-existing tree-sitter C scanner warnings, unrelated.

### Issues

#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)

- **checker_tests.rs (~line 511-513, diff 1445-1447):** three consecutive blank lines were inserted between `generic_function_...rejected` and the two new tests. Cosmetic; collapse to the usual single blank line.

- **monomorphize.rs:472-474 (`fn_bare_generic_refs`):** the check keys only on `t.name`, ignoring `t.generics`. A param written with explicit args, e.g. `o: Option[Int]`, would also be classified as bare-generic and resolved from the argument rather than the annotation. Harmless today (no such usage; the argument-inferred type would agree) and arguably desirable, but worth a one-line comment noting the annotation's explicit type args are intentionally ignored in favor of call-site inference, so a future reader doesn't mistake it for a bug.

### Recommendations

- **The test-coverage honesty holds up.** `same_bare_generic_enum_param_function_called_multiple_times_runs_correctly` and its comment accurately describe *same-type reuse/dedup*, not cross-type disambiguation — no overselling. I independently confirmed the untested cross-type-same-function-name case is reasoned-correct: `mangle` incorporates the concrete type arg, and `specialized`/`enqueued` are keyed on the full mangled name, so `unwrapOr$Option$Int` and `unwrapOr$Option$Str` would be two independent specializations with separately-substituted params and independently-rewritten match tables — no shared mutable state to alias. The stated reason it can't be directly tested (the return type would differ, and the language lacks Str operations to build a meaningfully different Int-returning body) is legitimate. `generic_enum_multiple_instantiations_coexist_and_run_correctly` already proves `Option$Int` and `Option$Str` coexist across *different* function names, which exercises the same mangle/dedup machinery.

- **Documentation is accurate.** README now states generic enums support any number of coexisting instantiations (with the internal `Some$Int`/`Some$Str` mangling noted as invisible to user code) and correctly narrows the residual limitation to a payload-free variant used as a bare value *outside* a `match` pattern. This matches the implemented behavior exactly.

### Assessment

**Ready to merge?** Yes

**Reasoning:** The previously-broken regression is genuinely fixed — verified by independent end-to-end tracing and a passing unmodified test, not by trusting the report — the bare-generic-ref mechanism is complete with no classification gaps that cause silent corruption, and the full workspace suite is green. Remaining items are two cosmetic/documentation-level minors.</result>
<usage><subagent_tokens>94088</subagent_tokens><tool_uses>8</tool_uses><duration_ms>199274</duration_ms></usage>
</task-notification>

**Claude:** Ready to merge — no Critical or Important findings, both remaining Minor items are cosmetic. Updating the ledger.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Both fixes clean. Wrapping up the branch.
  - `Skill: superpowers:finishing-a-development-branch`

**You:** Skill /superpowers:finishing-a-development-branch is already loaded above; instructions unchanged.
  - `Bash: cargo test --workspace 2>&1 | tail -5; echo ---; git branch --show-current; git status --short`

**Claude:** Tests pass, on `master` directly, everything committed.
  - `AskUserQuestion`

**Claude:** Left local. Moving on to the next piece: implementing closures.
  - `mcp__ccd_session__mark_chapter: Closures implementation`
  - `Skill: superpowers:brainstorming`

**You:** Skill /superpowers:brainstorming is already loaded above; instructions unchanged.
  - `Bash: grep -n "closure\|fn_type\b" tooling/tree-sitter-plum/grammar.js plum-core/src/ast.rs plum-core/src/parser.rs`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`

**Claude:** This is genuinely the biggest single feature request in this session — comparable in scope to the enum or generics work, maybe bigger, since wasm has no native closures. Let me lay out what's actually needed.

**Current state**: the grammar already has a `closure: "|" params "|" body` rule, but it's commented out of every reachable path — completely unwired. There's also no function-type annotation syntax at all (`cb: fn(v: a) -> Bool`, as used in the aspirational `list.plum`, doesn't parse today).

**What real closures require:**
1. **Grammar**: wire `closure` into expression position, and add a new function-type annotation syntax for param/field types (so `each(cb: fn(v: a))` can even be declared).
2. **AST/checker**: a new `Expr::Closure` node; `PlumType::TFun` already exists and mostly covers this — calling a closure-typed parameter by name should already type-check via the existing `FnCall` inference path.
3. **Codegen — the hard part**: wasm has no closures natively. The natural design, reusing patterns already built for classes/enums: each closure literal becomes a real wasm function (registered in a new function table), and a closure *value* at runtime is a single `i32` pointer to a heap-allocated `{table_index, captured_env_pointer}` pair — calling it means loading both fields and doing `call_indirect`, passing the env pointer as an implicit first argument (mirroring how methods already pass `self`). This needs: a wasm table + elements section (not yet supported by the current `WasmModule` wrapper), free-variable analysis to know what each closure needs to capture, and boxing captured variables into a heap struct.

This is genuinely real, capturing closures (the actual `list.plum` usage — e.g. `this.each() |v| { res.write(v, sep) }` — captures `res`/`sep` from the enclosing method), not just non-capturing function pointers.
  - `AskUserQuestion`
  - `AskUserQuestion`

**Claude:** A few more design specifics, then I'll write the spec.

**Function-type syntax**: I'll simplify to positional types only (no param names needed in a type annotation) — `fn(Int) -> Bool`, `fn(a) -> b` — rather than `list.plum`'s aspirational `fn(v: a) -> Bool` which redundantly names params in a type position.

**Free-variable capture**: each closure literal becomes its own wasm function with an implicit first param (the boxed captured-environment pointer, mirroring how methods already pass `self`). At the closure literal's call site, free-variable analysis finds every referenced name that isn't a closure param/global/function, boxes the *current* values of those variables into a heap struct (same bump-allocator convention as class fields), and produces the closure value as a single `i32` pointer to `{table_index, env_pointer}`.

**Why this doesn't need new monomorphization machinery**: a function taking a `fn(...)`-typed callback param (like `each`) doesn't need re-specializing per *closure*`call_indirect`'s type signature is resolved from the *outer* function's own concrete instantiation (already handled by existing monomorphization), and any closure passed to that specific call site is checked to match. The closure's captured-variable layout is internal/opaque to the caller.
  - `AskUserQuestion`

**Claude:** Let me present the complete design now, then write it up.

**1. Grammar**

- Uncomment `$.closure` in `expression`'s choice list — the rule already exists (`"|" params "|" body`), it's just unreachable today.
- Add a new function-type annotation rule, e.g. `fn_value_type: "fn" "(" commaSep(type) ")" ("->" type)?`, usable wherever a param type currently appears.

**2. AST**

- New `ast::Closure { params: Vec<String>, body: Block }`, added as `Expr::Closure(Box<Closure>)`.
- New `ParamType::Fn(Vec<Type>, Option<Box<Type>>)` variant (params only — class fields storing a callback are out of scope, not needed by any current example).

**3. Checker**

- `PlumType::TFun` already exists and fits directly. `infer_expr`'s new `Closure` arm: bind each closure param to a fresh `TVar` (matching this checker's existing permissive style — it never really unifies generics either), infer the body, return `TFun(param_types, body_type)`.
- Calling a closure-typed parameter by name (`cb(x)`) already type-checks via the *existing* `FnCall` inference path — `lookup(env, &call.name)` doesn't care whether the binding came from a top-level function or a local closure-typed param.

**4. Codegen**

- New wasm sections: a function table + elements section (`WasmModule` needs `add_table`/element-segment support — doesn't exist yet).
- Every closure literal compiles to its own real wasm function, registered in the table, with an implicit first param (the boxed captured-environment pointer) — exactly mirroring how methods already receive `self` as an implicit first param.
- At the closure literal's site: free-variable analysis (walk the body, collect every referenced name that isn't a closure param/global/function) → bump-allocate an env struct holding the current values of those variables (same convention as class fields) → bump-allocate a `{table_index, env_pointer}` pair → that pointer is the closure's runtime value (a single `i32`, consistent with every other reference-shaped value in this compiler).
- Calling a closure value: load `table_index`/`env_pointer` from the pointer, push `env_pointer` then the real arguments, then `call_indirect` using the type index resolved from the *caller's* own (already-monomorphized) concrete signature — no new monomorphization mechanism needed, since a closure's internals are opaque to its caller.

**5. Testing plan**

- Grammar: corpus tests for a closure literal parsing in expression position, and a `fn(...)`-typed param.
- Checker: a closure's inferred type flows correctly into a `TFun`-typed binding; calling a closure param type-checks.
- Codegen: a non-capturing closure passed and called; a closure capturing one enclosing variable, called after the enclosing scope's local has changed (proving it captured a snapshot, not a live reference — need to decide semantics here, see below); a closure passed through an already-generic higher-order function.
- Examples: a minimal, real usage (e.g. a simple `each`-like helper) exercised end-to-end.

**6. Out of scope**

- Class fields storing a closure value.
- Closures capturing by mutable reference (captures are a snapshot of the variable's value at closure-creation time, not a live binding — simpler, and matches what "boxing current values into a struct" naturally gives; mutation of a captured variable after closure creation won't be visible inside the closure). If you need live/shared mutable capture semantics, flag it now — that's a materially different (and harder) design.
- Multiple closures with genuinely different captured-variable sets passed to the *same* call site needing different wasm function signatures beyond the env-pointer convention — not expected to arise given the design, but noting as an edge to watch.
  - `AskUserQuestion`
  - `AskUserQuestion`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-closures-design.md`

**Claude:** Self-review: no placeholders, sections consistent (capture semantics stated once and referenced, not restated differently), scope boundaries explicit. Committing.
  - `Bash: git add docs/superpowers/specs/2026-07-20-closures-design.md && git commit -m "$(cat <<'EOF' docs: add design spec for closures Full, capturing closures (snapsh`

**Claude:** Spec written and committed to [2026-07-20-closures-design.md](docs/superpowers/specs/2026-07-20-closures-design.md). Please review it before I write the implementation plan.

**You:** Generate (or overwrite) this project's `context.md` summary now, using Recall's
local offline summarizer.

Run with the Bash tool:

```
python3 "/Users/pyrossh/.claude/plugins/cache/recall/recall/0.4.0/scripts/make_context.py" || python "/Users/pyrossh/.claude/plugins/cache/recall/recall/0.4.0/scripts/make_context.py"
```

(`make_context.py` defaults to the current directory, so no `--cwd` is needed —
and on Windows passing the shell's `$(pwd)` would hand it a `/c/...` path that
doesn't match the transcript dir. The `|| python` fallback covers Windows, where
the interpreter is normally `python`, not `python3`.)

Then report back:

- On success it prints the path it wrote and which summarizer path ran
  (numpy-accelerated TextRank if numpy is present, otherwise the pure-Python
  TextRank — both vendored, no install needed). Confirm the save and read back the
  **Goal** and **Where we left off** lines from `.recall/context.md`.
- If it errored, surface the exact message.
  - `Read: /Users/pyrossh/Code/plum/.recall/context.md`
  - `Bash: grep -n "Where we left off" -A6 /Users/pyrossh/Code/plum/.recall/context.md`

**Claude:** Recall context saved to `.recall/context.md` (numpy-accelerated TextRank summarizer).

- **Goal:** implementing the standard library (List/Map/Enum) — which led into general enum support, generics monomorphization, and now closures.
- **Where we left off:** spec written and committed for closures ([2026-07-20-closures-design.md](docs/superpowers/specs/2026-07-20-closures-design.md)), waiting on your review before I write the implementation plan.

Ready to continue whenever you are — just let me know if the closures spec looks good.

**You:** Looks good, but the List and Map libs/std have different syntax to bool.plum do you know why? Can we make them same?
  - `Bash: cat libs/std/bool.plum`
  - `Bash: for f in libs/std/*.plum; do echo "=== $f ==="; grep -nE "this\.|self\.|:=|\?\.|@[A-Za-z]|\bNil\b|\|.*\|.*=>|fn\(|\[[A-Z]\]\(|-> [A-Za-z]+\[|: Option\[|\.\.\."`
  - `Bash: echo "--- bool/int/float/option/result (clean, currently-parseable) ---"; grep -n "\bthis\b" libs/std/bool.plum libs/std/int.plum libs/std/float.plum libs/std/o`
  - `Bash: grep -n "\"this\"\|'this'\|self:" tooling/tree-sitter-plum/grammar.js; grep -n "Self_\|\"self\"" plum-core/src/parser.rs plum-core/src/ast.rs tooling/tree-sitte`

**Claude:** Good catch — let me clarify what's actually going on, since it's not quite "list/map differ from bool" so much as "none of `libs/std` fully matches current syntax, just to different degrees."

I checked: `this` (used throughout `bool.plum`, `int.plum`, `float.plum`, `result.plum`) **isn't a real keyword either** — only `self` is defined in the grammar. `bool.plum` also uses multi-subject `match this, o` (a documented, still-unimplemented gap) and `result.plum` has an `if PATTERN = expr` conditional that isn't valid `if` syntax at all. So `bool.plum` is closer to working, but not actually clean either — it just has far fewer distinct unsupported features than `list.plum`/`map.plum`, which lean heavily on closures, `Nil`, `?.` optional chaining, `:=`, `@decorators`, and colon-arrow return types — syntax from the *aspirational future language* the website docs describe, not what's implemented.

So yes, we can absolutely make them consistent — normalize **all** of `libs/std` to one syntax: `this``self` everywhere, rewrite multi-subject matches, fix the invalid `if`-pattern sugar in `result.plum`, and (once closures land, per the spec we just wrote) bring `list.plum`/`map.plum` down from the aspirational syntax to the real, currently-supported subset.
  - `AskUserQuestion`

**Claude:** Good, noted for when we get to the `libs/std` normalization step. Proceeding to write the closures implementation plan now.
  - `Skill: superpowers:writing-plans`

**You:** Skill /superpowers:writing-plans is already loaded above; instructions unchanged.
  - `Bash: grep -n "\"fn\"\|TableSection\|ElementSection\|CallIndirect\|Elements::" ~/.cargo/registry/src/*/wasm-encoder-*/src/*.rs 2>/dev/null | head -30 echo --- find /`
  - `Bash: cat plum-wasm-codegen/Cargo.toml | grep -i wasm-encoder`
  - `Bash: find ~/.cargo/registry/src -maxdepth 1 -iname "*wasm-encoder*" 2>/dev/null`
  - `Bash: find / -maxdepth 6 -ipath "*registry/src*" -iname "*wasm-encoder*" 2>/dev/null | head -5`
  - `Bash: cargo metadata --format-version 1 2>/dev/null | python3 -c " import json,sys d=json.load(sys.stdin) for pkg in d['packages']: if pkg['name']=='wasm-encoder': pr`
  - `Bash: grep -A2 'name = "wasm-encoder"' Cargo.lock`
  - `Bash: cd plum-wasm-codegen && cargo tree -p wasm-encoder 2>&1 | head -5`
  - `Bash: cd /Users/pyrossh/Code/plum && cargo tree -p plum-wasm-codegen -i wasm-encoder 2>&1 | head -10; echo ---; cargo tree -p plum-wasm-codegen 2>&1 | grep wasm-encod`
  - `Bash: grep -rn "pub struct TableSection\|impl TableSection\|pub fn table\b" /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/wasm-encoder-0.220.1/s`
  - `Read: /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/wasm-encoder-0.220.1/src/core/tables.rs`
  - `Bash: grep -rn "pub struct ElementSection\|pub fn active\|pub fn declared\|pub enum Elements\|Elements::" /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6`
  - `Read: /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/wasm-encoder-0.220.1/src/core/elements.rs`
  - `Bash: grep -rn "CallIndirect" /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/wasm-encoder-0.220.1/src/core/code.rs 2>/dev/null | head -10`
  - `Read: /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/wasm-encoder-0.220.1/src/core/code.rs`
  - `Bash: grep -n "struct WasmModule\|impl WasmModule\|pub fn add_\|fn finish\|module.section" plum-wasm-codegen/src/lib.rs | head -40`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Bash: grep -n "struct CompileCtx\|struct LocalCtx\|struct Collector\|struct ModuleState\|fn compile_fn_body\|fn compile_expr\|func_ids:\|func_sigs:\|classcall_scratch`
  - `Bash: grep -n "^\s*param:\|variadic_type:" tooling/tree-sitter-plum/grammar.js`
  - `Read: /Users/pyrossh/Code/plum/tooling/tree-sitter-plum/grammar.js`

**Claude:** One more scoping clarification before I write the plan: `list.plum`'s draft syntax uses "trailing closure" call syntax (`this.each() |v| ...`, the closure attached after a completed call) — that's a different, more elaborate grammar feature (Ruby/Kotlin-style trailing blocks) than a closure used as a normal expression/argument. Since `list.plum` will get rewritten to standard syntax later anyway, I'll scope this closures feature to closures as **ordinary expressions** — passable as a normal call argument like `each(|v| ...)` — not the trailing-block calling convention.
  - `AskUserQuestion`

**Claude:** Let me amend the spec with this scoping decision, then write the plan.
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-closures-design.md`
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/specs/2026-07-20-closures-design.md`
  - `Bash: git add docs/superpowers/specs/2026-07-20-closures-design.md && git commit -m "docs: scope closures spec to ordinary-argument call syntax, not trailing-closure"`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Bash: grep -n "ast::Expr::FnCall(call) =>" plum-wasm-codegen/src/lib.rs`
  - `Read: /Users/pyrossh/Code/plum/plum-wasm-codegen/src/lib.rs`
  - `Write: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-closures.md`

**Claude:** Let me fix a real gap I left in my own plan — Task 2's `parse_fn_value_type` had an ambiguous placeholder for disambiguating param types from the return type. I'll fix this properly using explicit grammar fields (the established idiom already used elsewhere in this grammar), removing the ambiguity entirely.
  - `Read: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-closures.md`
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-closures.md`
  - `Edit: /Users/pyrossh/Code/plum/docs/superpowers/plans/2026-07-20-closures.md`
  - `Bash: grep -rn "impl.*PartialEq.*for Node\|derive.*PartialEq" ~/.cargo/registry/src/*/tree-sitter-0.2*/src/lib.rs 2>/dev/null | head -5 python3 -c " import subprocess`
  - `Bash: cargo metadata --format-version 1 2>/dev/null | python3 -c " import json,sys d=json.load(sys.stdin) for pkg in d['packages']: if pkg['name']=='tree-sitter' and`
  - `Bash: grep -n "impl.*PartialEq.*Node\|pub fn id(" /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/tree-sitter-0.26.11/binding_rust/lib.rs 2>/dev/n`
  - `Read: /Users/pyrossh/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/tree-sitter-0.26.11/binding_rust/lib.rs`

**Claude:** Fixed. Now committing the plan.
  - `Bash: git add docs/superpowers/plans/2026-07-20-closures.md && git commit -m "$(cat <<'EOF' docs: add implementation plan for closures Six tasks: wire the already-exi`

**Claude:** Plan saved to [docs/superpowers/plans/2026-07-20-closures.md](docs/superpowers/plans/2026-07-20-closures.md) — 6 tasks. Task 5 (the core codegen: free-variable analysis, per-closure wasm functions, `call_indirect` dispatch) is flagged as the hardest, most judgment-requiring piece, similar to how the generics-monomorphization plan's hardest task needed a couple of iteration rounds.

Two execution options:

1. **Subagent-Driven (recommended)** — fresh subagent per task, review between tasks, fast iteration
2. **Inline Execution** — execute tasks in this session with checkpoints

Which approach?

**You:** 1
  - `Bash: cargo test --workspace 2>&1 | tail -5`

**Claude:** Baseline clean. Setting up a fresh ledger.
  - `Write: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`

**Claude:** Pre-flight scan clean (self-authored, self-reviewed, fixed the one issue found). Dispatching Task 1.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-c`
  - `Agent: Implement Task 1: closure grammar`

**Claude:** Task 1 implementer running in the background.

**You:** <task-notification>
<task-id>a5fb37e657affae0e</task-id>
<tool-use-id>toolu_014J9DQokxG51pKtofCC9KY3</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a5fb37e657affae0e.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 1: closure grammar" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

Commit: `906b92e` — "feat(tree-sitter-plum): wire closure literals into expression position; add fn(...) type syntax"

Test summary: tree-sitter corpus 26/26 passing (incl. 4 new cases, no ERROR/MISSING nodes, fields verified with `--show-fields`); `cargo test --workspace` fully green (no Rust code touched).

Concerns: none — no `conflicts` entry was needed for the `|` closure-delimiter vs. bitwise-or ambiguity the brief flagged as possible; `tree-sitter generate` resolved it cleanly on its own (closure only reachable at `expression`, not `primary_expression`, where the `|` binary operator lives).

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-1-report.md`</result>
<usage><subagent_tokens>59150</subagent_tokens><tool_uses>26</tool_uses><duration_ms>201734</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 1 reviewer.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 038eebd`
  - `Agent: Review Task 1 (spec + quality)`

**Claude:** Task 1 reviewer running in the background.

**You:** <task-notification>
<task-id>a2122ad3c2b407a72</task-id>
<tool-use-id>toolu_01D8huSEUsuWPBNsikGam744</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a2122ad3c2b407a72.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 1 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>### Spec Compliance
- ✅ Spec compliant

**Verification performed:**
- `tooling/tree-sitter-plum/grammar.js:275``$.closure` uncommented in `expression`'s choice list, exact position/text matches brief step 1.
- `tooling/tree-sitter-plum/grammar.js:31-42` (diff hunk) — `fn_value_type` rule added with `field("params", optional(commaSep1($.type)))` and `optional(seq("-&gt;", field("returns", $.type)))`, wired into `param`'s type choice as `choice($.type, $.variadic_type, $.fn_value_type)` — text matches the brief's prescribed rule verbatim.
- Confirmed the `params`/`returns` field-based disambiguation is genuine, not cosmetic: `field()` wrapping a `commaSep1($.type)` repetition is the same idiom already used for `fn`'s own `field("params", seq("(", optional(commaSep1($.param)), ")"))` (grammar.js line ~121, unchanged in this diff) — established precedent in this codebase that the field name propagates to every matched child, so Task 2's `node.child_by_field_name("returns")` / `children_by_field_name("params")` will work unambiguously.
- Traced the reported "no conflict needed" claim structurally: `closure`'s leading token is bare `"|"` only reachable from `expression`'s choice list at expression-start position; `binary_operator`'s `"|"` (bitwise-or) only appears *after* a already-parsed `field("left", $.primary_expression)`, inside `primary_expression`, a sibling arm of the same `expression` choice, not a shared parser state. The asymmetry (leading vs. infix position) is a legitimate LR disambiguator — the claim holds up under grammar reading, no hand-wave.
- Read the 4 new corpus cases in `tooling/tree-sitter-plum/test/corpus/function.txt` (lines 234-349 in the working tree): all four produce single-node trees (`closure`, `closure`, `fn_value_type`, `fn_value_type`) with no `ERROR`/`MISSING`. Field names (`params:`/`returns:`) are not printed in the corpus text — verified this is the file's universal convention, not an omission: grepped the whole file, zero existing entries anywhere show field-name annotations (e.g., `fn`'s own `name`/`params`/`returns` fields aren't shown either), so the new cases are consistent with existing style. The report's `--show-fields` verification is not directly checkable from the diff, but is unnecessary to trust blindly — grammar-level reasoning above already establishes field correctness independently.
- Confirmed the odd-looking `(type (a))` / `(type (b))` nodes in the "generic types and return" case are pre-existing project convention (`generic: ($) =&gt; choice($.a, $.b, $.c, $.d)` at grammar.js:514, inlined per `inline: [$.generic_type, $.generic]`), already exercised by an existing corpus case (`add(param: a, param2: List(b))`) — not a defect introduced here.
- Global constraints: closures only reachable via ordinary assignment + call syntax (`cb = |v| ...; cb(5)`), no trailing-closure syntax added — satisfied. `fn_value_type` params are positional types only, no param names in the rule — satisfied.

### Strengths
- Diff is minimal: exactly the two grammar edits + 4 corpus cases + regenerated binary artifacts (grammar.json/node-types.json/parser.c), 127 insertions / 2 deletions total — no scope creep.
- Field-based `params`/`returns` disambiguation correctly reuses an established idiom already present in the same grammar (`fn`'s own params/returns fields), giving Task 2 exactly the unambiguous access pattern the brief called for.
- Corpus cases cover both the base case (`fn(Int)`, no return) and the harder case (`fn(a) -&gt; b`, generic params + return), which is the pair that actually exercises the params/returns distinction.
- Report is transparent about a real deviation (brief's `-f` flag doesn't exist in the installed CLI; used `-i` instead) rather than silently papering over it.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- `fn_value_type` was placed immediately after `param` (before `return_type`) rather than "near `variadic_type`'s definition, a few lines above `param`" as the brief's parenthetical suggested. Purely stylistic/organizational — grammar.js:91 (`variadic_type`) and the new rule are now non-adjacent. No functional effect; not worth a follow-up.

### Assessment
**Task quality:** Approved
**Reasoning:** Both prescribed grammar changes match the brief exactly, the field-based `params`/`returns` wiring is verifiably correct by construction (reuses a proven in-codebase idiom), all 4 new corpus trees are clean single-node parses consistent with the file's existing conventions, and the diff carries no scope creep beyond the two rule changes plus regenerated artifacts.</result>
<usage><subagent_tokens>46988</subagent_tokens><tool_uses>11</tool_uses><duration_ms>128364</duration_ms></usage>
</task-notification>

**Claude:** Task 1 approved. Updating the ledger and moving to Task 2.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-c`
  - `Agent: Implement Task 2: closure AST and parser`

**Claude:** Task 2 implementer running in the background.

**You:** <task-notification>
<task-id>af31205b944fdfbc1</task-id>
<tool-use-id>toolu_016qrsbRBTQ5YEF9CHR4HoLS</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/af31205b944fdfbc1.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 2: closure AST and parser" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Committed successfully.

**Status:** DONE
Commits: `d7e5ff4` feat(plum-core): parse closure literals and fn(...) type annotations
Test summary: 4/4 new parser tests pass (RED confirmed before, GREEN after); full `cargo test --workspace` green with no regressions.
Concerns: none — 16 non-exhaustive-match compile errors were fixed with minimal/real fallback arms in `plum-checker` and `plum-wasm-codegen` (all TODO-tagged, inert, no test path reaches them).
Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md`</result>
<usage><subagent_tokens>97755</subagent_tokens><tool_uses>73</tool_uses><duration_ms>388795</duration_ms></usage>
</task-notification>

**Claude:** Let me check the report first — 16 fallback arms is a lot, want to make sure they don't preempt or conflict with Task 3's real implementation.
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/task-2-report.md`
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 906b92e`
  - `Agent: Review Task 2 (spec + quality)`

**Claude:** Task 2 reviewer running in the background, with specific scrutiny on whether the "real conversion instead of stub" choices could collide with Task 3's own planned work.

**You:** <task-notification>
<task-id>a77df8432b57d1e9c</task-id>
<tool-use-id>toolu_0154b9b1Lr5ddYZKooYfyki3</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a77df8432b57d1e9c.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 2 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>### Spec Compliance
- ✅ Spec compliant — `plum-core/src/ast.rs` and `plum-core/src/parser.rs` changes are byte-for-byte identical to the brief's prescribed code (`Closure` struct, `Expr::Closure`, `ParamType::Fn`, `parse_expression`'s `"closure"` arm, `parse_closure`, `parse_param`'s `fn_value_type` arm, `parse_fn_value_type`'s field-based `returns`/`params` disambiguation via `child_by_field_name("returns")` + node-identity filtering). `plum-core/tests/parser_test.rs` is the brief's 4 tests verbatim. RED/GREEN evidence in the report (4x `E0599` on `Expr::Closure`/`ParamType::Fn` not existing → all 4 tests passing) is consistent with what the diff would actually produce — traced `fn(Int) -&gt; Bool` by hand through `parse_fn_value_type`/`plum_type_from_ast` and it resolves to `ParamType::Fn([Type{Int}], Some(Type{Bool}))``PlumType::TFun([TInt], TBool)` correctly.

### Strengths
- All "real conversion" call sites (`plum-checker/src/lib.rs:build_global_tables`/`check_fn`, `monomorphize.rs:rewrite_fn_body` + both `PendingSpecialization` apply-paths, `plum-wasm-codegen/src/lib.rs:compile_fn_body`) use the *identical* expansion — `param_types.iter().map(plum_type_from_ast).collect()` plus `ret.as_ref().map(...).unwrap_or(PlumType::TUnit)` — so there is no drift in `ret: None` handling across sites (all six treat a missing return as `TUnit`, matching how ordinary `Fn.returns: None` is already handled elsewhere in this codebase, e.g. `plum-checker/src/lib.rs` line ~40).
- The `compile_expr` stub's `Err("codegen: closures are not yet supported".to_string())` matches this file's exact existing `"codegen: ..."` prefix convention (verified against ~15 other `Err`/`ok_or_else` sites in `plum-wasm-codegen/src/lib.rs`).
- `infer_expr`'s `Expr::Closure(_) =&gt; Ok(PlumType::TVar("_".to_string()))` reuses a convention already used 5+ times elsewhere in that same match (lines 136, 408, 507, 539, 565, 567 area) for "unmodeled, allow it" — not a novel pattern invented for this task.
- Confirmed none of the 16 fallback arms are reachable by any current test: `examples_test.rs` in both crates only scans `examples/*.plum` (not `libs/std/`), and `libs/std/list.plum` is the only place `fn(...)` syntax currently appears in the repo outside tests/examples — it's excluded from any exercised test path.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
- **Task 3 collision risk (point 3 of the scrutiny).** The actual plan (`docs/superpowers/plans/2026-07-20-closures.md`, Task 3, Step 4) instructs: *"Find every place in `plum-checker/src/lib.rs` that matches on `ast::ParamType::Type(t) =&gt; ... / ast::ParamType::Variadic(t) =&gt; ...` together (e.g. inside `build_global_tables`, `check_fn`) ... For each, add a third arm [`ParamType::Fn(...) =&gt; PlumType::TFun(...)`]."* Task 2 has already added exactly this arm, with the exact same conversion logic, at exactly those two call sites (`plum-checker/src/lib.rs` lines 31–35 and 57–61 in the diff). Task 3 Step 4 is therefore already done and must not be re-applied — when Task 3 is dispatched, its Step 4 will find nothing to do at those two sites (a no-op, not a conflict, since the code Task 2 wrote is functionally identical to what Task 3's brief prescribes), but the controller should strike or annotate Task 3 Step 4 before dispatch so the implementer doesn't waste effort re-deriving/duplicating it or get confused finding the "already done" state. Task 3 Step 3 (`infer_expr`'s real `Closure` arm) is unaffected — it will *replace* Task 2's placeholder `TVar("_")` arm rather than add a new one, which is a normal, expected edit, not a collision.
- This is a process/handoff flag, not a defect in Task 2's code — the conversions are correct and consistent — but it should be surfaced explicitly to whoever dispatches Task 3.

#### Minor (Nice to Have)
- The six duplicated `ParamType::Fn =&gt; PlumType::TFun(...)` conversion blocks (identical bodies) exist in `lib.rs` (x2) and `monomorphize.rs` (x4, one of which is copy-pasted twice more for the two `PendingSpecialization` arms) plus `plum-wasm-codegen/src/lib.rs` (x1) — a shared helper (e.g. `fn param_type_to_plum(pt: &amp;ast::ParamType) -&gt; PlumType`) would remove ~7x repetition of the same 4-line block and eliminate any future risk of the sites silently drifting apart. Not required for this task (brief only asked for compile-fixing fallbacks), but worth flagging since Task 3 will likely touch these same sites again.

### Assessment
**Task quality:** Approved
**Reasoning:** The brief's exact prescribed AST/parser code was reproduced faithfully and verified working (RED→GREEN evidence is credible and traceable); the "went beyond minimal stub" real `TFun` conversions are correct, mutually consistent (identical `ret: None → TUnit` handling everywhere), and reuse an existing type-conversion helper rather than inventing new logic, so they carry no correctness risk. The one thing worth acting on is informational, not a code defect: Task 3's brief (Step 4) already assumes `build_global_tables`/`check_fn` need a new `ParamType::Fn` arm, which Task 2 has already supplied — the controller should note this so Task 3 isn't dispatched expecting to still do that step.</result>
<usage><subagent_tokens>83465</subagent_tokens><tool_uses>16</tool_uses><duration_ms>109003</duration_ms></usage>
</task-notification>

**Claude:** Task 2 approved. One handoff note for Task 3: its Step 4 (adding a `ParamType::Fn` arm to `build_global_tables`/`check_fn`) is already done by Task 2 — I'll flag that in the dispatch so the implementer doesn't duplicate it.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-c`
  - `Agent: Implement Task 3: closure type checking`

**Claude:** Task 3 implementer running in the background.

**You:** <task-notification>
<task-id>a9ee878c9889d4eec</task-id>
<tool-use-id>toolu_0149LkdLaYEMU27Ks8do24YB</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a9ee878c9889d4eec.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 3: closure type checking" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

- Commit: `d1ea2ff` — feat(plum-checker): infer closure literal types; type-check fn(...)-typed params
- Test summary: all 3 new tests pass (RED confirmed first — `closure_literal_infers_as_a_function_type` failed with "'cb' is not a function" under the old placeholder); full `plum-checker` suite (42+1+9 tests) green; `cargo build --workspace` clean.
- Step 4 verified already-done: grepped `ast::ParamType::Fn` in `plum-checker/src/lib.rs`, confirmed both call sites (`build_global_tables`, `check_fn`) already had correct `TFun(...)` conversion from Task 2 — nothing duplicated or re-added.
- Concerns: none.
- Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-3-report.md` (note: this file previously held stale content from an unrelated earlier task; fully replaced with the correct Task 3 report).</result>
<usage><subagent_tokens>61001</subagent_tokens><tool_uses>30</tool_uses><duration_ms>169058</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 3 reviewer.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package d7e5ff4`
  - `Agent: Review Task 3 (spec + quality)`

**Claude:** Task 3 reviewer running in the background.

**You:** <task-notification>
<task-id>acefb2ec1851de9d8</task-id>
<tool-use-id>toolu_019r4aFDFERxbEgSiHG2KNW8</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/acefb2ec1851de9d8.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 3 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>These lines (72, 118, 173, 187, 482) are all absent from the diff hunks — confirming they exist in the repo (Task 2's work) but are truly unmodified context in this diff, matching the implementer's claim. The diff's only functional hunk is the `infer_expr` `Closure` arm at line 422-ish, and `FnCall` (line 482) is untouched, consistent with the design that closure calls flow through the existing unmodified `FnCall` path.

### Spec Compliance
- ✅ Spec compliant. The `infer_expr` `Expr::Closure` arm added (diff lines 30-43) is character-for-character identical to the brief's Step 3 prescribed code.
- ✅ Step 4 correctly left alone: `git diff` shows zero touches to `build_global_tables` (line 118) or `check_fn` (line 187) `ParamType::Fn` arms — confirmed these arms already exist pre-diff via grep, so the report's claim of "verification only, no duplication" checks out against the actual diff, not just the report's prose.
-`Expr::FnCall` (line 482) is untouched in the diff — closure calls type-check through the existing, unmodified path as designed; no unnecessary/duplicated call-site logic was added.
- ✅ Test additions (3 tests, diff lines 70-107) match the brief's Step 1 tests verbatim.

### Strengths
- Diff is minimal and precisely scoped: one functional hunk (`plum-checker/src/lib.rs` lines 422-443) plus pure test additions — no scope creep into Step 4 territory.
- Fresh-`TVar`-per-param approach is consistent with the file's existing permissive/gradual-typing style: `unify` (`plum-checker/src/lib.rs:19-36`) treats `(TVar(_), _)` as always `Ok(())` regardless of the variable's name, matching the same permissiveness already used for unmodeled type names (`ctx.enum_variants` fallback, diff line 23) and the prior placeholder itself.
- Report is candid about a real limitation (no cross-call-site strictness for closure param types) and correctly attributes it to the brief's specified design rather than hiding it.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- The `format!("_closure_{}", p)` naming for each param's `TVar` is dead weight — since `unify` ignores `TVar` identity entirely (line 21), the name has no semantic effect and could just as well be `TVar("_".to_string())` per param (as the old placeholder used). Not a defect, just a slightly misleading appearance of per-param distinctness that doesn't actually do anything. Harmless.

### Assessment
**Task quality:** Approved
**Reasoning:** The diff implements exactly the brief's Step 3 code, correctly leaves Step 4's pre-existing conversions untouched (verified against the diff itself, not just the report), and does not touch `FnCall`, matching the intended design where closure calls ride the existing inference path.</result>
<usage><subagent_tokens>40097</subagent_tokens><tool_uses>6</tool_uses><duration_ms>31093</duration_ms></usage>
</task-notification>

**Claude:** Task 3 approved. Updating the ledger and moving to Task 4.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-c`

**Claude:** Task 4 (WasmModule table/element infrastructure) is small and self-contained — dispatching on the standard tier.
  - `Agent: Implement Task 4: wasm table infra`

**Claude:** Task 4 implementer running in the background.

**You:** <task-notification>
<task-id>a957ccf5fcce417d0</task-id>
<tool-use-id>toolu_01T1xzWGVjiRmL6JYXkkrcmL</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a957ccf5fcce417d0.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 4: wasm table infra" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>**Status:** DONE

Commit: `db00bd9` — feat(plum-wasm-codegen): add wasm function table + element section support

Test summary: new `wasm_module_with_a_table_element_validates_and_call_indirect_works` passes (wasmparser validates, wasmtime executes `call_indirect` returning 42); full `cargo test --workspace` is green (no regressions).

Concerns: the brief's literal test snippet had two small bugs needed to actually compile/pass — a missing `use wasm_encoder::Encode;` import, and a missing leading `0u8` locals-decl-count byte in the hand-built function bodies (without it, wasm-tools reported "unknown type 11: type index out of bounds" — a garbage-locals decode cascade, not a Table/Element section-ordering problem). Both fixed in the test only; `WasmModule`'s implementation matches the brief exactly and the Table/Element section positions validated correctly on the first attempt once the test bodies were fixed.

Report: `/Users/pyrossh/Code/plum/.superpowers/sdd/task-4-report.md`</result>
<usage><subagent_tokens>98570</subagent_tokens><tool_uses>35</tool_uses><duration_ms>277593</duration_ms></usage>
</task-notification>

**Claude:** Now dispatching the Task 4 reviewer.
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package d1ea2ff`
  - `Agent: Review Task 4 (spec + quality)`

**Claude:** Task 4 reviewer running in the background.

**You:** <task-notification>
<task-id>a9df005c07e46c6ea</task-id>
<tool-use-id>toolu_011QGYiwDGjiNQRE9wYfcF7Y</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a9df005c07e46c6ea.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 4 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Wildcard import `use wasm_encoder::*;` — all the new types (TableSection, TableType, RefType, ElementSection, ConstExpr, Elements) are already in scope. Confirmed section order matches the diff context exactly: Table section block sits between Function section block and Memory section block; Element section block sits between Export section block and Code section block.

### Spec Compliance
- ✅ Spec compliant. `table_elements: Vec&lt;u32&gt;` field added exactly as specified (`plum-wasm-codegen/src/lib.rs`), initialized in `new()`, `add_table_element` method matches brief verbatim. Table section block inserted immediately after the `// Function section` block (`module.section(&amp;funcs); }`) and before `// Memory section`. Element section block inserted immediately after `// Export section` block (`module.section(&amp;exports); }`) and before `// Code section`. Both gated on `!self.table_elements.is_empty()` exactly as prescribed. `TableType`/`ElementSection`/`ConstExpr`/`Elements` all resolve via the pre-existing `use wasm_encoder::*;` glob import — no missing import in the implementation.

### Strengths
- Section-insertion diff hunks confirm correct binary section ordering by inspection of surrounding context lines, not just by trusting validation to catch it.
- Test exercises the actual `call_indirect` runtime path (wasmtime execution), not just structural validation, and uses a non-constant table-index operand as the brief intended (rules out coincidental validation).
- `add_table_element` and section-emission code are unchanged from the brief's prescribed snippets — no unrequested scope creep.

### Issues
#### Critical (Must Fix)
None.

#### Important (Should Fix)
None.

#### Minor (Nice to Have)
- The two test fixes (`use wasm_encoder::Encode;` import; `vec![0u8]` leading locals-count byte in both hand-built function bodies) are confirmed confined to `plum-wasm-codegen/tests/codegen_tests.rs` only — no corresponding change touches `WasmModule::add_function`, `finish()`, or the Code-section emission (`code.raw(body_bytes)` in `src/lib.rs` is unchanged and still treats the passed byte slice as raw, self-contained function body content, consistent with `compile_fn_body` elsewhere always prepending its own locals-count byte). The report's claim of "test-only" bugs is accurate — verified directly against the diff.

### Assessment
**Task quality:** Approved
**Reasoning:** The diff matches the brief's prescribed field/method/section-emission code exactly, section ordering (Table between Function/Memory, Element between Export/Code) is correct by direct inspection of the diff context, and the two self-reported test fixes are genuinely confined to test-construction code with no bleed into `WasmModule`'s implementation.</result>
<usage><subagent_tokens>39846</subagent_tokens><tool_uses>4</tool_uses><duration_ms>25744</duration_ms></usage>
</task-notification>

**Claude:** Task 4 approved. Updating the ledger and moving to Task 5 — the biggest task in this plan.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-c`

**Claude:** This is the big one — dispatching on Opus, with the same "design is a starting point, not verified code" framing that worked well for the generics monomorphization plan's hardest task.
  - `Agent: Implement Task 5: closure codegen`

**Claude:** Task 5 (the substantial closure-compilation codegen) implementer running in the background on Opus — this may take a while given the complexity, and could plausibly come back BLOCKED, which would be a reasonable outcome given the task's own framing.

**You:** <task-notification>
<task-id>ac7354b39654904da</task-id>
<tool-use-id>toolu_01WDyKJKPvbjUMLEgGVsq93j</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/ac7354b39654904da.output</output-file>
<status>completed</status>
<summary>Agent "Implement Task 5: closure codegen" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Committed as `4ba0db3`. Final report below.

---

**Status:** DONE_WITH_CONCERNS

**Commit:** `4ba0db3` feat(plum-wasm-codegen): compile closure literals and closure calls via a function table

**Tests:** All 3 new tests pass (incl. the capture-semantics one that matters most); full codegen suite 53/53, full `cargo test --workspace` green, tree-sitter 27/27 corpus parses.

**Divergence from the brief's design:** Minor and all justified in the report. Codegen followed the 6-part design closely — the substantive deltas: (1) `TFun` maps to `i32` (a closure is a pointer), (2) closure discovery records *every* closure literal (not just direct call-args) because test 2's closure lives in an assignment, and (3) `CompileCtx` gained a lifetime to hold closure ASTs. The codegen architecture itself held up fully.

**Concerns:**
- The brief's exact test 1/3 syntax `each(|v|\n    v)` (indented body, trailing `)` on the same line) **does not parse** with the merged grammar — a grammar/scanner gap in Task 1's domain, not codegen. I applied the minimal, precedented grammar fix (closure body `choice($.expression, $.body)`, mirroring `fn`/`case`) plus the matching `plum-core` parse change, enabling the plan's canonical `each(|v| v)` form, and set tests 1/3 to that single-line form. Assertions and verified semantics are unchanged; this is not a weakening. Test 2 passes exactly as written. This meant touching files beyond the brief's list (grammar.js + regenerated parser, plum-core/parser.rs, corpus).
- The trailing multi-line closure-arg form still needs an external-scanner change (emit DEDENT at a closing bracket) — flagged as follow-up, deliberately not attempted (risky, out of codegen scope).
- Nested closures (closure inside a closure body) are unsupported; they produce a clear compile error, not silent miscompilation.

**Report:** `/Users/pyrossh/Code/plum/.superpowers/sdd/task-5-report.md`</result>
<usage><subagent_tokens>178251</subagent_tokens><tool_uses>89</tool_uses><duration_ms>1646931</duration_ms></usage>
</task-notification>

**Claude:** Let me read the full report before dispatching review — this touched grammar territory beyond Task 5's stated scope, want to understand exactly what changed.
  - `Read: /Users/pyrossh/Code/plum/.superpowers/sdd/task-5-report.md`
  - `Bash: git log --oneline -3 && /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package db00bd9`
  - `Agent: Review Task 5 (spec + quality)`

**Claude:** Task 5 reviewer running in the background on Opus with heavy scrutiny on six specific risk points.

**You:** <task-notification>
<task-id>a631b6d1f397c4ecc</task-id>
<tool-use-id>toolu_01M6V9ZzqV7nf3DCNB8bJvAc</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a631b6d1f397c4ecc.output</output-file>
<status>completed</status>
<summary>Agent "Review Task 5 (spec + quality)" finished</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>### Spec Compliance
- ✅ Spec compliant. All 6 design parts implemented; the 6 flagged deviations are each technically justified. One latent soundness gap (point 2) is confined to a non-delivered case (non-`I64` closure signatures via the assignment path) and does not affect the delivered feature/tests.
- ⚠️ Cannot verify from diff: exact behavior of `plum_checker::infer_expr` on a bare closure (verified its consequences by reasoning about `plum_type_to_valtype` mappings, not by reading the checker).

### Verification of the six scrutiny points

1. **`plum_type_to_valtype(TFun) I64 → I32` — SOUND, and in fact necessary for consistency.** The fn-typed *param slot* was already `I32` before this task: `param_type_name(ParamType::Fn) = "Fn"` (lib.rs:310) → `ast_type_to_wasm("Fn") = I32` via the default arm (lib.rs:321). So `each(cb: fn(Int)-&gt;Int)`'s `cb` slot has always been declared `I32`. The old `plum_type_to_valtype(TFun)=I64` was the inconsistent path, and it was genuinely dead: the checker only produces `TFun` for closures (Tasks 3+5), so before this task nothing ever typed a local/param/field as `TFun` and reached `plum_type_to_valtype` with it. The "50 pre-existing tests unaffected" claim is credible — no prior test constructed or consumed a `TFun`-typed runtime value; it was a pure dead-path flip. All 9 `TFun` occurrences in the file (lib.rs:332/550/691/1088/1424/1820/2128) are either the mapping itself or new closure code.

2. **Generalized discovery (every literal, concrete-param-type OR `infer_expr` fallback) — works for delivered tests; contains a latent non-`I64` divergence.** The registration signature for an assignment-path closure comes from `infer_expr` on the closure *alone*, without the expected-type context (lib.rs:508). An under-constrained param infers `TVar``plum_type_to_valtype(TVar)=I64` (lib.rs:333). The call site (`compile_closure_call`, lib.rs:2122-2135) builds its `call_indirect` type from `infer_local_type` on the callee local, which inside `each` resolves to the *concrete* declared `fn(Int)-&gt;Int`. For `Int`, both `TVar` and `TInt` map to `I64`, so the registration key `[I32,I64]→I64` equals the call-site key and they share one type index — correct. The gap: a polymorphic closure (`|v| v`) assigned to a local and later used where the concrete type is `Float` (`F64`) or a pointer (`I32`) would be *registered* with an `I64` param but *called* through a `Float`/pointer type index → `call_indirect` runtime trap. Same applies to return type, and to the innermost hard fallback (lib.rs:518-520) which forces all-`I64`. Direct-argument closures (`each(|v| v)`) are always registered from the concrete `ParamType::Fn`, so they are never affected. This is confined to non-`Int` closures reaching a call site via the assignment path — outside the three delivered tests but reachable within the intended feature.

3. **`closure_call_types` dedup — SOUND, one type index per distinct wasm signature, and registration/call indices provably coincide.** The map is populated once from every fn-typed param (lib.rs:241-251) and reused-or-inserted per discovered closure (lib.rs:294-303); the closure *function* is registered with exactly that looked-up `type_idx` (lib.rs:304). The call site looks the same key up (lib.rs:2129-2134). Because it is a `HashMap` keyed by `(Vec&lt;ValType&gt;, Option&lt;ValType&gt;)`, equal signatures always yield the identical index — a mismatched register-vs-call index is impossible *when the keys are equal*; the only way to get divergence is the point-2 key divergence, not a dedup fault. A never-registered signature errors clearly ("internal codegen error: no call_indirect type", lib.rs:2135) rather than miscompiling.

4. **Grammar change — minimal and non-regressing.** `field("body", choice($.expression, $.body))` (grammar.js) exactly mirrors `fn`/`case`; the block form still routes through `parse_block` unchanged, the inline form normalizes into `Block { stmts: vec![Stmt::Expr(...)] }` (parser.rs:47-56) with no semantic change (a single expr statement is the block's value). The corpus test covers the inline-arg form. Deferring the multi-line-trailing-`)` scanner fix is the right scope call: no delivered artifact is left half-working — all three tests use forms that parse (tests 1/3 single-line, test 2 assignment-with-indented-block). The only cost is that brief tests 1/3 were rewritten to single-line; assertions (`==5`, `==14`) and the semantics are preserved.

5. **`CompileCtx&lt;'a&gt;` lifetime — SOUND, no dangling risk.** `let source = &amp;monomorphize_source(source)?;` (lib.rs:388) binds the owned monomorphized `Source` behind a reference in a `let`, so temporary-lifetime-extension keeps it alive for the whole `compile_source` body. `fns`, `raw_closures`, `closure_asts`, and `ctx` all borrow into it and are dropped before it. The borrowed source outlives every use of `CompileCtx`; this isn't the borrow checker merely tolerating a fragile structure — the owner genuinely lives longest.

6. **Nested closures — genuine named error path.** Discovery never recurses into closure bodies (`fv_collect_refs_expr` closure arm is `_ =&gt; {}`, lib.rs:702; `record_closure` doesn't walk the body), so a nested literal is absent from `ctx.closures`. Compiling it hits `compile_closure_literal`'s `.ok_or_else(|| "codegen: closure literal was not discovered (nested closures are not supported)")` (lib.rs:2036-2039) — a clear, specific error before any bytes are emitted, not an incidental failure or silent miscompile.

### Strengths
- Point 1's change is not just harmless but corrects a pre-existing latent inconsistency between `ast_type_to_wasm("Fn")` and `plum_type_to_valtype(TFun)`.
- Snapshot-by-value capture is structurally correct: env struct is filled from enclosing locals at construction (lib.rs:989-997), reloaded by index at closure entry (lib.rs:1158-1163), with consistent first-appearance ordering on both sides — test 2's `x==10`-not-`999` is a real proof, not a coincidence.
- Free-variable analysis correctly excludes params and locally-bound names (assignments, `for`/`match`/pattern bindings) via a two-pass bound/refs walk (lib.rs:549-704).
- `expr_has_result` correctly special-cases closure-typed locals so a `Unit`-returning closure call isn't wrongly dropped (lib.rs:1421-1428).
- Error paths throughout are specific and fail before emitting bytes (missing scratch slot, uncaptured variable, non-callable local).

### Issues
#### Critical (Must Fix)
- None.

#### Important (Should Fix)
- **Latent `call_indirect` type mismatch for non-`I64` closure signatures reached via the assignment path** (lib.rs:508-521, 2122-2135). A polymorphic closure assigned to a local, whose param/return is inferred as `TVar` (→`I64`) but which is used at a concrete `Float` (`F64`) or pointer (`I32`) type, registers its function under an `I64` signature yet is called through the concrete type index → runtime trap. Also the hard fallback (lib.rs:518-520) forces all-`I64` regardless of the real types. Safe today only because every delivered test is `Int`-only. Fix direction: resolve the closure's signature from the same expected-type/unification information the checker used at the call, rather than from a context-free `infer_expr` on the closure alone — or at minimum reject (clear error) an assignment-path closure whose inferred signature is not fully concrete, instead of silently registering an `I64` shape.

#### Minor (Nice to Have)
- The brief's multi-line-trailing-`)` closure form remains unparseable (correctly deferred as a scanner follow-up), but this means the delivered corpus/tests exercise only the single-line and indented-block forms; worth an explicit tracking item so the deferred scanner work isn't lost.
- `param_plum_type` logic is duplicated (lib.rs:543-551 and the inline construction at ~1088); could be unified.

### Assessment
**Task quality:** Approved
**Reasoning:** All six deviations hold up under scrutiny — five are fully sound, and the sixth (generalized discovery) is correct for every delivered test with a clearly-bounded latent gap that only affects non-`I64` closures on the assignment path, which is outside this task's delivered scope; that gap should be tracked/fixed before closures over `Float` or reference types are relied upon.</result>
<usage><subagent_tokens>73492</subagent_tokens><tool_uses>8</tool_uses><duration_ms>209613</duration_ms></usage>
</task-notification>

**Claude:** Task 5 approved, with one Important finding: a latent signature-mismatch risk for a polymorphic closure assigned to a local then used at a non-`Int` concrete type via the assignment-discovery path (would trap at runtime, not silently miscompute — wasm's own `call_indirect` type-check catches it, just late). Scoped outside all three delivered tests. I'll track this as a follow-up rather than block on a deeper bidirectional-inference fix, and make sure it's documented in Task 6.
  - `mcp__ccd_session__spawn_task: Fix closure signature resolution for non-Int types via assignment path`

**Claude:** Updating the ledger and moving to Task 6, with an added instruction to document this limitation.
  - `Edit: /Users/pyrossh/Code/plum/.superpowers/sdd/progress.md`
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/task-brief docs/superpowers/plans/2026-07-20-c`
  - `Agent: Implement Task 6: closures examples and docs`

**Claude:** Task 6 (final task) implementer running in the background.

**You:** <task-notification>
<task-id>a9df750f5018cf4fa</task-id>
<tool-use-id>toolu_014tHbwQu4ZERid6GnSKJ46b</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a9df750f5018cf4fa.output</output-file>
<status>failed</status>
<summary>Agent "Implement Task 6: closures examples and docs" failed: Agent stalled: no progress for 600s (stream watchdog did not recover)</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Now let's add the closures test to plum-wasm-codegen's examples_test.rs.</result>
</task-notification>

**Claude:** The Task 6 agent stalled. Let me check the current repo state before deciding whether to resume or restart it.
  - `Bash: git status --short; echo ---; git diff --stat`
  - `Bash: cat examples/closures.plum; echo ---DIFF---; git diff plum-wasm-codegen/tests/examples_test.rs`
  - `Bash: timeout 30 cargo test -p plum-wasm-codegen --test examples_test closures_example -- --nocapture 2>&1 | tail -30`
  - `Bash: timeout 60 cargo test -p plum-wasm-codegen --test examples_test closures_example -- --nocapture 2>&1 | tail -40`
  - `Bash: cargo test -p plum-wasm-codegen --test examples_test closures_example -- --nocapture 2>&1 | tail -40`

**Claude:** Confirmed — no actual hang, just an agent-harness hiccup. The real issue: `useCapturingClosure`'s multi-line `each(|v|\n v + offset)` form doesn't parse (exactly the known trailing-paren limitation flagged for this task). Resuming the same agent with that diagnosis so it can finish.
  - `SendMessage: a9df750f5018cf4fa`
  - `SendMessage: a9df750f5018cf4fa`

**Claude:** Resumed and running in the background.

**You:** <task-notification>
<task-id>a9df750f5018cf4fa</task-id>
<tool-use-id>toolu_01WrkiwxkX3ZpXSKTYvtSGzi</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a9df750f5018cf4fa.output</output-file>
<status>failed</status>
<summary>Agent "Implement Task 6: closures examples and docs" failed: Agent terminated early due to an API error: API Error: Connection closed mid-response. The response above may be incomplete.</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Both green, and now closures_example actually verifies both non-capturing and capturing (assignment-based) closure execution end to end, returning 120 as expected. Now let's update the README.</result>
</task-notification>

**Claude:** Another transient connection issue, but real progress landed (example now passes, returns 120). Checking state and resuming again.
  - `Bash: git status --short; echo ---; cat examples/closures.plum`
  - `Bash: grep -n "closures_example" -A15 plum-wasm-codegen/tests/examples_test.rs`

**Claude:** Good, already consistent. Resuming the agent to finish the remaining steps (README + full suite + commit).
  - `SendMessage: a9df750f5018cf4fa`

**Claude:** Resumed again, running in the background.

**You:** <task-notification>
<task-id>a9df750f5018cf4fa</task-id>
<tool-use-id>toolu_01RtCDtmy2KJFBKFLN6ACLyW</tool-use-id>
<output-file>/private/tmp/claude-501/-Users-pyrossh-Code-plum/61ed18b5-cab0-453c-9a65-dcb8cfc58361/tasks/a9df750f5018cf4fa.output</output-file>
<status>failed</status>
<summary>Agent "Implement Task 6: closures examples and docs" failed: Agent terminated early due to an API error: API Error: Connection closed mid-response. The response above may be incomplete.</summary>
<note>A task-notification fires each time this agent stops with no live background children of its own. The user can send it another message and resume it, so the same task-id may notify more than once.</note>
<result>Now write the report file.</result>
</task-notification>
  - `Bash: git status --short; echo ---; git diff --stat; echo ---README---; grep -n "Closures\|closure" README.md | head -20`
  - `Bash: git log --oneline -3; echo ---; cargo test --workspace 2>&1 | tail -10`
  - `Bash: cargo test --workspace 2>&1 | grep -E "test result|FAILED"`
  - `Bash: cd tooling/tree-sitter-plum && npx --yes tree-sitter-cli test 2>&1 | tail -3`

**Claude:** Everything is complete and committed despite the connection hiccups — full workspace green (all crates, 0 failures), tree-sitter corpus 27/27, everything committed as `6341c74`. Dispatching the Task 6 reviewer.
  - `Bash: /Users/pyrossh/.claude/plugins/cache/superpowers-dev/superpowers/6.1.1/skills/subagent-driven-development/scripts/review-package 4ba0db3 6341c74`
  - `Agent: Review Task 6 (spec + quality)`

**Claude:** Task 6 reviewer running in the background.