The University of Michigan's latest AI programming course, "Applied Agentic Software Engineering," packages its methodology as five Skills
/peanuts (legacy code only) → /elephant → /goldfish → /egm-implement → /mean-review
Skills address: https://github.com/eecs498-aase/materials/tree/main/egm/skills
1. /peanuts — Builds a README hierarchy for legacy code, bottom-up
Purpose: A bootstrap for legacy codebases (with no design docs). A million lines of code is "the whole jungle" to an AI, which gets lost and hallucinates. The fix is one README per directory: leaf directories are "peanuts," branch directories are "hay," and summaries are compressed recursively from the bottom up.
Core mechanisms:
· Strictly bottom-up: a parent directory's README can only be generated after all its direct child directories are approved, because bad leaves compound and amplify as they move up the tree.
· The human review gate is a hard rule: the AI-generated README can only be marked needs-human-review and can never mark itself approved. The skill states it outright: "The AI is not the judge."
· Honest error-rate expectations: it explicitly warns that "about 50% of leaf descriptions will be wrong," because the leaf layer only has code to look at, and it budgets 5–10 minutes of human correction per leaf. The branch layer is more accurate because child-directory summaries anchor the model.
· PEANUTS.md ledger: five states (pending / needs-human-review / approved / rolled-up / blocked) drive resumable runs, backed by a Python script for deterministic ordering and gate checks.
· Retirement clause: every README ends with a fixed line, "This file can be retired once design documentation covers it," which positions it as scaffolding rather than permanent documentation.
2. /elephant — Design long before code
Purpose: The soul of EGM. The output is not code but a four-section design document (Problem / Technical Plan / Alternatives / Detailed Implementation), which becomes the single source of truth.
Core mechanisms:
· Absolutely no code: the entire session never writes a line of code (not even "illustrative examples"), except for a few pseudocode snippets that explain mechanisms inside the document.
· The "No Code" interview: the opening is a fixed, verbatim second-person instruction ("I don't want you to write code... you should challenge my assumptions"), followed by a 20–30 minute design interview of follow-up questions, asking one or two questions at a time, challenging assumptions ("Why must this be real-time?"), and pushing toward edges (failure modes? rollback? what does it break?).
· AI drafts first: a hard rule that the user's first draft must not anchor the discussion. The AI must first produce an independent prose-plus-diagram proposal, because the user's blind spots are exactly what this skill is meant to expose.
· Anti-sycophancy safeguard: once it notices itself agreeing repeatedly, it must read a reset line verbatim: "You are not helping. Your greatest value is challenging my thinking. When you agree with me, you are not helping." Then it re-enters as a critic.
· Build section by section, never all at once: the four sections are written in order, and each is shown to the user and refined until satisfactory before the next is written. A large document produced in one pass is "shallow and internally inconsistent."
· The pacing expectation is counterintuitive: during the technical debate phase, "2–3 days of back-and-forth is normal, don't rush."
3. /goldfish — Verifying the design document with "fresh eyes"
Purpose: Elephant holds a lot of context in its head, but that context lives in the chat, not on the page. Goldfish is a new reviewer with zero shared context who knows only what is written on the page.
Pass criteria: a Goldfish reads the document and can (1) explain the system back, (2) identify gaps, and (3) confirm it could implement without further questions. Only then does the document carry the design; otherwise it is "a thin layer on top of context you'll lose."
Core mechanisms:
· Three reviewers spawned in parallel in a single message: A, the comprehension test (explain the system back); B, the critic (find omissions and wrong assumptions); C, implementation readiness (READY/NOT-READY plus a list of every question they would be forced to ask). Parallelism is not only faster, it also prevents collusion, such as feeding A's conclusions into B.
· The technical implementation of freshness is deliberately clear: the value comes from independent context, not from a persona. The orchestrator reads ELEPHANT.md itself to orient, but never passes any design conversation content to the reviewers; it even guards against document links leaking into the ledger directory.
· The 30% heuristic: about a third of the critic's suggestions are truly valuable, so the orchestrator must synthesize and rank them (real gaps first, wording nits at the bottom) and distinguish "the document is wrong" from "the document is right but the reviewer disagrees with the design choice," rather than accepting everything.
· Loop until a stopping condition: critiques drop to nit level and human_review_gate is marked passed by a human (or explicitly skipped-solo). This gate is never rewritten automatically; a person or teammate must personally write it into the header of GOLDFISH.md.
· Reports are never overwritten: each round is saved as a separate file (with an -r2 suffix), and GOLDFISH.md is appended to round by round.
4. /egm-implement — Pinning the approved document to the code
Purpose: The most engineering-disciplined stage in EGM. The gap between "document approved" and "code merged" is where AI projects quietly drift; a new session treats the document as "reference," re-derives context from chat history, and drifts.
Core mechanisms:
· Entry gate: run check_gate.py first; exit code 0 (GO) / 1 (BLOCKED) / 3 (CHECK-BY-HAND, where ambiguous ledger values are handed to a human to decide). Both readiness: ready and the human gate passed are required; neither can be skipped silently.
· File-level fence: files not enumerated in the design document are never changed. Want to change one? Stop and propose a document update first. "Also rename X while we're here" is drift and must be surfaced.
· Surface drift, don't absorb it: this is the most elegant rule in the whole skill. When reality diverges from the document (a wrong assumption, a needed new file, code that doesn't compile as written), it is not silently changed. Instead, a drift entry is recorded in IMPLEMENT.md (date/type/suggested disposition/open status), graded by impact: small clarifications are fixed in the document in place, substantive changes rewrite that section, and major changes go back to /elephant or /goldfish. Unresolved drift blocks further work. Otherwise the design document loses its status as source of truth one undocumented diff at a time.
· IMPLEMENT.md is a crash-recovery protocol: every file, every decision, and every drift is recorded in real time (not reconstructed afterward). A crashed session costs nothing; a new session reads the ledger and resumes from the breakpoint. Six months later, when someone asks "why does this code look like this?", the answer is a three-part bundle: design document + verification record + implementation trail. That is the organizational memory EGM compounds.
· Small-diff discipline: one logical change per file, because all output flows into mean-review.
5. /mean-review — Adversarial review of code
Purpose: A mirror of goldfish: goldfish reviews documents, mean-review reviews code. The premise is sharp: AI produces code faster than humans can carefully review it, and an AI reviewer that politely says "looks good!" is exactly how slop accumulates at scale.
Core mechanisms:
· A load-bearing framing prompt: the opening must read verbatim the user's line, "I have a strong intuition this code quality is poor, please tear it apart and tell me where it's rotten." This framing is the authorization to be sharp; without it the model drifts back into polite mode.
· Four mandatory scans (scripted): rules against 10 uncommented lines, functions over 50 lines, weak naming (a blacklist of data/tmp/result/helper/utils, etc.), and silently swallowed exceptions. Deterministic checks are handed to a Python script, with results tagged [enforced] to distinguish them from judgment-based findings. These four categories are exactly what human reviewers most easily skim past.
· A clever exemption: declarative code (constant tables, schemas, field lists, route tables) is exempt from the 10-line comment rule, because "the shape of the thing is itself the explanation." Wrongly flagging declarative code counts as a reviewer error.
· Reads whole files, not just diff hunks, since half of correctness bugs hide in the code surrounding the change.
· Output is a prioritized punch list (correctness > maintainability > style), each item tagged [correctness]/[readability]/[nit], with file:line, one sentence stating the problem, and one sentence stating the fix. Softening phrasing like "maybe consider" is prohibited. Everything is appended to MEAN-REVIEW.md, each round records dispositions (what was fixed, what was rejected and why), and the loop continues until only nits remain.
· Tone calibration: sharp but not abusive. "This is broken, here's why" is right; "What were you thinking?" is wrong.
· It only reviews and never changes code: the user fixes, then reruns the next round.
