LINUXOR.SK ... open source notes ...

Matt Pocock Skills - Quick & Dirty Starter Guide

category: howtoz · date: 2026-10-04 · author: LALA
noteThis guide describes the current main-branch architecture often discussed as the v1.3-era of Matt Pocock Skills. As of 2026-10-03, GitHub Releases still lists v1.2.3 as the latest formal release, so treat "1.3" here as the newer workflow architecture on main, not a published v1.3.0 tag.

Matt Pocock Skills are easiest to understand as a lightweight engineering operating system for coding agents, not as a bag of prompts. You choose the workflow; the agent can reach for reusable engineering disciplines such as TDD, debugging, domain modelling, research, architecture work and code review.

Matt Pocock Skills operating model: you choose a workflow, the agent applies engineering disciplines, changes produce feedback, and project knowledge survives across sessions.
Matt Pocock Skills operating model: you choose a workflow, the agent applies engineering disciplines, changes produce feedback, and project knowledge survives across sessions.

The architecture is intentionally composable. A workflow skill coordinates the job, engineering primitives provide reusable disciplines, repository-local documents carry context between sessions, and tests or reviews provide evidence back to the agent.

Quick start

For Claude Code, install the plugin:

bash
$ claude plugins install mattpocock-skills

For Codex and other compatible agents, use the skills installer:

bash
$ npx skills@latest add mattpocock/skills

Make sure setup-matt-pocock-skills is installed, then run it once in each repository:

output 1 line
/setup-matt-pocock-skills

Setup answers three repository-local questions: where issues live, which triage vocabulary the repository uses, and where domain documentation lives. It stores those answers under docs/agents/, so the same skills can work against GitHub, GitLab, local Markdown or another issue tracker without editing the skill files themselves.

That is a useful architectural pattern by itself: the skills stay stable; the environment becomes configuration.

The only four workflows you need at first

Four compact Matt Pocock Skills workflows for a small change, a normal feature, a multi-ticket feature and a hard bug.
Four compact Matt Pocock Skills workflows for a small change, a normal feature, a multi-ticket feature and a hard bug.

You do not need to run every skill for every change. The useful rule is simple: increase ceremony as uncertainty and scope increase.

For a small and obvious change, start with implement.

For a normal feature with unresolved assumptions, use grill-with-docs, then to-spec only if the decisions need to survive into later sessions, then implement.

For a larger feature that needs several independently useful slices, add to-tickets. If those tickets form a dependency graph large enough to benefit from multiple implementers, use implement-spec.

For a hard bug or performance regression, start with diagnosing-bugs, not with a request to "just fix it".

How much process does the work deserve?

SituationSmallest useful flow
Decision is already clear and the work fits one context windowimplement
Important assumptions are still unresolvedgrill-with-docs -> implement
Decisions must survive across sessionsgrill-with-docs -> to-spec -> implement
Several agent-sized vertical slices are neededto-spec -> to-tickets -> implement
Tickets expose independent ready workto-tickets -> implement-spec
Root cause is unknowndiagnosing-bugs
The whole initiative is larger than one planning horizonwayfinder

A useful threshold from the current to-spec documentation is the context-window boundary. If the work is already decided and comfortably fits one session, a spec can be unnecessary ceremony. If the work must survive several fresh sessions, the spec becomes the durable handoff.

Which skill should I use?

SituationStart with
I do not know exactly what we should buildgrill-with-docs
We already decided; capture it as a durable specto-spec
The work needs several tracer-bullet slicesto-tickets
This is one bounded implementation taskimplement
This is a dependency graph of ticketsimplement-spec
I do not know which workflow fitsask-matt
The bug is confusing or intermittentdiagnosing-bugs
Several technical approaches need experimental validationprototype
I need primary-source technical researchresearch
The architecture is becoming hard to changeimprove-codebase-architecture
The initiative is too large for one context windowwayfinder
Implementation is completecode-review
The reviewed work needs a useful PR descriptionpr
We should improve the agent environment after this workretro

User-invoked workflows vs engineering primitives

This separation is one of the most useful design ideas in the repository.

User-invoked skills are orchestration. You normally choose them explicitly: ask-matt, grill-with-docs, triage, improve-codebase-architecture, setup-matt-pocock-skills, to-spec, to-tickets, implement, implement-spec, wayfinder, and retro.

Model-invoked skills are reusable engineering disciplines. The agent can reach for them when the task fits: prototype, diagnosing-bugs, research, tdd, domain-modeling, codebase-design, code-review, pr, and wizard.

The mental model is: you choose the engineering workflow; the agent chooses appropriate engineering techniques inside it.

What good output looks like at each stage

This is a practical way to detect whether a workflow is working or merely producing more text.

StageGood output
grill-with-docsAmbiguities are resolved, vocabulary is sharpened, important decisions are captured in the glossary or ADRs
to-specA decision record that says what must be true, the chosen seams, testing decisions and explicit out-of-scope items
to-ticketsSmall vertical tracer bullets with acceptance criteria and blocking relationships
implementOne decided piece of work implemented with a tight TDD loop, checked and committed
implement-specThe whole ticket graph integrated on one branch, with ready tickets parallelized only where dependencies permit
code-reviewSeparate Standards and Spec findings against a known fixed point
retroConcrete improvements to navigation, checks, standards, steering files or tooling for the next session

A good skill run should leave behind better state, not just a longer transcript.

First real feature

Imagine the request is: add configurable per-customer API rate limits.

Do not begin with "implement customer rate limiting" unless the semantics are already obvious. Start with:

output 3 lines
/grill-with-docs

We need configurable per-customer API rate limits.

The useful questions are architectural, not cosmetic: what exactly is a customer, whether limits are per account or API key, how bursts behave, whether counters are distributed, what happens during datastore failure, and which terms belong in the project glossary.

grill-with-docs is deliberately small as an orchestrator: it combines the reusable grilling and domain-modeling disciplines so the interview sharpens both the plan and the project's language.

When the decisions are made and the work needs to survive across sessions, capture them:

output 1 line
/to-spec

to-spec is a decision record, not another interview. It synthesizes what has already been decided, uses the repository's domain vocabulary, records testing seams and avoids brittle implementation details that can become stale quickly.

If the feature needs several vertical slices, decompose it:

output 1 line
/to-tickets

Prefer tracer-bullet tickets that prove behavior end-to-end. Avoid tickets such as "database", "backend", "API" and "tests" that merely split the codebase into layers.

For one bounded ticket:

output 1 line
/implement

For the whole dependency graph:

output 1 line
/implement-spec

Why tracer-bullet tickets matter

to-tickets does not treat tickets as a conventional project-management checklist. Each ticket should be a narrow but complete vertical slice that can be verified on its own and sized for a fresh agent context.

Bad decomposition usually looks like this:

output 4 lines
Ticket 1: database
Ticket 2: backend
Ticket 3: API
Ticket 4: tests

Nothing is really proven until the final layer lands, so each ticket depends semantically on unfinished work elsewhere.

A better decomposition looks like this:

output 3 lines
Ticket A: enforce one default customer limit end-to-end
Ticket B: add per-customer override behavior end-to-end
Ticket C: expose limit status and rejection metadata end-to-end

Each slice can own its tests, implementation and acceptance criteria.

Parallelism should come from the task graph

implement-spec treats tickets as a dependency graph with blocking edges rather than as a sequential list. The currently unblocked tickets form a ready frontier. Independent tickets on that frontier can be implemented in separate worktrees and merged back into one integration branch.

Matt Pocock implement-spec task graph: independent ready tickets run in parallel, merge into an integration branch, then unlock the next frontier before final review.
Matt Pocock implement-spec task graph: independent ready tickets run in parallel, merge into an integration branch, then unlock the next frontier before final review.

This is a much stronger multi-agent pattern than choosing an arbitrary number of agents and hoping their work does not collide.

The implication is important: parallel agents reward good decomposition and clean module seams. If every ticket modifies the same central files, the task graph may say the work is logically independent while the codebase still makes it operationally coupled.

Bugs: evidence before edits

The debugging discipline is roughly:

output 1 line
reproduce -> minimise -> hypothesise -> instrument -> fix -> regression test

The key change is that the agent must build a feedback loop that goes red on the bug before it starts patching. That converts debugging from repeated guessing into a sequence of experiments.

Use:

output 1 line
/diagnosing-bugs

Then let tdd and code-review close the loop around the fix.

A useful warning sign is an agent editing several files before it can reliably reproduce the failure. At that point the workflow has become hypothesis-free patching rather than diagnosis.

TDD is the implementation heartbeat

The repository treats TDD as a fast feedback mechanism, not as a phase where the agent writes a large test suite first.

output 1 line
RED -> GREEN -> REFACTOR -> next vertical slice

The useful unit is one behavior at a time. That keeps the agent close to evidence, limits the amount of unverified code it can accumulate and makes failures easier to localize.

The same idea explains why to-spec records testing seams and why codebase-design favors deep modules with small, testable interfaces. Planning, architecture and TDD are supposed to line up around the same seams.

Persistent project memory is the hidden superpower

A strong workflow continuously moves important information out of ephemeral chat and into durable artifacts: GLOSSARY.md, ADRs, docs/agents/, issue tracker specs, tickets and committed code.

Matt Pocock Skills durable memory: temporary discussion becomes resolved decisions, versioned knowledge and context pointers consumed by future agent sessions.
Matt Pocock Skills durable memory: temporary discussion becomes resolved decisions, versioned knowledge and context pointers consumed by future agent sessions.

This is more robust than hoping the model "remembers the project". Important concepts become explicit, reviewable and reusable.

The complementary idea from writing-for-agents is the context pointer: a small reference that tells the agent both what information exists and when it should retrieve it. A perfect document behind a weak pointer is still unreliable because the agent may never load it.

That leads to a useful context-engineering rule: keep durable knowledge outside the always-loaded prompt, but make the pointer to it precise enough that retrieval is predictable.

Code review asks two independent questions

Matt Pocock's code-review is not a generic "find bugs" pass. It compares a diff against a fixed point and reviews it along two independent axes.

Matt Pocock code-review: the same diff is reviewed independently for repository standards and for fidelity to the originating specification.
Matt Pocock code-review: the same diff is reviewed independently for repository standards and for fidelity to the originating specification.

The Standards axis asks: is it built right? It reads the repository's documented coding standards and uses a code-smell baseline where local rules are silent.

The Spec axis asks: is it the right thing? It checks the change against the originating issue or specification.

The two analyses run separately so one line of reasoning does not contaminate the other. The result is intentionally not collapsed into one blended verdict because a change can be beautifully written and still implement the wrong requirement, or satisfy the requirement while violating important repository conventions.

Why deep modules matter more with coding agents

Coding agents can generate code much faster than humans can absorb the resulting complexity. That makes architectural depth more important, not less.

A shallow module exposes a large interface while hiding little complexity. A deep module exposes a small interface while hiding substantial internal complexity behind a stable seam.

codebase-design pushes toward deep modules because they give agents three advantages at once: a smaller context surface, clearer ownership boundaries and a high-level interface that can be tested without binding tests to implementation details.

That is also why multi-agent execution tends to work better in codebases with strong seams. Architectural modularity becomes a concurrency primitive.

What to learn first

Start with this small set:

Add the deeper primitives later:

A pragmatic first-week adoption path

Do not try to redesign your whole development process on day one.

The goal is not to invoke more skills. The goal is to make the engineering loop more explicit, more inspectable and more evidence-driven.

Common mistakes

Instead ofPrefer
Giving a vague large feature directly to an implementation agentgrill-with-docs then to-spec when a durable record is needed
Writing a spec for every tiny taskSkip directly to implement when the decision is clear and the work fits one context
Splitting tickets into database, backend, frontend and testsVertical tracer-bullet tickets
Launching many agents arbitrarilyLet ticket dependencies determine concurrency
Guessing at bug causesdiagnosing-bugs and an explicit feedback loop
Writing all tests before all implementationSmall Red -> Green -> Refactor slices
Keeping domain knowledge only in conversation historyGLOSSARY.md, ADRs and tracker artifacts
Feeding every document into every promptStrong context pointers and retrieval
Treating code review as one generic scoreKeep Standards and Spec as separate questions
Finishing a session without improving the environmentretro

The architecture in one sentence

Human-directed workflows + reusable engineering disciplines + persistent project knowledge + fast feedback loops.

The important lesson is not any individual prompt. It is the separation of orchestration, engineering discipline, durable knowledge and evidence-producing feedback.

noteIf you remember only three flows, remember these: grill-with-docs -> to-spec -> implement for a normal feature that needs a durable record, grill-with-docs -> to-spec -> to-tickets -> implement-spec for a larger initiative, and diagnosing-bugs -> tdd -> code-review for a hard bug.

If you want to go beyond this quick-start guide, these are the most useful next stops.

A simple learning path is: watch the "5 Agent Skills" overview -> read the AI Hero Skills catalog -> try one real feature with grill-with-docs -> to-spec -> implement -> return to the individual skill pages when something feels unclear.

Sources

← howtoz