Tutorial · Coding agents · Chapter two

Good defaults for Pi, chapter two: changes 11 to 20

The first ten covered configuration and habit. These ten cover the machinery — permission modes, diagnostics, context economics, checkpoints, and the telemetry that tells you what any of it is costing. Same ranking rule: consensus first, taste last.

← Chapter one Good defaults for Pi: ten changes to make on day one APPEND_SYSTEM.md, model lineups, web access, the session tree, sandboxing, and five more.

Where chapter one left off

Chapter one argued that Pi's minimalism is the product, not a gap, and that the highest-leverage changes are configuration rather than installs. That still holds. But it also left something unresolved. The recommendation to sandbox before turning off supervision sat at number five with a note that it probably belonged at two — and the reason it didn't get there is that "containerise everything" is a big ask for someone who just wants to try a coding agent on a Tuesday afternoon.

This chapter opens by closing that gap, then works outward. The rough shape:

One structural observation before the list. Since chapter one, I went through the full published package index rather than a handful of write-ups, and the shape of the ecosystem is clearer: several thousand packages, extremely long-tailed, with heavy duplication in a few categories. There are dozens of statuslines, dozens of subagent frameworks, dozens of memory layers. Where a category has ten near-identical implementations, that tells you the need is real and the solution is not yet settled — which is usually the signal to write thirty lines yourself rather than adopt someone's abstraction. I've flagged those cases.

The security note from chapter one still applies

Extensions run with your full system permissions and can execute arbitrary code, and Pi's own documentation says plainly to install only from sources you trust. That warning is more load-bearing in this chapter than the last, because several items below are packages whose entire job is to sit between the model and your filesystem. A compromised guard is worse than no guard. Read the source.


The next ten, ranked

11

Install a permission mode — the practical version of item 5

Near-universal consensus

Chapter one's sandboxing recommendation was correct and widely ignored, because containerisation is a project and most people want a switch. This is the switch, and it is the most duplicated category in the entire ecosystem — which is exactly what you'd expect for a need that is universal and a design Pi deliberately declined to make for you.

The pattern everyone converges on is a small number of named modes. One well-developed implementation offers three: yolo (everything allowed), safe (rule-based checks that ask about unknown bash commands), and read-only (no repo or home writes, built-in edits restricted to /tmp, bash limited to safe read-only commands), switched with /mode, with project rules merging over global ones. Others reimplement Claude Code's ask/plan/auto triad, add per-mode model profiles, or gate on a real bash AST rather than string matching.

# Widely used, actively maintained
pi install npm:@gotgenes/pi-permission-system

# OS-level sandboxing plus interactive prompts
pi install npm:pi-sandbox

# Declarative modes with bubblewrap / sandbox-exec and tree-sitter bash gating
pi install npm:pi-permission-modes

The mechanism underneath all of them is one event. Pi's tool_call hook fires before a tool executes and can block it, and event.input is mutable, so a handler can patch arguments or refuse outright. The documentation's own first example is a confirmation prompt on rm -rf. That is roughly fifteen lines:

pi.on("tool_call", async (event, ctx) => {
  if (event.toolName === "bash" && event.input.command?.includes("rm -rf")) {
    const ok = await ctx.ui.confirm("Dangerous!", "Allow rm -rf?");
    if (!ok) return { block: true, reason: "Blocked by user" };
  }
});

Which is the honest recommendation here: if you want real isolation, go back to chapter one and containerise, because a TypeScript handler in the same process as the agent is a speed bump, not a boundary. If you want a speed bump — and a speed bump genuinely prevents the common accident — write those fifteen lines yourself and know exactly what they do. Reach for a package when you want the mode-switching UI and merged project rules, which are more tedious than they look.

12

Give the agent your editor's diagnostics

Near-universal consensus

Pi ships four tools: read, write, edit, bash. Nothing in that set tells the model that the symbol it just referenced doesn't exist. Without diagnostics, the agent writes code, runs the test suite, reads a stack trace, and works backwards — an expensive round trip to discover something your editor underlines instantly.

# Broadest: LSP, linters, formatters, type-checking, structural analysis
pi install npm:pi-lens

# Focused, declarative LSP diagnostics and navigation
pi install npm:pi-lsp

# Pre-write syntax validation via tree-sitter WASM grammars
pi install npm:pi-tree-sitter

The third is worth separating out conceptually. LSP integration is reactive — the model writes, then learns what broke. Pre-write validation is preventive — syntactically invalid edits are rejected before they land. They compose well, and the preventive layer is cheap in a way the reactive one isn't, since it costs no model tokens at all.

A related family worth knowing about: several packages replace the built-in edit tool with hash-anchored editing, where each line gets a short stable identifier and stale or ambiguous anchors are rejected rather than fuzzy-matched. This targets a specific failure — the model editing a file based on a version it read six turns ago. Pi's extension API explicitly supports overriding built-in tools by registering a tool with the same name, and inherits the built-in renderer for any slot your override omits, so these swaps are less invasive than they sound.

If you write a custom tool that touches files, use withFileMutationQueue(). Tool calls run in parallel by default, and without the queue two tools can read the same original file, compute different updates, and have one silently overwrite the other.

13

Treat context as a budget, and manage it deliberately

Strong consensus on the problem

Pi compacts automatically as you approach the limit. That is a floor, not a strategy, and the density of packages in this category — among the most downloaded in the ecosystem — says the default leaves value on the table.

Three distinct approaches, worth telling apart before you pick:

ApproachWhat it doesTrade-off
Output compressionRoutes bash, read, grep, and find through a compressor so verbose tool output never reaches the model at full sizeCheapest and safest; savings depend on how noisy your tools are
Smarter compactionReplaces Pi's summarisation with layered, vector-backed, or observation-preserving schemesKeeps more of what matters, but you are trusting someone's heuristic about what matters
Reversible foldingFolds stale tool output out of the model's view while leaving it in the session, recoverable on demandThe conceptually cleanest — nothing is destroyed, only hidden

The reversible variety is the most interesting of the three, and fits Pi's architecture best: the session is plain JSONL on disk, so hiding something from the model's context window and deleting it are genuinely different operations. Several packages exploit this, and one advertises recovery of any original on demand through a context-tree query.

Before installing any of them, install nothing and look. A context inspector — there are several — shows you the hidden components of your window: the base system prompt, tool definitions, and extension injections. It is common to discover that the expensive thing is not your conversation but twenty tool schemas from packages you installed in chapter one and never use. That diagnosis is free; the fix might be pi uninstall rather than another package.

The hooks are all documented and stable, incidentally: context fires before each LLM call with a safe-to-modify deep copy of messages, and session_before_compact can cancel compaction or supply your own summary. Rolling your own pruning rule is a genuinely small extension.

14

Add a todo list — distinct from goals, and more useful day to day

Broad consensus

Chapter one recommended goal tracking for sessions that outlive your attention span. This is the smaller, more frequently useful sibling: a visible checklist for the current piece of work, which both you and the model can see and update.

The distinction matters. A goal is one durable objective with an audit at the end. A todo list is decomposition — the six things this refactor touches, ticked off as they land. Goal tracking answers "are we done?"; a todo list answers "what's left, and did it forget the third item?" That second question turns out to be the common failure mode in medium-length sessions.

pi install npm:@nguyenquangthai/pi-todo    # session checklist with a live TUI overlay
pi install npm:pi-todo-rail                # branch-aware, pinned above the editor

Implementations vary on two axes: where state lives (session-scoped, branch-scoped, or a durable TODO.md in the repo), and whether it renders as a widget. File-backed lists survive across sessions and can be committed, which some people love and others find noisy. Branch-aware lists follow your git branch, which matches how most work is actually organised.

This is another category where the sheer number of implementations is the message. Pi's docs list stateful tools — todo lists specifically — as an example use case for extensions, and state stored in tool result details is reconstructed on session start, so branching works correctly. If none of the published ones match how you think about tasks, this is a very good second extension to write.

15

Add file-level checkpoints — the session tree does not cover this

Broad consensus, commonly misunderstood

This is the correction I most want to make to chapter one. I praised the session tree for making risky refactors safe, because you can jump back up the tree and abandon a branch. That is true of the conversation. It is not true of your files.

Navigating the tree rewinds what the model remembers. The edits it already wrote to disk stay written. If you fork away from a bad branch after the agent has rewritten nine files, you have a clean conversation and a dirty working tree — arguably the worst of both, because the model no longer remembers doing the thing you now have to undo by hand.

pi install npm:pi-rewind              # per-tool snapshots, /rewind, Esc+Esc, redo stack
pi install npm:pi-workspace-history   # workspace-level undo and redo
pi install npm:pi-rewind-hook         # automatic git checkpoints per turn

Two philosophies here. Snapshot-based tools keep their own per-tool-call copies, which is precise and works in repos with messy git state. Git-based checkpointing commits or stashes at each turn, which means your existing tooling can inspect the history but adds noise to the repo. Pi's own documentation lists git checkpointing — stash at each turn, restore on branch — as an example extension use case, which suggests the maintainers consider it a reasonable thing to build rather than ship.

The cheap alternative, which costs nothing and works today: commit before you let the agent do anything ambitious. Most people who install a rewind package are buying granularity they could approximate with discipline. Buy it if you know your discipline is unreliable — mine is.

16

Make review a separate pass with a different model

Broad consensus among heavy users

A model reviewing its own output in the same context is not reviewing; it is agreeing with itself with extra steps. The convergent practice is a /review command that runs against a diff in an isolated context, ideally on a different model.

pi install npm:pi-pr-review        # parallel tiered reviewers, COMMENT-only publishing
pi install npm:@zephyrdeng/pi-review

The good implementations share three properties. They are read-only — review produces findings, not edits, which keeps the pass honest and cheap. They cover several targets: a GitHub PR checked out locally, a diff against a base branch, uncommitted changes, or a specific commit. And they respect a project-level guidelines file, so "we don't do that here" is enforced once rather than re-explained per review.

One package publishes review comments to GitHub as COMMENT-only — never approving or requesting changes. That restraint is correct and worth insisting on: an agent should not be able to approve a pull request, and any tool that offers to is solving the wrong problem.

A related and slightly odd genre has emerged: reviewers that specifically detect AI-generated slop — stub implementations, placeholder comments, hedging language, deferrals, characteristic sentence patterns — and report findings without editing. Whether you find that useful depends on how much agent-written code you are shipping, but the existence of the category is itself a data point about what unsupervised agents produce.

The complementary practice: make review part of the deal in APPEND_SYSTEM.md from item 1, so the agent proposes review at natural boundaries rather than waiting for you to remember.

17

Let the agent ask you structured questions

Underrated; growing fast

The most underrated item in either chapter. A structured questionnaire tool lets the model put a question to you when it would otherwise guess, with typed options instead of free-form prose — and answering three multiple-choice questions is faster and more accurate than writing a paragraph you didn't know you needed to write.

pi install npm:pi-ask-user        # searchable split-pane selection, multi-select, freeform
pi install npm:pi-interview       # interactive form overlay

Why this beats the obvious alternative: without it, a model that is uncertain either guesses — badly, silently, and expensively — or asks in prose, which invites a prose answer, which invites ambiguity. Structured options make uncertainty cheap to resolve. One extension inverts the flow entirely, extracting the questions from Pi's response and presenting them one at a time rather than making you answer five embedded questions in a single reply.

This pairs directly with chapter one's item 8. Armin Ronacher's /discuss planning interviewer covers the same ground conversationally — inspect the project, ask focused questions in short rounds, stop when the plan is clear. If you tried /discuss and liked the interview but wanted it during implementation rather than only before it, this is that.

Add a line to APPEND_SYSTEM.md telling the agent to ask rather than assume when a decision is architecturally significant or hard to reverse. The tool without the instruction goes unused; models default to confident.

18

Notifications and background tasks, once sessions outlast your attention

Situational; obvious once needed

Chapter one described a 285,000-URL scrape running roughly ninety minutes. Nobody watches a terminal for ninety minutes. The moment your sessions get long, the bottleneck stops being the agent and starts being your attention.

pi install npm:task-complete-notify     # desktop notification and chime on completion
pi install npm:pi-background-tasks      # durable background shell tasks

Two capabilities, often conflated. Notification tells you the agent finished or needs input; the better implementations fire only when you are actually away and distinguish "done" from "blocked on a question." Background execution lets long shell work run without holding the session hostage, which matters because a ninety-minute task that blocks the loop also blocks your ability to steer it.

The event to hook, if you write this yourself, is agent_settled rather than agent_end. Pi's docs are explicit: agent_end fires when a run ends, but Pi may still auto-retry, auto-compact and retry, or continue with queued follow-ups. agent_settled is the one that means Pi will not continue on its own. Getting this wrong produces notifications that fire three times per task, which is how people end up disabling notifications.

Beyond desktop alerts there is a whole remote-control tier — web UIs, phone pairing, Telegram, Slack, Matrix bridges. Genuinely useful if you supervise long runs away from your desk. Also a considerably larger attack surface on a tool that already runs with your full permissions. Ranked here rather than higher for that reason.

19

Instrument cost and quota before you tune anything

Preference, with one strong argument

Pi shows tokens and cost for the current session via /session, and the footer carries context usage. What it does not give you is the cross-session picture: what this feature cost, which model is eating your budget, how close you are to a subscription limit before you hit it mid-task.

pi install npm:@mtrojnar/pi-usage      # usage and rate limits, startup report and footer
pi install npm:@latentminds/pi-quotas  # remaining quota across several providers

The strong argument for instrumenting: chapter one's item 2 recommended a model lineup with cheap models for bulk work, and item 3 recommended pointing search summarisation at a cheap model. Both are optimisations you cannot evaluate without measurement. Installing the lineup and never checking whether the routing helps is cargo cult.

Two adjacent categories that show up in the same neighbourhood, with a caution attached to each. Spending guards track cost per task and pause at a threshold so you can continue, refine, or stop — sensible for autonomous runs. Multi-account rotation automatically fails over between accounts on the same provider when one hits its limit; this is popular, and it is worth reading your provider's terms before using it, since "rotate through several subscriptions to exceed the limits of one" is precisely what the limits exist to prevent.

There is also a local-analytics genre that turns the session transcripts Pi already writes into a dashboard — cost attribution per feature or PR, and usage patterns over time. Since sessions are plain JSONL organised by working directory, this needs no instrumentation at all; the data is already on your disk. That property is the same one that makes Pi auditable in regulated environments, noted in chapter one.

20

Sync your config, then write your own extension

The one item that makes the other nineteen optional

Two halves, and the second is the real recommendation.

Sync. Once you have accumulated an APPEND_SYSTEM.md, a model lineup, a permission policy, and a handful of extensions, that configuration is worth more than any single package in it. Packages exist to sync ~/.pi/agent through a private git repo — the responsible ones with explicit guards to keep credentials out, which matters, because auth.json lives in that tree and a config repo that quietly commits your API keys is a bad day. Verify the guard yourself rather than trusting the README.

Write one. Pi's extension documentation opens with a line that reads like a thesis statement for the whole tool:

pi can create extensions. Ask it to build one for your use case.

The API is genuinely small. An extension is a TypeScript module exporting a default function that receives ExtensionAPI; it is loaded through jiti so TypeScript works without a build step; drop it in ~/.pi/agent/extensions/ and /reload picks it up. The event list covers the lifecycle at a useful granularity — tool_call to block or patch, context to prune, before_agent_start to inject, session_before_compact to summarise your own way, agent_settled to notify.

There are progressive learning guides published as skills — install one into ~/.pi/agent/skills/ and it loads when you mention extension development — which is a pleasingly recursive way to learn: ask Pi to teach you to extend Pi, using a skill that exists to teach Pi to extend Pi.

The reason this is last is the reason chapter one warned about lists like these. Twenty recommendations is nineteen more than the tool's design implies you need. Every item in both chapters is a place where someone hit a gap and packaged their answer. Yours may differ. The measure of whether you have understood Pi is not how many of these twenty you installed — it is whether, the next time you notice a gap, your first instinct is to search npm or to open an editor.


The combined shortlist

Across both chapters, if I had to defend a minimum viable setup to someone starting today — six things:

#ChangeWhy it survives the cut
1APPEND_SYSTEM.mdShapes every session; costs nothing; survives everything
2A model lineupSwitching by keystroke changes what work you attempt
4Session tree fluencyPi's genuinely distinctive feature; free
11A permission modeThe realistic version of "sandbox it"
12DiagnosticsRemoves the most expensive feedback loop the agent has
15File checkpointsBecause item 4 protects the conversation, not the files

Three of those six are configuration or habit rather than installs. That ratio is not an accident, and it has held up across both chapters: the changes that matter most are the ones you make to how you work, and the ones that matter least are the ones with the highest download counts.

Caveats, restated and sharpened

The package churn is worse than chapter one implied. Many entries in the index were updated within the last day or two, package names move between scopes, forks proliferate, and several of the most-downloaded packages are forks of forks. Verify anything here against pi.dev/packages before installing. Treat the categories as durable and the specific names as volatile.

Download counts are a weak signal and I have leaned on them lightly. They reflect installs, not retention, and the pattern documented in chapter one — install broadly, prune back to four or five — means the gap between installs and daily use is probably enormous. A high count tells you a category matters. It does not tell you that particular implementation is good.

Finally: neither chapter is a substitute for two weeks of use. The list you would write after those two weeks is the only one calibrated to your work, and it will be shorter than forty items.

Sources

  1. Pi documentation — Extensions — the primary source for this chapter. Event lifecycle, tool_call blocking and input mutation, agent_end vs agent_settled, context and session_before_compact hooks, built-in tool overriding and renderer inheritance, withFileMutationQueue(), extension locations and /reload, the security warning, and the "ask pi to build one" framing.
  2. Awesome Pi Coding Agent — extensions index — the full published package list with download figures and update dates; the basis for every claim about which categories are crowded and what the convergent implementations do.
  3. tomsej/pi-ext — the three-mode permission system (yolo / safe / read-only) with project-to-global rule merging, and a /review command supporting PRs, base-branch diffs, uncommitted changes, and specific commits with project-level review guidelines.
  4. NelsonBrandao/pi-agent-extensions — file-based cross-session todos with a TUI, a Codex-inspired review command, and the question-extraction pattern behind item 17. Explicitly notes that much of it was built by Pi itself.
  5. Dwsy/pi-extensions-skill — progressive extension-development guide distributed as a skill; the beginner-to-expert path referenced in item 20, and the extension-vs-skill distinction.
  6. creating-pi-extensions skill — command registration, keybinding conflict resolution, TUI overlay lifecycle, and XDG config storage for extension authors.
  7. "Pi Coding Agent: A Self-Documenting, Extensible AI Partner" — skills loading from ~/.pi/agent/skills/ and the aggressive-extensibility positioning quoted from Pi's own docs.
  8. qualisero/awesome-pi-agent — an earlier community index; useful mainly as a record of how much smaller the ecosystem was a few months ago.
  9. earendil-works/pi — main repository, carried over from chapter one for session storage, containerisation patterns, and the absence of a built-in permission system.
  10. Chapter one of this tutorial — for the items referenced by number throughout: APPEND_SYSTEM.md (1), model lineups (2), web access (3), the session tree (4), sandboxing (5), goal tracking (7), and /discuss (8).
← Back to chapter one Good defaults for Pi: ten changes to make on day one Configuration and habit before installs — the four changes that survive every reinstall.