The first ten covered configuration and habit. These ten cover the machinery — permission modes, diagnostics, context economics, checkpoints, and the telemetry that tells you what any of it is costing. Same ranking rule: consensus first, taste last.
Chapter one argued that Pi's minimalism is the product, not a gap, and that the highest-leverage changes are configuration rather than installs. That still holds. But it also left something unresolved. The recommendation to sandbox before turning off supervision sat at number five with a note that it probably belonged at two — and the reason it didn't get there is that "containerise everything" is a big ask for someone who just wants to try a coding agent on a Tuesday afternoon.
This chapter opens by closing that gap, then works outward. The rough shape:
One structural observation before the list. Since chapter one, I went through the full published package index rather than a handful of write-ups, and the shape of the ecosystem is clearer: several thousand packages, extremely long-tailed, with heavy duplication in a few categories. There are dozens of statuslines, dozens of subagent frameworks, dozens of memory layers. Where a category has ten near-identical implementations, that tells you the need is real and the solution is not yet settled — which is usually the signal to write thirty lines yourself rather than adopt someone's abstraction. I've flagged those cases.
Extensions run with your full system permissions and can execute arbitrary code, and Pi's own documentation says plainly to install only from sources you trust. That warning is more load-bearing in this chapter than the last, because several items below are packages whose entire job is to sit between the model and your filesystem. A compromised guard is worse than no guard. Read the source.
Chapter one's sandboxing recommendation was correct and widely ignored, because containerisation is a project and most people want a switch. This is the switch, and it is the most duplicated category in the entire ecosystem — which is exactly what you'd expect for a need that is universal and a design Pi deliberately declined to make for you.
The pattern everyone converges on is a small number of named modes. One well-developed implementation offers three: yolo (everything allowed), safe (rule-based checks that ask about unknown bash commands), and read-only (no repo or home writes, built-in edits restricted to /tmp, bash limited to safe read-only commands), switched with /mode, with project rules merging over global ones. Others reimplement Claude Code's ask/plan/auto triad, add per-mode model profiles, or gate on a real bash AST rather than string matching.
# Widely used, actively maintained
pi install npm:@gotgenes/pi-permission-system
# OS-level sandboxing plus interactive prompts
pi install npm:pi-sandbox
# Declarative modes with bubblewrap / sandbox-exec and tree-sitter bash gating
pi install npm:pi-permission-modes
The mechanism underneath all of them is one event. Pi's tool_call hook fires before a tool executes and can block it, and event.input is mutable, so a handler can patch arguments or refuse outright. The documentation's own first example is a confirmation prompt on rm -rf. That is roughly fifteen lines:
pi.on("tool_call", async (event, ctx) => {
if (event.toolName === "bash" && event.input.command?.includes("rm -rf")) {
const ok = await ctx.ui.confirm("Dangerous!", "Allow rm -rf?");
if (!ok) return { block: true, reason: "Blocked by user" };
}
});
Which is the honest recommendation here: if you want real isolation, go back to chapter one and containerise, because a TypeScript handler in the same process as the agent is a speed bump, not a boundary. If you want a speed bump — and a speed bump genuinely prevents the common accident — write those fifteen lines yourself and know exactly what they do. Reach for a package when you want the mode-switching UI and merged project rules, which are more tedious than they look.
Pi ships four tools: read, write, edit, bash. Nothing in that set tells the model that the symbol it just referenced doesn't exist. Without diagnostics, the agent writes code, runs the test suite, reads a stack trace, and works backwards — an expensive round trip to discover something your editor underlines instantly.
# Broadest: LSP, linters, formatters, type-checking, structural analysis
pi install npm:pi-lens
# Focused, declarative LSP diagnostics and navigation
pi install npm:pi-lsp
# Pre-write syntax validation via tree-sitter WASM grammars
pi install npm:pi-tree-sitter
The third is worth separating out conceptually. LSP integration is reactive — the model writes, then learns what broke. Pre-write validation is preventive — syntactically invalid edits are rejected before they land. They compose well, and the preventive layer is cheap in a way the reactive one isn't, since it costs no model tokens at all.
A related family worth knowing about: several packages replace the built-in edit tool with hash-anchored editing, where each line gets a short stable identifier and stale or ambiguous anchors are rejected rather than fuzzy-matched. This targets a specific failure — the model editing a file based on a version it read six turns ago. Pi's extension API explicitly supports overriding built-in tools by registering a tool with the same name, and inherits the built-in renderer for any slot your override omits, so these swaps are less invasive than they sound.
If you write a custom tool that touches files, use withFileMutationQueue(). Tool calls run in parallel by default, and without the queue two tools can read the same original file, compute different updates, and have one silently overwrite the other.
Pi compacts automatically as you approach the limit. That is a floor, not a strategy, and the density of packages in this category — among the most downloaded in the ecosystem — says the default leaves value on the table.
Three distinct approaches, worth telling apart before you pick:
| Approach | What it does | Trade-off |
|---|---|---|
| Output compression | Routes bash, read, grep, and find through a compressor so verbose tool output never reaches the model at full size | Cheapest and safest; savings depend on how noisy your tools are |
| Smarter compaction | Replaces Pi's summarisation with layered, vector-backed, or observation-preserving schemes | Keeps more of what matters, but you are trusting someone's heuristic about what matters |
| Reversible folding | Folds stale tool output out of the model's view while leaving it in the session, recoverable on demand | The conceptually cleanest — nothing is destroyed, only hidden |
The reversible variety is the most interesting of the three, and fits Pi's architecture best: the session is plain JSONL on disk, so hiding something from the model's context window and deleting it are genuinely different operations. Several packages exploit this, and one advertises recovery of any original on demand through a context-tree query.
Before installing any of them, install nothing and look. A context inspector — there are several — shows you the hidden components of your window: the base system prompt, tool definitions, and extension injections. It is common to discover that the expensive thing is not your conversation but twenty tool schemas from packages you installed in chapter one and never use. That diagnosis is free; the fix might be pi uninstall rather than another package.
The hooks are all documented and stable, incidentally: context fires before each LLM call with a safe-to-modify deep copy of messages, and session_before_compact can cancel compaction or supply your own summary. Rolling your own pruning rule is a genuinely small extension.
Chapter one recommended goal tracking for sessions that outlive your attention span. This is the smaller, more frequently useful sibling: a visible checklist for the current piece of work, which both you and the model can see and update.
The distinction matters. A goal is one durable objective with an audit at the end. A todo list is decomposition — the six things this refactor touches, ticked off as they land. Goal tracking answers "are we done?"; a todo list answers "what's left, and did it forget the third item?" That second question turns out to be the common failure mode in medium-length sessions.
pi install npm:@nguyenquangthai/pi-todo # session checklist with a live TUI overlay
pi install npm:pi-todo-rail # branch-aware, pinned above the editor
Implementations vary on two axes: where state lives (session-scoped, branch-scoped, or a durable TODO.md in the repo), and whether it renders as a widget. File-backed lists survive across sessions and can be committed, which some people love and others find noisy. Branch-aware lists follow your git branch, which matches how most work is actually organised.
This is another category where the sheer number of implementations is the message. Pi's docs list stateful tools — todo lists specifically — as an example use case for extensions, and state stored in tool result details is reconstructed on session start, so branching works correctly. If none of the published ones match how you think about tasks, this is a very good second extension to write.
This is the correction I most want to make to chapter one. I praised the session tree for making risky refactors safe, because you can jump back up the tree and abandon a branch. That is true of the conversation. It is not true of your files.
Navigating the tree rewinds what the model remembers. The edits it already wrote to disk stay written. If you fork away from a bad branch after the agent has rewritten nine files, you have a clean conversation and a dirty working tree — arguably the worst of both, because the model no longer remembers doing the thing you now have to undo by hand.
pi install npm:pi-rewind # per-tool snapshots, /rewind, Esc+Esc, redo stack
pi install npm:pi-workspace-history # workspace-level undo and redo
pi install npm:pi-rewind-hook # automatic git checkpoints per turn
Two philosophies here. Snapshot-based tools keep their own per-tool-call copies, which is precise and works in repos with messy git state. Git-based checkpointing commits or stashes at each turn, which means your existing tooling can inspect the history but adds noise to the repo. Pi's own documentation lists git checkpointing — stash at each turn, restore on branch — as an example extension use case, which suggests the maintainers consider it a reasonable thing to build rather than ship.
The cheap alternative, which costs nothing and works today: commit before you let the agent do anything ambitious. Most people who install a rewind package are buying granularity they could approximate with discipline. Buy it if you know your discipline is unreliable — mine is.
A model reviewing its own output in the same context is not reviewing; it is agreeing with itself with extra steps. The convergent practice is a /review command that runs against a diff in an isolated context, ideally on a different model.
pi install npm:pi-pr-review # parallel tiered reviewers, COMMENT-only publishing
pi install npm:@zephyrdeng/pi-review
The good implementations share three properties. They are read-only — review produces findings, not edits, which keeps the pass honest and cheap. They cover several targets: a GitHub PR checked out locally, a diff against a base branch, uncommitted changes, or a specific commit. And they respect a project-level guidelines file, so "we don't do that here" is enforced once rather than re-explained per review.
One package publishes review comments to GitHub as COMMENT-only — never approving or requesting changes. That restraint is correct and worth insisting on: an agent should not be able to approve a pull request, and any tool that offers to is solving the wrong problem.
A related and slightly odd genre has emerged: reviewers that specifically detect AI-generated slop — stub implementations, placeholder comments, hedging language, deferrals, characteristic sentence patterns — and report findings without editing. Whether you find that useful depends on how much agent-written code you are shipping, but the existence of the category is itself a data point about what unsupervised agents produce.
The complementary practice: make review part of the deal in APPEND_SYSTEM.md from item 1, so the agent proposes review at natural boundaries rather than waiting for you to remember.
The most underrated item in either chapter. A structured questionnaire tool lets the model put a question to you when it would otherwise guess, with typed options instead of free-form prose — and answering three multiple-choice questions is faster and more accurate than writing a paragraph you didn't know you needed to write.
pi install npm:pi-ask-user # searchable split-pane selection, multi-select, freeform
pi install npm:pi-interview # interactive form overlay
Why this beats the obvious alternative: without it, a model that is uncertain either guesses — badly, silently, and expensively — or asks in prose, which invites a prose answer, which invites ambiguity. Structured options make uncertainty cheap to resolve. One extension inverts the flow entirely, extracting the questions from Pi's response and presenting them one at a time rather than making you answer five embedded questions in a single reply.
This pairs directly with chapter one's item 8. Armin Ronacher's /discuss planning interviewer covers the same ground conversationally — inspect the project, ask focused questions in short rounds, stop when the plan is clear. If you tried /discuss and liked the interview but wanted it during implementation rather than only before it, this is that.
Add a line to APPEND_SYSTEM.md telling the agent to ask rather than assume when a decision is architecturally significant or hard to reverse. The tool without the instruction goes unused; models default to confident.
Chapter one described a 285,000-URL scrape running roughly ninety minutes. Nobody watches a terminal for ninety minutes. The moment your sessions get long, the bottleneck stops being the agent and starts being your attention.
pi install npm:task-complete-notify # desktop notification and chime on completion
pi install npm:pi-background-tasks # durable background shell tasks
Two capabilities, often conflated. Notification tells you the agent finished or needs input; the better implementations fire only when you are actually away and distinguish "done" from "blocked on a question." Background execution lets long shell work run without holding the session hostage, which matters because a ninety-minute task that blocks the loop also blocks your ability to steer it.
The event to hook, if you write this yourself, is agent_settled rather than agent_end. Pi's docs are explicit: agent_end fires when a run ends, but Pi may still auto-retry, auto-compact and retry, or continue with queued follow-ups. agent_settled is the one that means Pi will not continue on its own. Getting this wrong produces notifications that fire three times per task, which is how people end up disabling notifications.
Beyond desktop alerts there is a whole remote-control tier — web UIs, phone pairing, Telegram, Slack, Matrix bridges. Genuinely useful if you supervise long runs away from your desk. Also a considerably larger attack surface on a tool that already runs with your full permissions. Ranked here rather than higher for that reason.
Pi shows tokens and cost for the current session via /session, and the footer carries context usage. What it does not give you is the cross-session picture: what this feature cost, which model is eating your budget, how close you are to a subscription limit before you hit it mid-task.
pi install npm:@mtrojnar/pi-usage # usage and rate limits, startup report and footer
pi install npm:@latentminds/pi-quotas # remaining quota across several providers
The strong argument for instrumenting: chapter one's item 2 recommended a model lineup with cheap models for bulk work, and item 3 recommended pointing search summarisation at a cheap model. Both are optimisations you cannot evaluate without measurement. Installing the lineup and never checking whether the routing helps is cargo cult.
Two adjacent categories that show up in the same neighbourhood, with a caution attached to each. Spending guards track cost per task and pause at a threshold so you can continue, refine, or stop — sensible for autonomous runs. Multi-account rotation automatically fails over between accounts on the same provider when one hits its limit; this is popular, and it is worth reading your provider's terms before using it, since "rotate through several subscriptions to exceed the limits of one" is precisely what the limits exist to prevent.
There is also a local-analytics genre that turns the session transcripts Pi already writes into a dashboard — cost attribution per feature or PR, and usage patterns over time. Since sessions are plain JSONL organised by working directory, this needs no instrumentation at all; the data is already on your disk. That property is the same one that makes Pi auditable in regulated environments, noted in chapter one.
Two halves, and the second is the real recommendation.
Sync. Once you have accumulated an APPEND_SYSTEM.md, a model lineup, a permission policy, and a handful of extensions, that configuration is worth more than any single package in it. Packages exist to sync ~/.pi/agent through a private git repo — the responsible ones with explicit guards to keep credentials out, which matters, because auth.json lives in that tree and a config repo that quietly commits your API keys is a bad day. Verify the guard yourself rather than trusting the README.
Write one. Pi's extension documentation opens with a line that reads like a thesis statement for the whole tool:
pi can create extensions. Ask it to build one for your use case.
The API is genuinely small. An extension is a TypeScript module exporting a default function that receives ExtensionAPI; it is loaded through jiti so TypeScript works without a build step; drop it in ~/.pi/agent/extensions/ and /reload picks it up. The event list covers the lifecycle at a useful granularity — tool_call to block or patch, context to prune, before_agent_start to inject, session_before_compact to summarise your own way, agent_settled to notify.
There are progressive learning guides published as skills — install one into ~/.pi/agent/skills/ and it loads when you mention extension development — which is a pleasingly recursive way to learn: ask Pi to teach you to extend Pi, using a skill that exists to teach Pi to extend Pi.
The reason this is last is the reason chapter one warned about lists like these. Twenty recommendations is nineteen more than the tool's design implies you need. Every item in both chapters is a place where someone hit a gap and packaged their answer. Yours may differ. The measure of whether you have understood Pi is not how many of these twenty you installed — it is whether, the next time you notice a gap, your first instinct is to search npm or to open an editor.
Across both chapters, if I had to defend a minimum viable setup to someone starting today — six things:
| # | Change | Why it survives the cut |
|---|---|---|
| 1 | APPEND_SYSTEM.md | Shapes every session; costs nothing; survives everything |
| 2 | A model lineup | Switching by keystroke changes what work you attempt |
| 4 | Session tree fluency | Pi's genuinely distinctive feature; free |
| 11 | A permission mode | The realistic version of "sandbox it" |
| 12 | Diagnostics | Removes the most expensive feedback loop the agent has |
| 15 | File checkpoints | Because item 4 protects the conversation, not the files |
Three of those six are configuration or habit rather than installs. That ratio is not an accident, and it has held up across both chapters: the changes that matter most are the ones you make to how you work, and the ones that matter least are the ones with the highest download counts.
The package churn is worse than chapter one implied. Many entries in the index were updated within the last day or two, package names move between scopes, forks proliferate, and several of the most-downloaded packages are forks of forks. Verify anything here against pi.dev/packages before installing. Treat the categories as durable and the specific names as volatile.
Download counts are a weak signal and I have leaned on them lightly. They reflect installs, not retention, and the pattern documented in chapter one — install broadly, prune back to four or five — means the gap between installs and daily use is probably enormous. A high count tells you a category matters. It does not tell you that particular implementation is good.
Finally: neither chapter is a substitute for two weeks of use. The list you would write after those two weeks is the only one calibrated to your work, and it will be shorter than forty items.
tool_call blocking and input mutation, agent_end vs agent_settled, context and session_before_compact hooks, built-in tool overriding and renderer inheritance, withFileMutationQueue(), extension locations and /reload, the security warning, and the "ask pi to build one" framing./review command supporting PRs, base-branch diffs, uncommitted changes, and specific commits with project-level review guidelines.~/.pi/agent/skills/ and the aggressive-extensibility positioning quoted from Pi's own docs.APPEND_SYSTEM.md (1), model lineups (2), web access (3), the session tree (4), sandboxing (5), goal tracking (7), and /discuss (8).