
Planning with Files
計画を3つのファイルとしてディスクに置き、毎ターン読み直させる。文脈が消えても計画は残る
Planning with Filesとは
長い作業には特有の壊れ方があります。エージェントは計画を持っていたのに、文脈が消去または圧縮され、戻ってきたエージェントが自信満々に別の仕事を仕上げてしまう、という壊れ方です。このスキルは計画を会話から外し、ディスク上に移します。計画のファイル、分かったことのファイル、済んだことのファイルの3つで、毎ターンそれらをエージェントの目の前に戻します。計画がファイルである以上、消去や異常終了、圧縮を越えて残り、実行中に人が自分で読むこともできます。任意で有効にできる関門は、計画がそう示すまでエージェントに完了を宣言させません。開発元は盲検のA/B比較結果を公開しており、これは多くのスキルにはない点です。
Planning with Filesで何ができますか?
- 文脈の消去を越えて残る — 計画は会話ではなくディスク上にあるため、文脈の消去や異常終了、圧縮が起きても、何をしていたかが失われません。
- 毎ターン計画を差し戻す — エージェントの記憶に任せず、フックが毎ターンファイルを読み直させます。計画と単なる提案の違いはここにあります。
- 実行中に人が計画を読む — リポジトリ内の普通のMarkdownファイルなので、途中で開いて、エージェントが今どう認識しているかを確認できます。
- 分かったことと進捗を分ける — 学んだ内容と完了した内容が別のファイルになっているため、途中の発見がチェックリストの中に埋もれません。
- 早すぎる完了を拒む — 任意の関門を有効にすると、計画ファイルが認めるまでエージェントは完了を宣言できません。自信を伴う早期終了を防げます。
Planning with Filesを選ぶ前に
- 作業用のファイルがプロジェクト内に書き出されます。版管理に含めるかどうかを早めに決めてください。誤ってコミットされた計画ファイルは履歴の雑音になります。
- 毎ターン計画を読み直させる分、毎ターン費用がかかります。迷走する長い作業に比べれば安いものですが、計画の要らない短い作業では割に合いません。
よくある質問
Planning with Filesは商用利用できますか?
Planning with FilesはMITライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。
Planning with Filesはどの形で使えますか?
Planning with Filesはローカル実行の形で利用できます。
ドキュメント
OthmanAdi/planning-with-files のREADMEより転載(MIT)。 原文を読む ↗
Before and after /clear
Every coding agent loses its working memory when the context window resets. The plan does not have to die with it.
Without planning files
The agent re-reads the repo, asks you to restate the goal, and rediscovers work it already finished.
With planning-with-files
The transcript is illustrative; the ===BEGIN PLAN DATA=== block is the skill’s real injection format, written into context by the UserPromptSubmit hook from task_plan.md on disk. In the project’s internal recovery benchmark, a fresh session with the files on disk resumed in 5.0 turns on average against 13.3 for a raw agent (internal v1, author-run; method and limits in docs/evals.md).
| At a glance | |
|---|---|
| Plan files | 3 |
| Agents covered | 60+ |
| Pass rate (with skill) | 96.7% |
| Test suite | 417 green |
Survives /clear | yes |
The Problem
Claude Code and most AI agents suffer from:
- Volatile memory: the TodoWrite list disappears on context reset
- Goal drift: after 50+ tool calls, the original goals get crowded out
- Hidden errors: failures are not tracked, so the same mistakes repeat
- Context stuffing: everything crammed into the window instead of stored
The Solution: 3-File Pattern
For every complex task, create THREE files:
task_plan.md → Track phases and progress
findings.md → Store research and findings
progress.md → Session log and test results
The Core Principle
Context Window = RAM (volatile, limited)
Filesystem = Disk (persistent, unlimited)
→ Anything important gets written to disk.
In your project, exactly this lands on disk and nothing else:
your-project/
├── task_plan.md ← phases + checkboxes; the resume point after /clear
├── findings.md ← research notes and decisions, appended as you go
└── progress.md ← session log and test results
Parallel tasks get isolated directories instead: .planning/YYYY-MM-DD-slug/ with the same three files, selected via .active_plan (v2.36.0+). Plain markdown, gitignored by default, no runtime state anywhere else.
Why This Skill?
On December 29, 2025, Meta acquired Manus for $2 billion. In just 8 months, Manus went from launch to $100M+ revenue. Their secret? Context engineering.
“Markdown is my ‘working memory’ on disk. Since I process information iteratively and my active context has limits, Markdown files serve as scratch pads for notes, checkpoints for progress, building blocks for final deliverables.” — Manus AI
This skill packages that exact pattern for your coding agent.
The Manus Principles
| Principle | Implementation |
|---|---|
| Filesystem as memory | Store in files, not context |
| Attention manipulation | Re-read plan before decisions (hooks) |
| Error persistence | Log failures in plan file |
| Goal tracking | Checkboxes show progress |
| Completion verification | Stop hook checks all phases |
Benchmark Results
Methodology note: the 96.7% figure comes from the v2.21.0 evaluation run on
claude-sonnet-4-6(2026-03-06). It measures file-pattern fidelity (does the agent create and maintain the 3-file structure), not goal-drift over long autonomous runs. Newer models and the autonomous-mode work are not yet covered by this number. Full methodology, dataset, and assertion list: docs/evals.md.
Evaluated with Anthropic’s skill-creator framework: skill v2.21.0, model claude-sonnet-4-6, 2026-03-06. 10 parallel subagents, 5 task types, 30 objectively verifiable assertions, 3 blind A/B comparisons.
| Test | with_skill | without_skill |
|---|---|---|
| Pass rate (30 assertions) | 96.7% (29/30) | 6.7% (2/30) |
| 3-file pattern followed | 5/5 evals | 0/5 evals |
| Blind A/B wins | 3/3 (100%) | 0/3 |
| Avg rubric score | 10.0/10 | 6.8/10 |
Recovery after a context wipe
Internal benchmark, v1 (2026-07-06). Author-run against v3.4.0, harness-authored tasks, deterministic grading, no LLM grades anything. Treat it as the project’s own measurement, not an independent comparison. Full method, arms, disclosed limits, and grader validation: docs/evals.md.
Protocol: the session is hard-stopped at roughly half done, and a fresh session is told only “Continue the work in this directory.” Every graded run across every arm ended pytest-green (77/77), so the difference is re-orientation cost, not correctness.
With the planning files on disk, a resume took 5.0 turns on average; a raw agent took 13.3. Session catchup plus hook injection put phase state in front of the model before its first tool call, and the same run found no correctness penalty anywhere. An animated summary lives at docs/benchmark/index.html (rendered view).
Full methodology and results · Technical write-up
Quick Install
Claude Code, plugin route (ships everything: skill, hooks, slash commands):
/plugin marketplace add OthmanAdi/planning-with-files
/plugin install planning-with-files@planning-with-files
Every other agent, one line, 60+ agents via the Agent Skills standard:
npx skills add OthmanAdi/planning-with-files --skill planning-with-files -g
npm, to pin an exact version into a project or vendor it:
npm install planning-with-files
The package carries SKILL.md, scripts/ and templates/, so this is the route for locking a version into a repo’s dependencies or copying the skill in yourself. It does not register hooks on its own.
Pi Coding Agent, same npm package, wired up for you (skill, extension, status bar):
pi install npm:planning-with-files
Under a minute. Safe to re-run. Trigger it by typing /plan (plugin) or asking the agent to “plan this task”; the skill also self-triggers on multi-step tasks.
What each route actually ships:
| Route | Skill + scripts + templates | Slash commands | Hooks |
|---|---|---|---|
| Claude Code plugin | yes | yes | yes |
npx skills add | yes | no | frontmatter hooks, see note |
npm install | yes, under node_modules/ | no | no, copy the skill in yourself |
pi install npm: | yes | yes, Pi commands | yes, via the Pi extension |
| ClawHub / manual copy | yes | no | frontmatter hooks, see note |
Skill-route installs can end up silently hook-less (project trust not accepted, or frontmatter hooks not registering on project-level installs). The hooks are the differentiating mechanism, so if they matter to you, use the plugin route, then verify with /plan-doctor. Full matrix and the two silent killers: docs/installation.md.
Install acting up? Open your agent and say: “Read docs/installation.md and docs/troubleshooting.md from OthmanAdi/planning-with-files and fix my install.” Then run /plan-doctor.
🇸🇦 العربية / Arabic
npx skills add OthmanAdi/planning-with-files --skill planning-with-files-ar -g
🇩🇪 Deutsch / German
npx skills add OthmanAdi/planning-with-files --skill planning-with-files-de -g
🇪🇸 Español / Spanish
npx skills add OthmanAdi/planning-with-files --skill planning-with-files-es -g
🇨🇳 中文版 / Chinese (Simplified)
npx skills add OthmanAdi/planning-with-files --skill planning-with-files-zh -g
🇹🇼 正體中文版 / Chinese (Traditional)
npx skills add OthmanAdi/planning-with-files --skill planning-with-files-zht -g
These are real translations, not an English body with a translated description: the SKILL.md prose, the templates, and the user-facing output of check-complete, init-session and session-catchup are all localized. The status tokens stay literal English (**Status:** complete) on purpose, because check-complete.sh matches them with grep -F, so translating them would disable the completion gate.
Since v3.10.0 the variants also ship the full script surface: attestation, the Stop gate, the ledger, phase status and plan-doctor used to be canonical-only, which quietly made every non-English install a subset install. Full details, including what changed on the plugin route in v3.11.0, are in docs/languages.md.
They live under skills/i18n/, one directory deeper than the canonical skill. The install commands above are unchanged, because npx skills add resolves --skill by skill name across the whole repository. The Claude Code plugin scan reads skills/*/SKILL.md without recursing, so the plugin route registers the canonical skill alone and no longer carries five extra descriptions in every session’s system prompt. On that route the /plan-ar, /plan-de, /plan-es, /plan-zh and /plan-zht commands read the translated skill from disk instead of invoking it by name.
Copy the skill to your local folder:
macOS/Linux:
cp -r ~/.claude/plugins/cache/planning-with-files/planning-with-files/*/skills/planning-with-files ~/.claude/skills/
Windows (PowerShell):
Copy-Item -Recurse -Path "$env:USERPROFILE\.claude\plugins\cache\planning-with-files\planning-with-files\*\skills\planning-with-files" -Destination "$env:USERPROFILE\.claude\skills\"
All install methods: docs/installation.md.
Reference
Everything below is the technical half: how the hooks fire, every command, every supported platform, and the release history.
| How It Works | The hook loop, injection, and session recovery |
| Commands | All 13 slash commands |
| Works across 18+ platforms | Per-IDE setup and discovery paths |
| v3 Long-Running Agent Features | Modes, the completion gate, attestation, env vars |
| Key Rules · When to Use | The four rules, and when the pattern pays off |
| File Structure | What lands in your project, and the repository layout |
| FAQ | Context rot, plan mode, agent memory tools |
| Releases · Community | Version history and community forks |
| Documentation | Every guide in docs/ |
How It Works
The agent stops at the first rung that applies:
1. Task needs 3+ steps or 5+ tool calls? → create the three files first
2. Learned something? → append it to findings.md
3. Did something? → log it in progress.md
4. Phase done? → check it off in task_plan.md
5. Context died (/clear, crash)? → session catchup re-reads all three
6. Every phase complete? → only then does the Stop gate release (gated mode)
Hooks make steps 2 to 6 mechanical rather than optional: 5 lifecycle hooks on Claude Code, 7 on Codex, 8 on Pi re-inject the plan each turn, remind after writes, and check completion before stopping.
flowchart LR
A["agent works"] -->|"writes decisions, findings, errors"| F["task_plan.md<br/>findings.md<br/>progress.md"]
F -->|"hooks re-inject the plan<br/>at the start of each turn"| A
K["/clear · crash · compaction"] -.->|"wipes the context window"| A
F ==>|"session catchup re-reads the files"| R["fresh session resumes<br/>at the current phase"]
Session Recovery
When your context fills up and you run /clear, the skill recovers the previous session automatically:
- Checks the active IDE’s session store for previous session data (
~/.claude/projects/for Claude Code,~/.codex/sessions/for Codex) - Finds when the planning files were last updated
- Extracts the conversation that happened after (potentially lost context)
- Shows a catchup report so you can sync
Pro tip: disable auto-compact to maximize context before clearing:
{ "autoCompact": false }
Maintainer depth (hook architecture, dispatcher layout, parity tooling) lives in AGENTS.md and docs/.
Commands
Slash commands ship with the Claude Code plugin route (see the install matrix above).
| Command | Autocomplete | What you get |
|---|---|---|
/planning-with-files:plan | type /plan | Creates the three planning files and starts the session (v2.11.0+) |
/planning-with-files:pwf | type /pwf | Short alias for /plan; --autonomous / --gated init (v3.0.0+) |
/planning-with-files:status | type /status | One-glance report: current phase and phase totals (v2.15.0+) |
/planning-with-files:plan-doctor | type /plan-doctor | Self-check for the failure modes that are silent by design: one PASS/WARN/FAIL line each for resolution, injection, attestation, install surfaces, and per-fire latency (v3.6.0+) |
/planning-with-files:plan-attest | type /plan-attest | Locks task_plan.md with a SHA-256; hooks refuse a tampered plan body; --show / --clear (v2.37.0+) |
/planning-with-files:plan-goal | type /plan-goal | Runs until the plan reports complete, composing with Claude Code /goal (v2.38.0+) |
/planning-with-files:plan-loop | type /plan-loop | Planning-aware cadence on /loop, default 10 minute tick (v2.38.0+) |
/planning-with-files:plan-de | type /plan-de | Start planning in German; also -ar, -es, -zh, -zht (v2.33.0+) |
/planning-with-files:start | type /planning | Original start command |
Typing /plan prefix-matches every plan* command in autocomplete; /planning-with-files:status autocompletes as /status (the older /plan:status label predates the rename).
Pi extension commands
Install the Pi extension with pi install npm:planning-with-files; it registers these commands, typed with no /planning-with-files: prefix.
| Command | What it does | Version |
|---|---|---|
/plan-execute | Pi only. Approve the active plan to ACTIVATE all Pi hooks; hooks stay passive until you run this; reset returns to passive review | v3.3.0+ |
/plan-status | Active plan path, scope, and phase totals | v2.39.0+ |
/plan-goal <text|default|clear> | Set or clear the goal string appended to auto-continue prompts | v2.39.0+ |
/plan-loop [interval] [prompt|stop] | Start or stop a planning tick (default 10m) that re-reads the plan and nudges progress | v2.39.0+ |
/plan-attest [--show|--clear] | Run the attest-plan helper; shares the .attestation file with Claude Code | v2.39.0+ |
On Pi there is no /plan command to create the files; the skill creates them, then /plan-execute approves and activates the hooks. Pi plan-goal/plan-loop run their own logic, while the Claude Code commands of the same name forward to native /goal and /loop. The doctor ships as a script in every mirror since v3.7.0: run sh scripts/plan-doctor.sh directly on platforms without the command.
Command names vs skill names
| Platform | You type | Examples |
|---|---|---|
| Claude Code | /planning-with-files:<verb>, autocompletes from the short form | /plan, /pwf, /plan-attest, /plan-de |
| Pi | bare form, no prefix | /plan-status, /plan-execute, /plan-goal |
| Continue.dev | /planning-with-files |
On the plugin route the model-invocable SKILL is planning-with-files:planning-with-files; the doubled form is the skill id, not a command you type. The five language variants live under skills/i18n/, which the plugin scan does not reach, so there is no planning-with-files:planning-with-files-de to invoke by name — reach a translation through its /plan-ar, /plan-de, /plan-es, /plan-zh or /plan-zht command, or install it as its own skill with npx skills add OthmanAdi/planning-with-files --skill planning-with-files-de -g, which registers it under its own name. There is no /pwf-de and no /planning-with-files:planning-with-files-goal; /pwf is just a short alias for /plan.
このREADMEは一部を省略しています。全文はGitHubにあります。 原文を読む ↗