I plan before I execute. A stronger model writes the plan, and a cheap model like deepseek-flash carries it out. That works better than I expected, because the plan spells out things cheap models tend to skip, like writing the tests. But a plan can’t cover everything, and an executor deep in a task still lets things slip.
The obvious fix is a review pass after the work is done. That works, but it’s one more step to remember, and by then the problems are already in the code. I once had a task run for an hour. When I reviewed it, the agent had wandered far past what I asked for, and I only found out after the hour was gone.
What I wanted was a nudge while changing course is still cheap. In Oh My Pi, that’s what an advisor does. A second model watches the coding agent work and speaks up when something slips, with no separate review turn.
My last post gave the advisor role one line. It has earned more than that, and it has cost more than one line suggests. Advisors are an Oh My Pi feature. Upstream pi doesn’t ship them.
What an advisor is
An advisor is a separate agent with its own model, context, and isolated tool session. After each update from the primary agent, it receives only the new transcript delta: the agent’s reasoning, its tool calls, and the results, with secrets obfuscated. It also gets your project’s AGENTS.md context. By default it can investigate with read, grep, and glob.
It never approves anything and never changes the primary’s state directly. It talks back through one tool, advise, and every note carries a severity:
| Severity | What happens |
|---|---|
nit | A non-interrupting aside, delivered at the next turn boundary. |
concern | Delivered when the primary yields. Wakes it if it stopped mid-work. If it already gave its final answer, the note shows up as a card instead. |
blocker | Can interrupt work in progress, and can wake a turn that already ended with a final answer. |
OMP wraps each note in an advisory tag that tells the primary to weigh the advice and not blindly obey it. If you hit Esc on purpose, auto-resume stops and the note becomes a visible card instead. Plan mode keeps every steer as a card.
The built-in system prompt tells an advisor to stay quiet. It says “Silence preferred when agent on track.” An emission guard drops content-free notes like lgtm and nothing to add, along with duplicates, and advisor.immuneTurns, which defaults to 3, downgrades further concerns to asides for 3 turns after a note steers the primary. Blockers skip that cooldown.
It is not passive about drift, though. The same prompt tells the advisor to “Enforce user ask; flag drift immediately” and to advise before wrong-direction work. It objects when a bounded request gains features nobody asked for. A big diff alone is not a problem to it.
The idea isn’t unique to OMP. Anthropic’s advisor tool pairs a faster executor with a stronger advisor model, Claude Code exposes it as /advisor, and The advisor strategy explains the thinking behind it. The difference is who starts the conversation. There, the executor decides when to consult the advisor, so a model that might skip something has to know it needs help. In OMP the advisor watches every update and speaks up on its own.
Turning it on
You need an advisor role and the enable flag. This is straight from my config:
modelRoles:
advisor: anthropic/claude-opus-5-5:auto
advisor:
enabled: true
Point it at something stronger than whatever does your executing.
The :auto suffix makes the advisor follow the primary’s resolved thinking effort. With no suffix, it runs at medium. If you don’t set the role at all, it falls back to the slow role or a built-in strong-model chain.
Inside a session:
/advisortoggles it for the current session, and/advisor on|offsets it explicitly. Neither changes your saved config./advisor statusshows state, model, context size, tokens, and cost./advisor dumpcopies the advisor’s transcript to your clipboard./advisor dump rawadds the system prompt, tool calls, and thinking./advisor configureopens an editor forWATCHDOG.yml, which we’ll get to.omp -p --advisor "…"turns it on for a headless run.
If the reviewer’s provider fails, retry.fallbackChains.advisor lets you name backup reviewers.
What it caught
On September 30 I had deepseek-flash working through a build plan for a Django service, with Opus 5 as the advisor. I had approved the first wave of a multi-wave plan. Flash finished it, asked me whether to start wave 2, and stopped. Then OMP’s todo reminder fired, because wave 2 was still an open item:
You stopped with 1 incomplete todo item(s):
- Next
- Implement build wave 2: items 2, 3, 4, 6, 7, 8, 11, 12, 13, 23, 33, 41, 51, 56
Please continue working on these tasks or mark them complete if finished.
Flash took that as permission. It marked wave 2 in progress and started creating test databases for it. The advisor’s next review sent this:
<advisory severity="concern" guidance="weigh, don't blindly obey">
Your todo now has "Implement build wave 2" in_progress, but the user only ever selected wave 1 — starting 14 more items of production code unprompted is unrequested scope. Leave it pending and confirm before touching wave 2; the actionable state is that the plan set and wave 1 are both complete and awaiting a commit decision.
</advisory>
Flash’s reasoning shows it had been torn: “That’s a fair point and matches my own hesitation. The system reminder pushed me to continue, but the advisory argues the todo item is unrequested scope.” It replied:
The advisory is right and I’ll stop here rather than start wave 2 unprompted — you selected wave 1, and 14 more items is scope I shouldn’t infer.
Then it blocked the todo item with the reason “not started: you approved wave 1 only; say “go” and this starts immediately”. That’s the hour-long story from the top of this post, caught as it started.
It catches plain bugs too. In the same session it raised a blocker because a new hook used two names that didn’t exist: “tasks is never imported in ingest.py, and source isn’t a parameter of _persist (you removed it back in item 1).”
What it costs
Those catches aren’t cheap. From September 29 to October 7, across 190 sessions with the advisor on, the advisor cost $2,003.50 in real API spend. The primary agents in those same sessions cost $1,759.78 and their subagents $133.71. The reviewer was 51% of my bill.
The executor is not where that money goes. deepseek-flash ran 8,022 turns for $15.30. The wave 2 session above cost $36.90 for the primary and its subagents, and $307.86 for the advisor. Most of my non-advisor spend is plan mode and Sonnet sessions. My last post said this setup gets the most out of every dollar. For the executor that’s true. An Opus advisor watching every turn is the expensive part.
It’s also mostly quiet, and you pay for that silence. Of 13,956 reviews, 2,134 produced a note, about 15%. Every review still costs a model call.
The knobs that cut it, all per advisor in WATCHDOG.yml (below):
model: a Sonnet reviewer costs a fraction of Opus. My Rust demo reviewer cost three to five cents a run.reviewMode: agent-endreviews only when the agent yields, once per run. It would have seen wave 2 after the work, not as it started.reviewInterval: Nreviews every Nth update and batches the skipped ones into the next review./advisor offfor throwaway sessions.
Steering advisors with WATCHDOG.md
WATCHDOG.md is guidance for advisors only. OMP appends it to every advisor’s system prompt, including the ones you define in WATCHDOG.yml below, and the primary agent never sees it. That makes it the right place for review priorities you don’t want cluttering your AGENTS.md.
OMP loads every file it finds, in this order:
~/.omp/agent/WATCHDOG.mdWATCHDOG.mdand.omp/WATCHDOG.mdin each directory from your cwd up to the git root
It doesn’t stop at the nearest one. Narrower files go later in the prompt. Here’s what I’d put in a Rust project’s .omp/WATCHDOG.md:
- Watch for tests that mock the thing under test.
- Flag `unwrap()` and `expect()` in library code.
- Call out any "done" claim with no build or test output behind it.
- Ignore formatting and naming. The linter owns those.
Building your own advisors with WATCHDOG.yml
WATCHDOG.md shapes what advisors watch for. WATCHDOG.yml (or WATCHDOG.yaml) decides which advisors run. It uses the same search path, and each entry is a separate agent with its own focus:
| Field | What it does |
|---|---|
name | Required. Shows up in notes and status. |
model | Optional. Falls back to the advisor role if omitted. |
tools | Optional list. Defaults to read, grep, glob. [] grants none. |
instructions | Optional. Extra guidance for this advisor only. |
enabled | Optional. Set false to park an entry. |
maxNotesPerUpdate | Optional. Caps non-blocker notes per review. Defaults to 4. |
reviewMode | Optional. turn (default) reviews every turn. agent-end reviews only when the agent yields. |
reviewInterval | Optional. Review every Nth update. Defaults to 1. |
syncBacklog | Optional. How long the primary waits for this advisor to catch up: off, 1, 3, 5, or strict. Inherits advisor.syncBacklog, which defaults to off. |
The top level takes its own instructions, maxNotesPerUpdate, and the advisors list. The top-level instructions are shared by every advisor and concatenated across files. If two files define an advisor with the same name, the file closest to your working directory wins.
Here is the gotcha. Once a roster exists, the default advisor stops running. Only the entries you list run. To keep the generalist, add it back as an entry. With no model, it uses your advisor role:
advisors:
- name: General
Each entry can set its own model. A narrow focus doesn’t need your most expensive model, so a mid-tier reviewer can own one concern while a stronger advisor reviews the work as a whole. OMP writes each advisor’s transcript next to the session as __advisor.<slug>.jsonl.
You can grant an advisor any built-in tool, including edit, write, bash, and eval. They still honor approval mode, but I’d keep reviewers read-only until there’s a reason not to.
Giving an advisor skills
For the Rust reviewer I use two skills I already had installed: rust-best-practices, based on Apollo GraphQL’s Rust handbook, and rust-testing. They’re written-down standards a cheap executor may never open. Pointing the reviewer at them means it checks the work against specific rules instead of its own taste.
Loading the skills into the executor helps too, but it’s a different job. In my demo below, the executor had the skill installed and never opened it. The reviewer was told to read it, and did. The advisor’s one job is to check what just happened against the rules, with fresh context and a stronger model. A model that skipped something is unlikely to notice it skipped it.
There’s no skills: field in WATCHDOG.yml, but there are two ways to get the same effect.
Lazy skill:// reads
This is the option I recommend. The advisor shares your session’s skills, so it can read a skill:// URL on demand. It doesn’t know which skills exist, so your instructions have to name them:
instructions: |
Review for correctness and verification, not style.
advisors:
- name: General
- name: Rust Reviewer
model: anthropic/claude-sonnet-5-5:medium
instructions: |
You review Rust changes. Before your first advise on any Rust code,
read skill://rust-best-practices. When tests are added, changed, or
missing for new public functions, also read skill://rust-testing.
Cite the specific skill rule your note is based on.
The General entry keeps the default advisor running next to the Rust reviewer. Without it, only the Rust reviewer runs.
The skill only lands in the advisor’s context when it’s relevant. The last line matters most to me. A note that cites its rule is one I can check. I can look up the rule and tell a real catch from a false alarm, which matters when a stronger model is second-guessing a cheaper one.
@ imports
The alternative is to inline the skill with an @ import in the instructions:
advisors:
- name: Rust Reviewer
model: anthropic/claude-sonnet-5-5:medium
instructions: |
You review Rust changes against these standards.
@~/.agents/skills/rust-best-practices/SKILL.md
Two details to watch. A plain YAML value can’t start with @, so use a | block scalar. And an import written inside a fenced code block in the instructions text stays literal, so keep the path outside any fence. The import also inlines the whole file, frontmatter included, into every review prompt. For my skills that’s 4.4 KB for rust-best-practices and 11.8 KB for rust-testing.
Which one
Imports are deterministic, so the standards are in every review prompt. Lazy reads are cheaper and scale to a long list of skills, but two things can go wrong. The advisor has to follow the instruction to read them. And by default OMP evicts old read, grep, and glob results from the advisor’s context (advisor.evictStaleResults), so a skill read early on gets swapped for a placeholder a review or two later. In my demo runs the advisor read the skill once, on its first review, and never again. That’s fine for short tasks. For long sessions, tell the reviewer to re-read the skill before any note that cites it, or use an import. I use lazy reads and name the skills explicitly.
The Rust reviewer in action
I tested this in a throwaway git repo with only the Rust Reviewer entry from the WATCHDOG.yml above, no General. The primary was deepseek-v4-pro at xhigh, not my usual flash, and the reviewer was claude-sonnet-5-5 at medium. So this run shows the skill read and the note, not a cheap-versus-strong gap. Anthropic’s docs note that the benefit shrinks as the executor’s capability approaches the advisor’s. The wave 2 story above is the cheap-executor case.
My first two prompts dictated the implementation, with s.parse().unwrap(), no tests, and in one run an exit(1) on empty input. In both, the advisor read skill://rust-best-practices, then stayed silent through every step, reasoning that the user had asked for exactly that code. That matches the built-in prompt, which says “NEVER nitpick what user accepts.”
For the third run I gave only the signature:
Create a Cargo library crate in this directory named portcfg. Implement pub fn parse_port(s: &str) -> u16 in src/lib.rs that parses a TCP port number from a string. Keep it small and quick. Run cargo build, then reply done.
The primary wrote:
//! Minimal TCP port parsing.
/// Parses a TCP port number from a string.
///
/// Accepts only an unsigned decimal number in the range `1..=65535`;
/// anything else (empty, whitespace, sign, overflow, `0`) panics.
pub fn parse_port(s: &str) -> u16 {
let port: u16 = s.parse().expect("invalid TCP port");
assert!(port != 0, "invalid TCP port: 0");
port
}
As in the first two runs, the advisor’s first move was read skill://rust-best-practices.
This time it sent one note. Here it is as it landed in the primary’s transcript, with the note text exactly as the advisor wrote it in its advise call:
<advisory advisor="Rust Reviewer" severity="nit" guidance="weigh, don't blindly obey">
rust-best-practices (Error Handling): "avoid panic!/never expect() outside tests; return Result". The mandated `-> u16` forces a panic or sentinel here, so keep the `# Panics` doc. In your "done" reply, say that invalid input panics and that `Result<u16, _>` would be the idiomatic signature. Don't change the signature unasked.
</advisory>
It’s a nit, so it was never going to interrupt anything, and it also arrived too late. By default the primary doesn’t wait for the advisor (advisor.syncBacklog: off). The primary replied done., and the review of that last turn landed afterward. Print mode keeps late notes but never starts another primary turn, so the run just exited. The primary never saw the note, and its done. doesn’t mention panics. The note was also off on one detail. It said to keep the # Panics doc, but the doc comment only mentions panics in prose.
The quoted rule isn’t verbatim either. The skill has two bullets, “Return Result<T, E> for fallible operations; avoid panic! in production” and “Never use unwrap()/expect() outside tests”, and the advisor merged them. The citation still took me to the right section in seconds, which is why I ask for it.
If you want a reviewer that gates “done”, the docs describe a setup for it. I haven’t run it yet:
advisors:
- name: Rust Final Review
model: anthropic/claude-sonnet-5-5:medium
reviewMode: agent-end
syncBacklog: strict
instructions: |
Read skill://rust-best-practices, then review the finished run
against it. Use severity concern for any rule violation.
agent-end reviews once, when the agent yields. strict makes the primary wait for that review. An agent-end concern can ask for one more turn so the primary acts on it. A nit still won’t. In an interactive session, that’s the gate this demo was missing.
Advisors for subagents
Subagents run without an advisor by default. To attach one, set advisor: true in the agent’s frontmatter to use your advisor role’s model, or give it a model pattern:
---
name: rust-implementer
advisor: "anthropic/claude-sonnet-5-5:medium"
---
task.agentAdvisor overrides the frontmatter per agent. It takes "on", "off", or a model pattern, and you can edit it in the /agents hub.
The full reference lives in the Oh My Pi advisor and watchdog docs.