
This past week I retooled my development environment after close to two years on Cursor. I wanted something that could coordinate agents and let me switch models per task, so I could maximize my tokens while minimizing my spend.
After a little research, I landed on a setup with two layers. Oh My Pi is the harness that runs my agents. Orca is where I run and interact with many of them at once.
Oh My Pi
Oh My Pi (OMP) lets you configure multiple models, assign them “roles”, and use them within sessions.
Model roles
OMP has several model roles for different tasks. Here are the ones I’ve defined, what each one does, and which model I use for it.
defaultruns normal session turns. I usedeepseek-flash:maxbecause most turns are implementation, and a cheap, fast model with max thinking handles them well.planruns plan mode. I useopus-5-5:highbecause planning is where mistakes are cheapest to fix.advisorwatches every turn and injects notes, like a code review that “shifts left.” I useopus-5-5:autohere. A stronger model reviews the cheap one’s work, and because it runs in its own context, it catches what the doer rushed past.taskcovers subagents spawned to do work.deepseek-flash:mediumhandles this delegated work. The plan already scopes it, so it doesn’t need a stronger model.smolcovers cheap subagent fan-out. I usedeepseek-flash:lowhere. These agents act as scouts that read code and report back, so speed and cost matter more than depth.commitwrites commit messages and splits commits. I runllama3.2:3blocally through Ollama. It’s a small task, it costs nothing, and my diffs never leave the machine.judgemakes typed yes/no and scoring decisions. I’m using the new hotness,jev-latest, withjev-previewas a fallback. It returns decisions instead of text and is cheap enough to run constantly.
BTW, if you’re just now hearing about Jev, this post gives a really good overview of the different ways agents can use it to make decisions.
Here’s a peek at my full model config:
setupVersion: 2
modelRoles:
commit: ollama/llama3.2:3b:auto
default: deepseek/deepseek-flash:max
advisor: anthropic/claude-opus-5-5:auto
plan: anthropic/claude-opus-5-5:high
task: deepseek/deepseek-flash:medium
smol: deepseek/deepseek-flash:low
judge: typesafe/jev-latest
display:
hideToolActivity: false
advisor:
enabled: true
task:
agentModelOverrides:
reviewer: anthropic/claude-sonnet-5-5:medium
retry:
fallbackChains:
judge:
- typesafe/jev-preview
Subagents
OMP also launches and monitors subagents of different types. Scouts investigate and gather details, while task agents execute the work the coordinator assigns them.

Prewalk
One feature I use a lot now is /prewalk. It has a stronger model plan the work and make the first edit, then hands the context to a cheaper agent. The first edit comes out higher quality and gives the cheaper agent a pattern to follow.
Orca
I used Cursor for quite a while, and I even have a free year of Cursor Pro from a promotion, but I really wanted to try something different.
In Orca, you add a repository and then create worktrees from it. Each worktree cuts a new branch and spawns your chosen coding agent in that directory, so several agents can run in parallel without stepping on each other’s toes.

You can also spawn a worktree straight from a GitHub issue or Linear ticket. From the CLI, run:
> orca worktree create --linear-issue ABC-123 --name fix-race-condition
Orca creates a new child worktree with the Linear ticket tracked to it.

Killer feature: mobile development

The Orca feature I love most is relay, which lets me run sessions from my phone. I scan a QR code, and then I can see my agents’ terminal output, send commands, and even spin up new work. Plenty of times I’ve pulled out my phone, requested work, and then heard the dings later as PRs opened and CI gates passed.
How it all adds up
On the Oh My Pi side, it’s clear how this setup gets the most out of every dollar I spend on AI. Expensive models plan the work and review it as an advisor, while cheap models execute it. Jev handles the decisions that would be wasteful to run through an LLM.
Orca doesn’t make a dent in the model bill, but it lets many worktrees run at once. Time that would otherwise sit idle goes to shipping features.