oh-my-pi on Claude Code
Proposal · harness engineering

oh-my-pi on Claude Code

What it takes to rebuild oh-my-pi's best features as Claude Code mods (TypeScript plugins that hook the engine from inside), and how natural each one feels. Fourteen of the twenty features land on hooks built for that job. Four need a workaround, one is a large build, and two are blocked today. A comparison of the two tools' subagents is near the end.

Status: this is the research written before the spike. What was built from it, and what the spike found, is in the claude-mods README; the mods there have names of their own rather than an omc prefix.

Based on the mod API types shipped with Claude Code 2.1.287 (the CLI installed here reports 2.1.288, so small drift is possible) and the oh-my-pi repository at 18.5.0 (your installed omp is 18.3.0). Every API name below is taken from those types. Claims that need a test are listed under Open questions.

oh-my-pi (omp)An open-source coding agent for the terminal, built on the pi agent. Its strength is the harness (everything around the model: which model runs when, what goes into context, which rules and reviewers watch the output). It runs models from many vendors.
Claude Code modsTypeScript plugins that run inside Claude Code and hook its engine: each model request, each tool call, the system prompt, compaction, the UI. A mod can watch an event, rewrite it, or answer it in the engine's place. This page asks how much of omp fits into mods.
NativeThe API has a hook built for this. Small code, no tricks.
Supported, heavyThe hook is documented, but you rebuild a whole subsystem on top of it.
WorkaroundIt works, but bends an API around a limit. Likely to break on upgrades.
BlockedThe API can't do it today.
Scope

The features, in plain words

What each feature does in omp, as this proposal understands it. If a description is off, the verdict is too, so this table is the first thing to agree on. Names link to the section that covers each one.

FeatureWhat it doesFitSize
Other vendors · gateway (B)Run Claude Code's main loop and subagents on non-Claude models (OpenAI, Google, OpenRouter, local), by pointing Claude Code at a proxy that translates between APIs. omp can log in with existing subscriptions such as ChatGPT or Gemini plans.NativeM
Other vendors · own client (A)Same goal, without a proxy: the mod itself sends each model request to the other vendor and feeds the answer back to Claude Code.HeavyL
Model rolesNamed slots, each holding a model: default, smol (fast and cheap), slow (strongest), tiny, task, advisor, judge, commit. Features ask for a role rather than a model, so swapping a model is one line of config.NativeS
Fallback chainsWhen a model errors or hits a usage limit, retry with the next model in a list. omp can also switch accounts or wait for the limit to reset.NativeS
Eval kernel + PTCA tool that runs JavaScript (or Python) in a long-lived process that keeps its variables between calls, like a notebook. Code in it can call the agent's other tools, models and subagents (PTC: programmatic tool calling), so the model writes one program instead of making twenty separate tool calls.WorkaroundL
PrewalkPlan with the strong model, then switch automatically to the cheap model once the first file edits begin.NativeS
JEV decisionsSmall yes/no or multiple-choice calls inside the harness, made by Jev (a model built for judgments, not chat). Examples: how much thinking this prompt needs, and whether the agent stopped before doing what it promised.NativeM
TTSR tool rulesTTSR (time-traveling stream rules) are rules written as Markdown files. A tool rule is a regex or code pattern checked against what the agent is about to write to a file. A match blocks the write and tells the model the rule.NativeS
TTSR question rulesA rule phrased as a yes/no question ("did the answer skip the tests?"). The judge checks every answer against it.NativeS
TTSR stream rulesA rule checked against the model's text while it is still being written. A match stops the answer and makes the model redo it with the rule in front of it.WorkaroundM
Compaction: shake · handoff · soft · idleWays to shrink a conversation that has grown too long. shake replaces old tool output with file references and needs no model call. handoff has the model write a hand-over note that replaces the history. soft is a normal summary. idle compacts while you're away.NativeM
Notes rolloverThe model keeps a small notebook. When the context fills, it starts a fresh window that carries only the notebook and its most recent work, with no summary.WorkaroundM
snapcompactOld history is printed onto images, and a vision model reads them back instead of the original text.BlockedM
Mid-turn compactionCompact while a long chain of tool calls is still running, not only between turns.BlockedS
Always-on advisorsA second model reviews the agent's work in the background as it happens. It interrupts only for a real problem (a blocker); smaller notes wait. omp can make the agent pause until the review catches up.NativeM
Hashline editAn edit tool where every edit cites a short tag taken when the file was last read. If the file changed since then, the edit is rejected instead of silently landing in the wrong place.NativeM
Commands and settingsSlash commands and settings that switch each feature on or off and configure it, plus one control pane that shows the state of all of them.NativeS
Structured subagent resultsThe parent gives a child a JSON schema; the child must return data that matches it, and is told to fix it until it does.NativeM
Agent hub paneA live list of running and finished child agents with their model, cost and current activity. From it you can steer any child.NativeM
Copy-on-write workspacesEach child works in an instant copy of the repo (a file-system clone, not a full copy), and its changes are merged back automatically.WorkaroundM
The whole picture

Where each feature plugs in

A Claude Code turn passes through a fixed set of engine events. A mod hooks an event with on(event, hook) and can watch it, rewrite it, or answer it instead of the engine. Each omp feature hooks one or two of them. The bottom row shows the services a mod calls through $; there is no Node inside a mod, so everything outside it goes through these.

JEV · pick effort Always-on rules Prewalk nudge Roles · model per step Advisor wait gate TTSR stream rules Other vendors · A TTSR tool rules Hashline edit Eval kernel + PTC Advisor notes in Advisor review JEV stop check TTSR question rules prompt.submitprompt typed prompt.composesystem prompt built turn.stepone model request, streamed tool.calleach tool use session.appendrow stored turn.completeturn ends tool result → next step (the agent loop) context full session.compacthistory rebuilt shake · soft · idle notes rollover snapcompact handoff services a mod calls through $ $.clock $.http $.process $.model.fork $.agent.spawn timers: idle compaction other vendors, Jev judge eval kernel child process side question on cached history advisors, role subagents
Most features cluster on turn.step (one model request, streamed) and turn.complete. Dashed lines show which service a hook relies on: $.http reaches other vendors, $.process runs the eval kernel, $.model.fork writes the handoff on cached history.
Natural or hacking?

Fit against size

Columns are how natural the fit is. Rows are rough size: S is a single hook and a few hundred lines at most, M is a small mod with its own state, L is a subsystem. Top-left is cheap and safe. Right-hand columns are where upgrades will hurt.

Native
Supported, heavy
Workaround
Blocked
S
RolesPrewalkTTSR tool rulesTTSR question rulesFallback for side callsCommands and settings
Mid-turn compaction
M
Gateway vendors (B)AdvisorsJEV decisionsshake · handoff · soft · idleHashline editStructured subagent resultsAgent hub pane
TTSR stream rulesNotes rolloverCopy-on-write workspaces
snapcompact
L
Own vendor client (A)
Eval kernel + PTC
The short answer
  • Feels designed for it: roles, prewalk, advisors, JEV decisions, tool-level TTSR, and the compaction methods that return text. The API docs literally describe these moves: rewrite model per step, answer session.compact with your own messages, append a note into a running turn.
  • Feels like hacking: the eval kernel's transport, TTSR on streamed prose, and notes rollover. Each one routes around a stated limit: a child process's stdin can only be written once, streamed text can't be taken back, and compaction can't run mid-turn.
  • Supported but expensive: running the main loop on another vendor's model. The API explicitly lets a hook answer a model request itself, so it isn't a hack. The cost is that you rewrite a provider layer.
1 · Models

Other vendors, roles and fallback

omp talks to 88 providers and assigns models to roles (smol, slow, tiny, task, advisor, judge and more). Claude Code's own client speaks only the Anthropic Messages API, and a mod can't change its base URL. That leaves two routes, and the choice between them shapes everything else.

A · omc answers every step itself B · the engine's client points at a gateway Claude Code agent loop turn.step hook (omc) engine's Anthropic client next() never called: no request translate history + toolsschemas shipped with the mod curl -N via $.process$.http.fetch can't stream OpenAI · Gemini · OpenRouter chunks yielded back Claude Code agent loop omc rolessmol → gpt-5.6-luna model name per step engine's Anthropic client ANTHROPIC_BASE_URL gatewayomc router + omp auth-gateway OpenAI · Gemini · OpenRouter translation, OAuth and fallback live in the gateway, not in the mod
In A the mod owns the whole provider layer: it reads history with $.session.messages({as:'api'}), sends it out, and yields text, thinking and tool-call chunks that the engine trusts as the model's answer. In B the mod only picks model names; the gateway does the translation.

Other vendors via gateway (B)

Native · M
In omp
OAuth subscriptions (Codex, Antigravity, Copilot…), 88 providers.
Here
A fixed ANTHROPIC_BASE_URL to a local router that launchd keeps running; the router sends Claude traffic straight through and everything else to omp auth-gateway. The mod installs it, health-checks it and picks the model per step.

Feel: mostly configuration plus a small helper. Main loop, subagents and side calls all switch together, using your omp logins. The one rough edge: switching it on or off applies from the next session start. See Route B without the user noticing.

Own vendor client (A)

Supported, heavy · L
In omp
Native wire clients per provider.
Here
turn.step answers without next. Streams via curl -N, translates the vendor's event stream (SSE) into Claude Code's chunks (TurnStepChunk).

Feel: the hook is documented for exactly this ("an answer without next sends no request"). The cost is everything around it: tool input schemas aren't readable at runtime, so you ship them; thinking you invent is shown but not recorded; each provider's quirks are yours.

Model roles

Native · S
In omp
modelRoles.*, picked per call site.
Here
Role table in userConfig (shows in /config). next({...e, model}) in turn.step; $.agent.register for scout/reviewer/task agents.

Feel: natural. The model and effort fields on each step are rewritable by design, main thread and subagents alike. Shipped as a $.roles noun other mods depend on.

Fallback chains

Native · S
In omp
retry.fallbackChains, usage-aware account switching.
Here
For calls the mod makes (judge, advisor): retry down the chain on isAnswered:false. Main loop: only through A or a gateway.

Feel: natural for side calls; every $.model.complete resolves with a reason instead of throwing. Usage-aware switching belongs in the gateway.

Roleomp uses it forWhere it hooks in Claude Code
defaultmain session modelthe session's model
smolprewalk target, scout agent, edit auto-repairturn.step model rewrite; scout agent via $.agent.register
slowreviewer agent, advisor fallbackreviewer agent via $.agent.register
taskgeneric task subagentagent type with model set
tinysession titles$.model.complete
advisoralways-on revieweradvisor mod (section 5)
judgetyped decisions (Jev)$.http.fetch to OpenRouter's decisions API
commitomp commit messages/commit command answered by command.run

A role can name a non-Anthropic model only once route A or B works. The judge and advisor are the exception: the mod calls them directly over $.http.

Route B without the user noticing

omp already ships the gateway. omp auth-gateway serve answers the Anthropic Messages format on /v1/messages, plus OpenAI Chat Completions, Responses and Jev decisions. It translates each request for the target vendor and signs it with your omp logins. It gets those logins from a second omp process, omp auth-broker serve, which holds the tokens and refreshes them. Nobody needs to write a translation layer.

A mod can't change Claude Code's base URL while a session runs, and a mod can't listen on a port. So the design fixes the address once and keeps a small helper always running, with the mod in charge of both:

Claude Code sessionany number of them omc modruns inside every session ~/.claude/settings.jsonenv.ANTHROPIC_BASE_URL omc router127.0.0.1:4141 · small Bun script launchd user agentstarts at login, restarts on crash api.anthropic.com omp auth-gatewaytranslates wire formats omp auth-brokerlogins, token refresh OpenAI · GeminiOpenRouter · Jev every model request claude-*: as is other models credentials health check, restart,push role → model map installs once writes once keeps it running
The base URL in settings.json never changes; the router behind it decides where each request goes. Claude models skip the translation, so Claude Code's caching, thinking and beta features reach Anthropic untouched.

What the mod does

  • /omc gateway setup, once. It writes a launchd agent (macOS's per-user service manager) for the router and the two omp processes, starts it, then asks before adding ANTHROPIC_BASE_URL to settings.json. It asks because that one line affects every future Claude Code session, with or without the mod.
  • At every session start. It checks the router's health before the first prompt (session.start runs before the first prompt is sent). If the router is down it restarts it with launchctl kickstart, and if that fails it says so in a toast.
  • On every model request. It rewrites the model name to the role's real model, e.g. openai-codex/gpt-5.6-luna for smol. The router reads the name and routes the request.
  • On every role change. It pushes the new role map to the router over $.http, so the router can also resolve @smol-style names.
  • Status and control. The status line shows "gateway ok". /omc gateway status|restart|logs|off. off removes the settings line, and the change applies from the next start.

Why a router in front of omp's gateway

  • omp's gateway has no raw passthrough. Even Claude requests are parsed and re-encoded. That risks dropping things Claude Code relies on: prompt-cache markers, signed thinking blocks, beta headers. The router sends claude-* traffic to Anthropic byte for byte.
  • The HTTP gateway needs the broker. omp auth-gateway stdio runs without a broker, but it returns a streamed answer in one piece at the end. That's too slow for an interactive session.
  • One owner for login tokens. If the broker and your everyday omp both refresh the same OAuth tokens, a refresh by one can invalidate the other's token. Pointing omp at the broker too (auth.broker.url) leaves the broker as the only one refreshing.

What can still break

  • Model names. Claude Code checks model names against an allowlist. Whether it accepts openai-codex/… behind a custom base URL needs a test. If not, the router maps alias names instead.
  • Mixed history. After prewalk switches models mid-session, the history holds Claude's signed thinking blocks and another vendor's reasoning. Both directions need a test.
  • Context size. Claude Code doesn't know a foreign model's context window, so its auto-compaction threshold may be wrong for it.
  • Sign-in. With a custom base URL, Claude requests go out with whatever credential Claude Code sends. Which header that is under a subscription login needs a test.
2 · Code execution

The eval kernel with full programmatic tool calling

omp's eval tool runs a cell in a JavaScript (Bun) or Python kernel that keeps its variables between calls. Inside a cell, await tool.grep(...) calls any session tool, and completion(), agent(), judge() and workpool() reach models and subagents. This is programmatic tool calling (PTC): the model writes a program that calls tools, instead of making one tool call per turn.

In Claude Code the tool side is natural: $.tool.register adds eval, and $.tool.call runs any built-in or MCP tool with the normal permission check. The kernel is the awkward part. $.process.spawn writes stdin once and closes it, so you can't keep a REPL open over stdin. The kernel serves HTTP on a Unix socket instead, and the mod talks to it with $.http.fetch({socketPath}).

Model mcp__omc__eval Bun kernelchild process, Unix socket Claude Code toolsbuilt-in + MCP eval { code } POST /run paused: tool.Grep(args) $.tool.call · permission check result POST /resume { result } done { value, logs } tool result variables, imports, open handlesstay until the mod reloads steps 3–6 repeat per tool call
The kernel can't reach $, so every tool call inside a cell travels back through the mod. That's the right shape anyway: permissions, hooks and transcripts stay in force. Waiting inside $ calls doesn't count against the mod's 10-second budget for its own code, so long cells are fine.

Eval kernel + PTC

Workaround · L
In omp
Retained Bun VM / IPython kernel, tool.*, completion, agent, workpool, judge.
Here
$.tool.register + $.tool.call (native); kernel over a Unix socket (workaround); agent() via $.agent.spawn; completion() via $.model.complete.

Feel: the core is natural; the transport is the hack. The kernel restarts whenever the mod reloads after an edit, which only happens while developing it. Open questions: whether a plugin tool can return an image (for display()), and Python later or not at all.

3 · Harness decisions

Prewalk and Jev-judged decisions

Prewalk plans on a strong model, then switches to a cheap one for the edits. In the mod it's a small state machine kept in $.state (session values that survive a hot reload): turn.complete moves it forward and turn.step rewrites the model once it flips.

Armedplan nudge in prompt.compose Gate openstill on @default Switchedmodel → @smol, checklist steer Doneroles untouched todo succeeds turn with edit/writeends disarm
Every transition is something a hook already sees: a tool result, a turn ending, the next model request.

Jev is TypeSafe's judgment model (served through OpenRouter's decisions API). omp asks it typed questions instead of using a chat model: which effort level a prompt deserves, whether a turn promised an action and then stopped, whether an answer breaks a question-style rule. Each one maps to a hook:

  • Pick effort: at prompt.submit, one choice question over the prompt; then turn.step rewrites effort.
  • Unexpected stop: at turn.complete, a yes/no question on the last answer; if yes, $.prompt.submit queues "continue". It arrives as a new turn, not a continuation of the old one.
  • Question rules: at turn.complete, all candidate questions in one judge request (covered under TTSR).

Prewalk

Native · S
In omp
Plan on default, switch to smol after the first edit turn.
Here
$.state gate + turn.complete + turn.step model rewrite; /prewalk command.

Feel: natural, maybe 100 lines. For subagents it depends on whether spawned subagents pass through the same hook (open question).

JEV decisions

Native · M
In omp
Auto thinking level, unexpected-stop detection, semantic find.
Here
$.http.fetch to the decisions API; effort rewrite; $.prompt.submit to resume.

Feel: natural. A small typed client plus three hooks. $.model.classify is a built-in fallback that picks one label using a small Claude model.

4 · Rules

TTSR: rules that fire while the model writes

omp's time-traveling stream rules watch the model's output as it streams. When a regex, an ast-grep pattern or a judge question matches, omp aborts the response, throws the partial output away, injects the rule as a hidden interrupt and lets the model try again. Rules are Markdown files with a frontmatter condition, scope and globs. omp's rule format can be reused as is.

Three kinds of rule, three different fits:

  • Tool rules (on edit or write arguments): tool.call returns {deny} with the rule body before anything touches disk. Natural, and arguably better than omp's, because the deny happens before any side effect.
  • Question rules: one judge request at turn.complete, and the verdict comes back as a note. Natural.
  • Stream rules (on prose or thinking): a workaround, shown below.
omp Claude Code (omc mod) partial output dropped (contextMode: discard) regex hits <system-interrupt> + rule model tries again already yielded: shown and kept in history regex hits yield tool chunkrule_interrupt tool result = rulenext step reads it
Two limits force the detour. A chunk a hook has yielded stays in the response, so text already shown can't be taken back. And a retried request carries the same messages, so the rule can't be injected into it. To drop the text as omp does, the mod has to buffer the stream before showing it, which costs latency.

TTSR tool rules

Native · S
In omp
Regex/AST on edit and write arguments, blocks before side effects.
Here
tool.call deny with the rule body; ast-grep via $.process.run.

Feel: natural. Same rule files, discovered from .omp/rules and .claude/rules.

TTSR question rules

Native · S
In omp
Judge checks yes/no questions after each output.
Here
turn.complete → one judge batch → $.session.append note.

Feel: natural. Depends on the JEV client.

TTSR stream rules

Workaround · M
In omp
Abort, discard, inject interrupt, retry.
Here
turn.step generator watches chunks; on a hit, stop forwarding and yield a synthetic call to rule_interrupt.

Feel: hacky. It fakes a tool call to get a message in. Untested: whether a hook can drop next mid-stream and yield its own tool call.

5 · Context

Custom compaction

Here Claude Code is more open than expected. When the engine compacts, the session.compact event lets a hook return any list of messages instead of the engine's summary. omp walks a method order and falls through on failure. The same chain fits inside one hook.

context fullor /compact, or idle timer session.compactomc hook, methodOrder 123 shakeno model: old results → files on disk handoff$.model.fork writes a doc, cache reused softnext(e): the engine's own summary snapcompactneeds images back in context { messages }first that succeeds
Answering session.compact without calling next replaces the engine's algorithm. Returned messages can be plain text and tool blocks. That rules out snapcompact (old history printed onto images for a vision model to read back): a mod can append and return text only.

The catch is timing. $.session.compact rejects while a turn is running. omp's notes-backed context (a context_notes tool plus new_context to roll over to a fresh window) wants to roll over in the middle of a tool loop. In Claude Code the mod has to do it in four steps:

  1. The model calls new_context.
  2. The tool replies "rolling over" and the turn ends.
  3. A turn.complete hook compacts down to the notebook and the recent tool calls.
  4. $.prompt.submit sends "continue", which starts a new turn.

shake · handoff · soft · idle

Native · M
In omp
compaction.methodOrder, idle compaction, /handoff, /shake.
Here
session.compact answer; $.model.fork for handoff; $.clock.after for idle; slash commands.

Feel: natural. The event is built for replacement, including vetoing with {skip}.

Notes rollover

Workaround · M
In omp
context_notes notebook, new_context at a safe tool boundary.
Here
End turn → compact in turn.complete → $.prompt.submit.

Feel: hacky. A rollover becomes a visible turn break. Untested: whether compaction is allowed from inside turn.complete.

snapcompact

Blocked
In omp
History printed onto PNG frames, read back by a vision model.
Here
Compact answers and plugin appends are text only.

Feel: not possible unless a plugin tool result can carry an image (one test settles it).

Mid-turn compaction

Blocked
In omp
Compacts inside a long tool loop.
Here
$.session.compact rejects while a turn runs.

Feel: blocked. Claude Code's own auto-compaction still runs; the mod's methods apply when it does.

6 · Oversight

Advisors that are always on

omp's advisor reviews each new stretch of the transcript as the primary works, in its own context, and speaks only through an advise tool with a severity: nit, concern or blocker. Blockers steer the live turn. With syncBacklog: 1 (your setting), the primary waits up to 30 seconds whenever a review is pending.

Every piece has a direct hook here. The review runs in the background with $.agent.spawn (or over $.http for a non-Claude advisor). $.session.append adds a note that a running turn reads at its next step. The wait is a turn.step hook that awaits the review before letting the next request through, and waiting inside $ calls doesn't count against the 10-second budget.

main advisor step 1 step 2 waits: syncBacklog = 1 (≤ 30 s) step 3 reads note step 4 review Δ1 review Δ2 → blocker review Δ3 Δ Δ $.session.append Δ
The advisor sees only the delta (Δ) since its last review, in its own context. Nits and concerns can be held back and shown as toasts or in a side pane, so only blockers interrupt the main thread.

Always-on advisors

Native · M
In omp
advisor.*, WATCHDOG.yml roster, severities, dedupe, rate limits.
Here
turn.step / turn.complete → background review → $.session.append; roster file read with $.fs; a pane listing notes.

Feel: the most natural fit on the list. The only rough edge: $.session.append takes text blocks only, which is all a note needs.

Hashline edit

Native · M
In omp
[PATH#TAG] edits checked against a tag taken when the file was last read; stale edits rejected.
Here
$.tool.register an edit tool; tool.call deny on the built-in Edit; tags in $.state.

Feel: natural. You lose the built-in diff view unless you redraw it with a ui.render hook on ToolResult.

7 · Control

Commands and settings to drive all of this

omp puts each feature behind a slash command and a settings key (/advisor on, /prewalk restart, /handoff, /omfg, /model → Roles). Mods get the same three controls, and all of them are native:

  • Slash commands: $.command.register at session start, answered by a command.run hook. A command marked immediate runs even while a turn is in progress.
  • Settings: fields declared in the mod's manifest (userConfig) appear as rows in Claude Code's own /config menu. A field with a fixed list of values becomes a picker. Changing one reloads the mod with the new value.
  • A control pane: $.ui.open opens a side pane with buttons and pickers, redrawn from $.state whenever a value changes. Status-line text and toasts show what is running.
claude — ~/code/app … transcript … ● Edit src/auth.ts ⎿ advisor (concern): token TTL never checked > /omc omc ROLES smolgpt-5.6-luna ▾ slowclaude-fable-5-1 ▾ judgejev-1.13 ▾ FEATURES prewalkarmed advisoron · 1 ttsr12 rules COMPACTION ORDER [shake] → [handoff] → [soft] fable-5-1 · advisor reviewing Δ3 · backlog 1 · ttsr ok
A mock of how it could look, not a screenshot. The pane is drawn by a ui.render hook on Pane. The bottom line is $.ui.status. Every value shown also lives in /config, so it can be scripted.
CommandWhat it does
/omcOpens the control pane above: roles, feature switches, compaction order.
/roles [role model]Shows or sets role assignments.
/prewalk [restart]Arms prewalk, or switches back to default and arms it again.
/advisor on|off|status|dumpControls the advisor; dump shows its notes so far.
/ttsr list|test|scanLists rules, tests one against a text, or scans recent history for hits.
/omfg <complaint>Drafts a new rule from a complaint ("stop adding try/catch everywhere") and checks it against recent history, as omp does.
/handoff, /shakeRuns that compaction method now (only between turns).

Commands and settings

Native · S
In omp
Slash commands per feature, config.yml keys, /model Roles view, /advisor configure.
Here
$.command.register, userConfig rows in /config, a control pane, $.ui.status.

Feel: natural. Each mod registers its own commands, and the roles mod adds the shared /omc pane.

8 · Subagents

Subagents: omp against Claude Code

Both tools let the main agent hand work to child agents with their own prompt, tools and model, running in parallel or in the background. omp goes further on orchestration: structured results, a model from any vendor, workspaces that merge back automatically, nesting, and live control. Claude Code is stronger on context and safety: a child can be forked with the whole conversation, can run in the cloud, and stays under your permission settings.

CapabilityompClaude CodeEdgeCan a mod close the gap?
Model per childAny vendor; an ordered list with role aliases (@smol); tag a model in the prompt with ^Claude models (others only through a gateway)ompThrough route B, or route A if test 2 passes
Structured resultoutputSchema per spawn, validated with retries; child must finish through a yield toolFree text from the Agent tool; schemas only inside workflow scriptsompYes: a yield plugin tool that validates the schema in tool.call
Shared contextChild starts fresh; a batch shares one context blockFresh child, or a fork that inherits the whole conversationClauden/a
Isolated workspaceCopy-on-write clone (APFS, btrfs, overlayfs…); changes merged back as a patch or a branch automaticallyGit worktree (left for you to merge), or a remote cloud sandboxomp locally; Claude for cloudPartly: clone with $.process, spawn with cwd, apply the diff afterwards
NestingChildren spawn children (depth 2 by default), with a per-agent spawns allow-listNot confirmed for subagents; workflows handle multi-level fan-outompUnclear; needs a check
Watch and steerAgent Hub (Alt+A): live roster, cost, tokens, open a child's transcript and type into itBackground tasks list; SendMessage to continue or steer an agentompYes: a pane over $.agent.list(), steering with $.session.append({agentId})
Revive laterIdle children parked after 7 min, revived from their saved transcriptSendMessage resumes a finished agent with its contextEvenn/a
Per-child harnessPrewalk, advisor and thinking level per agent typeModel, effort, tools, hooks, MCP servers, permission mode per agent typeEven, different knobsPrewalk and advisor per child, if test 2 passes
Code-driven fan-outagent() and workpool() from the eval kernel; children can call functions defined in the parent kernelWorkflow scripts (parallel, pipeline, schemas)omp for kernel callbacksYes, inside the eval mod
Start-up speedChildren start while the spawn call is still streamingChildren start once the call is completeompNo
PermissionsChildren always run with approvals off (no screen to ask on)Children follow the session's permission modeClauden/a

Structured subagent results

Native · M
In omp
outputSchema per spawn; child ends with yield, validated, up to 3 reminders.
Here
$.agent.register agents that must call a yield plugin tool; tool.call validates against the schema and replies with errors until it passes.

Feel: natural. Only the per-spawn schema is awkward: it travels in the task text, because the Agent tool has no schema field.

Agent hub pane

Native · M
In omp
Live roster with cost and tokens, open and steer any child.
Here
Pane over $.agent.list() and turn.complete usage; steer with $.session.append({agentId}).

Feel: natural for the roster and steering. Opening a child's full transcript inside the pane would mean drawing it yourself.

Copy-on-write workspaces

Workaround · M
In omp
APFS/btrfs clone per child, auto-merged as a patch or a branch.
Here
cp -c clone via $.process, $.agent.spawn({cwd}), then git diff and apply at turn.complete.

Feel: hacky next to Claude Code's built-in worktrees. Worth it only if worktrees are too slow for your repos.

Skip

Already in Claude Code

Rebuilding these would duplicate what's there:

  • Subagents with their own model, tools and prompt (agent files, and $.agent.register for roles).
  • Steering: messages typed during a turn are delivered into it.
  • Rewind to an earlier point, and /compact with instructions.
  • Skills, MCP servers, plan mode, background shells, worktrees.
  • Dynamic workflows (multi-agent scripts), which cover much of omp's workpool use outside a kernel.
Hard limits

What a mod can't do

  • Change the session client's provider or base URL. Hence route A or a gateway.
  • Read tool input schemas at runtime. Route A has to ship them.
  • Stream with $.http.fetch. Streaming needs a spawned curl.
  • Talk to a child process over stdin across calls, or keep one alive through a hot reload.
  • Take back streamed text, or inject into a retried request.
  • Edit past rows except through a compaction answer, or compact mid-turn.
  • Run its own code for more than 10 seconds per dispatch. Waiting inside $ calls is free.
Before building

Open questions, each settled by one test

Each one is a single mod written as a test file run with claude plugin test, or a single curl call. Together they decide eight of the verdicts above.

QuestionHow to settle itDecides
Does Claude Code work end to end through omp auth-gateway? (Its docs list a /v1/messages route, so the format is covered.)Run broker and gateway locally, point ANTHROPIC_BASE_URL at them, run one claude -p with a non-Claude modelRoute B
Does Claude Code accept a non-Claude model name behind a custom base URL, and which auth header does it send?Same run, with the router logging the requestRouter naming, Claude passthrough
Does a history mixing Claude thinking blocks and another vendor's reasoning survive a model switch both ways?Prewalk from Claude to gpt-5.6-luna and back in one sessionPrewalk across vendors
Can a plugin tool result include an image?A test tool returning an image blocksnapcompact, eval display()
Does a subagent from $.agent.spawn pass through the same plugin's turn.step?Spawn one and log e.agentId in the hookOther-vendor subagents, prewalk in subagents
Can a turn.step hook stop reading next mid-stream and yield its own tool call?Test with a regex on the word "TODO"TTSR stream rules
Is $.session.compact allowed from a turn.complete hook?Call it there and check the resultNotes rollover
Can a Claude Code subagent spawn its own subagents?Give a test agent the Agent tool and ask it to delegateNesting row in the subagent table
Proposal

Order of work

Small mods, not one big one. A roles mod adds a $.roles noun to $ (published as a typed contract), and the others list it as a dependency. Each mod reloads on its own while being built.

Wave 0

Tests

The eight open questions. Their answers change the verdicts before any real code is written.

Wave 1

Natural fits

Roles + gateway (if the test passes), commands and the control pane, prewalk, advisors, JEV decisions, TTSR tool and question rules, shake · handoff · soft · idle.

Wave 2

Workarounds worth it

Eval kernel + PTC, structured subagent results, agent hub pane, TTSR stream rules, notes rollover, hashline edit.

Wave 3

Only if the gateway falls short

Own vendor client on turn.step (route A).