On This Page8 sections
Key Takeaways
- Oh My Pi v18.4.4 is the latest release as of September 30, 2026, continuing one of the fastest release streaks in the terminal coding-agent ecosystem.
- The update adds GPT-6.1 Sol discovery and support, including extended-context handling, cost metadata, and OpenAI service-tier integration.
- OMP now exposes Ultrafast routing through
/fast ultra, extending its model router into a combined model, role, reasoning, and service-tier router. - Claude Sonnet 5.5 support has been strengthened with between-tools thinking and related Anthropic/Bedrock compatibility work.
- Recent releases also added Command Code Jev as a judge model, speculative subagent execution, live steering, predictive terminal text, prompt-cache warming, and a redesigned statistics system.
- The larger story is not a single model integration. Oh My Pi is evolving from a feature-rich Pi fork into a programmable AI coding workstation and agent runtime for the terminal.
What Changed in Oh My Pi v18.4.4?
Oh My Pi, commonly shortened to OMP, has been shipping releases at an unusually aggressive pace. The v18.4.4 release is important because it expands three areas at once: frontier-model support, inference routing, and agent-loop behavior.
The headline addition is support for GPT-6.1 Sol through the OpenAI Codex ecosystem.
OMP now recognizes GPT-6.1 Sol in Codex model discovery and carries model-specific metadata such as pricing, context limits, and service-tier support. That sounds like a routine catalog update, but the implementation reveals something more important about the project.
OMP is increasingly integrating with the behavior of provider-specific coding systems rather than treating every model as a generic OpenAI-compatible endpoint.
For GPT-6.1 Sol, the project had to account for Codex discovery behavior tied to the reported Codex CLI client version. OMP updated its discovery request behavior so the model would be exposed correctly by the backend.
That is a small implementation detail with a large product implication: OMP is becoming a compatibility layer across coding-agent ecosystems, not merely a model switcher.
GPT-6.1 Sol Comes to OMP
The GPT-6.1 Sol integration is the most searchable update in v18.4.4.
OMP's model metadata includes support for:
- GPT-6.1 Sol
- GPT-6.1 Sol WM
- Codex model discovery
- Service-tier selection
- Extended context
- Cost calculation
- Ultrafast mode
The published catalog metadata for GPT-6.1 Sol lists approximately:
- $2 per 1M input tokens
- $10 per 1M output tokens
- $0.10 per 1M cached input tokens
- 272K default context
- Up to roughly 922K extended input context
The important part for developers is not simply that another OpenAI model appears in a dropdown.
OMP can route GPT-6.1 Sol into a broader orchestration system containing model roles, subagents, advisers, judges, thinking levels, and service tiers.
A task can therefore be treated as more than:
prompt -> model -> answer
It can become:
task
↓
agent role
↓
provider
↓
model
↓
reasoning level
↓
service tier
↓
tools and subagentsThat architecture is one of the clearest ways OMP is differentiating itself from simpler terminal assistants.
/fast ultra Turns Speed Into a Routing Decision
OMP v18.4.4 also adds support for an OpenAI Ultrafast service tier.
Users can select it with:
/fast ultraThe tier is also wired into the broader routing system, including OpenAI configuration, subagents, advisers, and command-line service-tier selection.
This matters because OMP's routing problem is no longer only:
Which model should handle this task?
It is becoming:
Which model, provider, role, reasoning effort, and latency tier should handle this task?
For interactive coding, that distinction can be significant.
A developer may prefer a fast response for repository search, a stronger reasoning configuration for architecture work, and a lower-cost worker for parallel subagent tasks. A coding harness that can express those choices explicitly has more room to optimize for both responsiveness and cost.
Claude Sonnet 5.5 Gets Better Agent-Loop Support
The same release line also expands support for Claude Sonnet 5.5.
One of the most technically interesting additions is a capability flag for between-tools thinking.
A simplistic agent loop often behaves like this:
reason
↓
tool
↓
tool
↓
tool
↓
final answerBetween-tools thinking enables a more adaptive pattern:
reason
↓
tool
↓
reason again
↓
tool
↓
reason again
↓
tool
↓
final answerThat difference matters most in workflows where new evidence appears after every tool call.
Examples include:
- Debugging an intermittent test failure
- Exploring an unfamiliar repository
- Inspecting runtime state
- Refactoring across package boundaries
- Investigating dependency behavior
- Comparing multiple implementation strategies
The agent can revise its working hypothesis as new information arrives instead of relying on a plan made before the first tool result exists.
Recent OMP releases also include broader Sonnet 5.5 and Anthropic compatibility work, including multimodal handling and Bedrock-related improvements.
Jev Becomes a Judge Model
Another notable recent change is the addition of Command Code's Jev decision model for the judge role.
This is more strategically important than adding another general-purpose coding model.
OMP already separates model responsibilities into roles such as:
defaultsmolslowplantaskadvisorvisioncommitjudge
Jev fits naturally into this architecture because it is designed around lightweight decision tasks rather than long-form generation.
The agent runtime can therefore move toward a pattern where expensive frontier models are reserved for work that needs them, while specialized decision models handle narrow classification or judging tasks.
Conceptually:
GPT-6.1 Sol / Claude
↓
main implementation and reasoning
Jev
↓
judge, classify, route, decide
smaller worker model
↓
search, narrow subtasks, mechanical workThis is one of the strongest arguments for treating OMP as an agent orchestration system rather than a conventional CLI chatbot.
Speculative Subagents Reduce Agent Latency
OMP has also been working on speculative execution.
Recent agent-core changes introduced authorization hooks that allow a host to begin approved side-effect work, including certain subagent launches, before the full tool-call lifecycle has completed.
The traditional sequence is roughly:
model finishes tool call
↓
parser validates request
↓
tool dispatcher starts
↓
subagent begins workA speculative workflow can begin moving toward:
model output becomes sufficiently certain
↓
host authorizes the operation
↓
subagent starts early
↓
formal tool call completesThis is a form of latency hiding.
The benefit becomes larger as agent workflows contain more parallel workers. Saving a small amount of startup latency on a single tool call may not matter much, but saving it repeatedly across a large multi-agent task can materially improve perceived responsiveness.
Live Steering Makes Long-Running Agents More Interactive
OMP's recent releases have also strengthened live steering and message queues.
Instead of forcing users into a strict cycle of:
prompt -> wait -> result
the runtime increasingly supports something closer to:
agent working
├── user steering
├── queued instruction
├── follow-up request
├── background result
└── subagent resultQueue-management APIs have been expanded, and users can steer an active session while generation is still in progress.
This matters for real software work because developers often discover new requirements while an agent is already investigating the repository.
For example:
Do not change the database schema.
or:
The failing test is only reproducible on Windows.
or:
Keep the public API backward compatible.
A coding agent that can absorb that information without discarding the entire active task is more useful for long-running engineering workflows.
Auto Thinking Now Focuses on Solution-Space Openness
OMP also changed how automatic reasoning effort is conceptualized.
A previous task-complexity concept was replaced by solutionSpace, shifting attention from the amount of work to how open-ended the problem is.
That is an important distinction.
A task can be large but straightforward:
Rename this API across 120 files.
The workload is substantial, but the solution space is narrow.
Another task may involve only a small patch:
Find the cause of this intermittent distributed-cache consistency bug.
The final code change may be ten lines, but the space of plausible causes is large.
Reasoning effort should therefore track uncertainty and openness more closely than raw task size.
OMP's direction can be summarized as:
solution-space openness
↓
reasoning effortrather than:
amount of work
↓
reasoning effortThat is a more sophisticated approach to automatic inference budgeting.
Predictive Text Pushes OMP Into Terminal UX
One of the more surprising recent additions is a complete predictive-text system for the terminal composer.
OMP has added multiple prediction engines, including:
- N-gram prediction
- SmolLM2-based prediction
- Native macOS prediction
- Multi-engine blending
- Cross-process prediction services
- Background model downloading
- Ghost-text completion
The project also introduced an omp predict workflow for evaluating prediction behavior.
More interestingly, OMP can bootstrap predictions from existing Claude Code and Codex prompt history.
That changes the nature of the product.
An agent that merely runs inside a terminal is still a CLI tool. An agent that begins improving the terminal's actual text-entry experience is moving closer to the interaction model of an editor.
This reinforces OMP's broader direction: bring IDE-like intelligence into the terminal rather than forcing users to adopt a new graphical IDE.
Prompt Cache Warming Targets the Economics of Long Sessions
Recent OMP versions also introduced prompt-cache warming.
The idea is straightforward.
Long coding sessions can build expensive prompt prefixes. If the provider's prompt cache expires, the next request may need to reconstruct that context at full input-token cost.
OMP can proactively refresh the cache before expiration by replaying enough of the request to keep the prefix warm, then stopping generation almost immediately.
Available modes include:
off
streaming
idleThe default is designed around idle-time warming.
Importantly, the feature is cost-aware. OMP can compare the estimated cost of losing the cache against the cost of refreshing it and avoid warming when the expected saving is too small.
This is another sign of how coding-agent optimization is changing.
The question is no longer simply:
How many tokens did the agent use?
It is increasingly:
How efficiently did the agent use provider caching, reasoning budgets, model roles, and service tiers?
Web Search Now Includes Cost-Aware Fallbacks
OMP has also expanded its web-search routing.
Recent releases added OpenAI Responses-based search as part of the fallback chain.
The noteworthy part is the ordering.
The system can prefer search capacity available through coding subscriptions before falling back to separately billed API-key search.
That means routing decisions are influenced by billing path, not just technical availability.
For a power user running many coding-agent sessions, this can matter significantly. The cheapest usable route may depend on whether capacity comes from a subscription, API key, gateway, or local model.
OMP's multi-provider architecture increasingly treats cost as part of runtime orchestration.
The Stats Dashboard Is Becoming an Agent Analytics Tool
OMP's /stats work has also evolved beyond a basic token counter.
Recent changes include faster historical queries, asynchronous session synchronization, and a new Frustration view.
The Frustration analysis can use the configured judge model to classify user interactions and estimate analysis cost before running.
This represents an unusual direction for a coding agent.
Traditional developer tooling tracks:
- Requests
- Tokens
- Cost
- Models
- Session length
OMP is beginning to ask another question:
Where does the agent experience create user friction?
That could become useful for comparing model configurations, prompt strategies, extensions, or agent policies over long periods.
Privacy Controls Are Becoming More Explicit
Recent releases also added a clearer telemetry privacy control for OpenTelemetry exports.
Users can explicitly disable OTLP traces, logs, and metrics from the privacy settings even when OTEL endpoints exist in the environment.
This matters because OMP's growing feature set can involve multiple external systems:
- Model providers
- Search providers
- MCP servers
- Extensions
- Collaboration services
- Telemetry endpoints
- Browser automation
- Remote gateways
Privacy in a multi-provider coding agent is therefore not determined by one setting.
A local model can keep inference local, but that does not automatically make the entire workflow offline if other configured tools still communicate with external services.
Developers working with sensitive repositories should evaluate the complete execution graph, not only the selected language model.
Performance Work Is Happening Below the Feature Layer
Recent OMP updates also include substantial runtime optimization.
Examples include reducing allocation in streaming scanners, avoiding unnecessary raw SSE buffering when listeners are absent, improving content-block lookup paths, optimizing queue handling, and speeding syntax highlighting.
This matters because terminal coding agents increasingly operate in very long sessions.
A feature-rich agent can become unpleasant if every streamed token, tool call, syntax-highlight update, or session-history operation adds incremental overhead.
OMP's recent optimization work suggests that maintainers are addressing both sides of the product:
- Adding more agent capabilities
- Reducing the cost of keeping those capabilities active
That balance will matter as the system continues to grow.
Why the Recent Release Pace Matters
From v18.3.3 through v18.4.4, OMP has shipped a dense sequence of releases in only a few days.
The upside is obvious: new models and agent capabilities arrive very quickly.
The downside is equally important.
Fast-moving coding-agent infrastructure can create:
- Configuration churn
- Provider-specific regressions
- Behavior changes between minor releases
- Extension compatibility issues
- Short-lived bugs
- Unexpected changes to defaults
Developers using OMP for production work should therefore consider pinning known-good versions rather than automatically treating every release as risk-free.
Rapid iteration is a competitive advantage for experimentation. It is also an operational tradeoff.
Oh My Pi's Position Is Changing
OMP was initially easy to describe as a batteries-included fork or extension of Pi.
That description is becoming less useful.
The current architecture increasingly looks like this:
code intelligence
├── LSP
├── AST tooling
└── DAP debugging
agent runtime
├── main agent
├── subagents
├── advisor
├── judge
└── live steering
inference routing
├── providers
├── model roles
├── thinking effort
├── service tiers
└── local models
interaction
├── terminal UI
├── predictive text
├── ghost text
├── browser tools
└── queued steering
platform
├── CLI
├── SDK
├── RPC
├── ACP
├── MCP
└── extensionsA more accurate description in late September 2026 is:
Oh My Pi is becoming a programmable AI coding workstation and multi-model agent runtime for the terminal.
Oh My Pi vs Claude Code
Claude Code remains the simpler comparison for developers already committed to Anthropic's ecosystem.
Its advantage is a more vertically integrated model-and-agent experience.
OMP takes the opposite approach.
It emphasizes:
- Provider independence
- Model-role routing
- Local model support
- Extensibility
- Subagent orchestration
- IDE-style language intelligence
- Debugger integration
- Custom agent infrastructure
The choice is therefore not only about which agent can write better code.
It is also about whether the developer wants a vendor-integrated coding agent or a programmable multi-provider coding harness.
Oh My Pi vs Codex CLI
Codex CLI is a natural choice for developers centered on OpenAI's coding ecosystem.
OMP can use OpenAI models, including GPT-6.1 Sol, while keeping the surrounding runtime independent of OpenAI.
That makes OMP attractive to developers who expect their preferred models to change.
The model can change from OpenAI to Anthropic, Gemini, DeepSeek, a gateway, or a local backend without forcing the user to replace the entire agent environment.
The v18.4.4 Ultrafast work makes the comparison more interesting because OMP is now exposing OpenAI-specific service-tier behavior while still remaining multi-provider.
What Developers Should Watch Next
OMP's recent trajectory suggests several areas worth monitoring.
More Specialized Model Roles
Jev's judge integration shows that OMP does not assume every agent task belongs to a frontier coding model.
More specialized routing could follow for classification, review, planning, vision, commit generation, or tool selection.
More Speculative Execution
Early subagent launch is likely only one place where agent latency can be hidden.
Future runtimes may speculate on searches, repository reads, model calls, or alternative implementation branches.
Service-Tier-Aware Scheduling
With Ultrafast entering the router, latency tier may become another first-class scheduling variable alongside model quality and price.
Better Long-Session Economics
Prompt-cache warming points toward more optimization around cache lifetime, compaction, persistent memory, and context reuse.
IDE Capabilities Without an IDE
Predictive text, LSP, DAP, AST editing, and browser automation are gradually reducing the gap between a terminal agent and an AI-native editor.
Conclusion
Oh My Pi v18.4.4 is important, but the version number is not the real story.
GPT-6.1 Sol support, OpenAI Ultrafast routing, Claude Sonnet 5.5 between-tools thinking, Jev judging, speculative subagents, predictive text, live steering, prompt-cache warming, and richer analytics all point in the same direction.
OMP is evolving from a coding assistant into an agent operating environment.
For developers who want a minimal terminal tool, that expanding surface area may be unnecessary.
For developers who want to combine Claude, GPT, Gemini, DeepSeek, local models, subagents, debugger state, language-server intelligence, custom tools, and cost-aware routing inside one programmable terminal workflow, OMP is becoming increasingly distinctive.
The most important metric to watch now is not how many providers Oh My Pi supports.
It is how effectively the runtime can decide which model, role, tool, reasoning level, and service tier should handle each part of a software-engineering task.
Continue Reading
More articles connected to the same themes, protocols, and tools.
Referenced Tools
Browse entries that are adjacent to the topics covered in this article.






