| title | Guide: Per-Agent Model Configuration |
|---|---|
| sidebar_label | Agent Models |
| description | Configure which AI model each agent uses via model_preset in oma-config.yaml. Covers built-in presets, per-agent overrides, inline model definitions, custom presets with extends, oma doctor --profile, and migration from legacy agent_cli_mapping. |
model_preset: auto is the default for new installs. Unconfigured agents use the current vendor's native agent definitions and model settings. Choose a fixed preset to pin models, or override individual agents when you need a different model or vendor. Existing explicit presets are preserved on reinstall and update.
Shared configuration lives in .agents/oma-config.cue or .agents/oma-config.yaml. An optional Git-ignored local file overrides settings for your machine.
For the full top-level key and precedence reference, see Configuration reference.
This page covers:
- The built-in presets
- Overriding individual agents with the
agents:map - Inlining custom model slugs with
models: - Defining custom presets with
custom_presets:andextends: - Inspecting resolved configuration with
oma doctor --profile - Migration from legacy
agent_cli_mapping
Set model_preset to one of the built-in keys:
# .agents/oma-config.yaml
language: en
model_preset: auto| Key | Description | Best for |
|---|---|---|
auto |
Follow the current runtime's agent/model settings without injecting a model or effort flag | Default for new installs |
free |
Special gateway mode for OMA-spawned Codex, Claude, or Qwen processes; it is resolved separately from the built-in preset registry. | Local FreeLLMAPI gateway |
antigravity |
All agents use Antigravity CLI (agy): Gemini 3.1 Pro for implementation/architecture and Gemini 3.6 Flash for orchestration, documentation, and explore. Model selection is config-driven inside agy — no --model or --thinking-budget flags are exposed. |
Antigravity CLI users |
claude |
All agents use Claude (Sonnet/Opus) | Claude Max subscription holders |
codex |
All agents use OpenAI Codex (GPT-5.5 for most roles, GPT-5.4-mini for explore) with effort levels | ChatGPT Plus/Pro users |
qwen |
All agents use Qwen Code; matching Qwen sessions can use generated native agents, and other runtimes use CLI dispatch | Local / self-hosted inference |
kiro |
All agents use Kiro CLI; Sonnet handles implementation/architecture and Haiku handles orchestration/explore | Kiro users |
cursor |
All agents use Cursor composer-2.5 (composer-2.5-fast for orchestrator/qa/pm/docs/explore) |
Cursor Pro / Pro Student users |
mixed |
Mixed: impl roles use Codex, architecture/qa/pm use Claude, explore uses Gemini | Cross-vendor strengths without managing per-agent config |
Built-in presets ship inside the CLI package and update automatically when you upgrade oh-my-agent. gemini is a compatibility alias that redirects to antigravity; it is not a separate current preset. No local preset file is required.
With auto, explicit agents.<id> model overrides take priority. Otherwise, OMA detects the current runtime and uses its native subagent path when available. Cross-vendor agents and runtimes without native dispatch use oma agent spawn. Auto does not expand to a fixed vendor preset.
For CLI dispatch, --vendor explicitly selects the target. Without it, OMA uses the detected runtime, then default_cli when detection fails (claude if omitted). Inherited plans do not inject OMA model or effort flags; the vendor's own agent/session configuration supplies them. An external CLI process uses that CLI's persisted defaults, which may differ from a model selected only in the parent session.
oma doctor --profile displays (vendor agent default) for inherited agents and the resolved model for explicit overrides. Native agent files keep their vendor definitions; same-vendor overrides in auto mode are applied when those files are generated by install/update.
Create one of .agents/oma-config.local.cue or .agents/oma-config.local.yaml next to the shared configuration. Install, link and update add both paths to .gitignore; update preserves existing local files, including with --force.
OMA selects the nearest project configuration directory. Within that directory, shared CUE takes priority over shared YAML, and the local file overrides the shared values. CUE files are evaluated independently before merging, so a shared model_preset: "auto" can be replaced by local "free". Objects merge recursively; arrays, scalars and null replace the shared value. A malformed local file, a missing CUE executable for local CUE, or both local formats being present is an error, not permission to use shared defaults.
Command options and supported environment overrides take precedence over the effective file configuration. oma doctor --profile shows which files were used. Local files do not travel with Git clones or new worktrees. Free-mode subprocesses inherit OMA_MODEL_PRESET=free and the resolved gateway environment so nested OMA spawns preserve the route; independently launched sessions need their own local config or environment. Settings saved by install/setup commands still target the shared configuration; a local override continues to win at runtime.
Keep model_preset: auto in the shared file and opt in locally:
// .agents/oma-config.local.cue
model_preset: "free"
free: {
base_url: "http://127.0.0.1:31415/v1"
api_key_env: "FREELLM_API_KEY"
model: "auto"
}The equivalent YAML file is:
# .agents/oma-config.local.yaml
model_preset: free
free:
base_url: http://127.0.0.1:31415/v1
api_key_env: FREELLM_API_KEY
model: autoStart FreeLLMAPI separately and export its unified key as FREELLM_API_KEY. OMA also accepts upstream's FREELLMAPI_API_KEY when the default key variable is selected; the canonical variable wins when both are set. A custom api_key_env reads only that variable. Never put the key itself in the config. OMA_MODEL_PRESET overrides the preset. FREELLM_BASE_URL and FREELLM_MODEL override the respective file settings. The values in the example are the defaults, so model_preset: free alone is enough when the server and key are ready.
oma doctor --profile
oma agent spawn backend "Review the API error handling" free-review --vendor codex --read-onlyFree mode uses free.model for every OMA-dispatched role, including roles with existing agents.*.model pins. It does not resolve those pins into paid subscriptions. Choose auto, a gateway model ID, or a named gateway chain such as auto:coding (create that chain in FreeLLMAPI first).
The transport is --vendor, then OMA_RUNTIME_VENDOR, then a supported detected runtime, then default_cli, then codex. Only Codex, Claude and Qwen transports are supported. An explicitly selected unsupported transport is an error.
| Transport | Gateway endpoint | CLI base URL |
|---|---|---|
| Codex | /v1/responses |
Includes /v1 |
| Claude | /v1/messages |
Server root; OMA removes the /v1 suffix |
| Qwen | /v1/chat/completions |
Includes /v1 |
Use oma agent spawn even when the parent uses the same vendor. OMA injects the gateway connection and credentials into that subprocess only; changing the preset does not change the model of an already-open host session or host-native subagent tool. Codex gets a custom Responses provider through invocation arguments, while the key stays in the child environment. Claude and Qwen receive their compatible endpoint settings. Conflicting Claude/Qwen settings that would override the route or key are reported before execution; OMA does not rewrite those files.
Spawn and review check authenticated GET /v1/models before starting the agent. Missing keys, connection failures and HTTP authentication errors stop execution. oma doctor --profile shows the effective URL/model, environment overrides, key presence and server readiness without printing the key. Readiness does not guarantee a model has enough quota to complete a task.
FreeLLMAPI owns request-level provider failover. OMA's explicit checkpoint-based vendor failover remains a separate process recovery mechanism; every successor in free mode must still use a supported FreeLLMAPI transport. No automatic return to a paid vendor configuration occurs.
The free preset configures agent inference. It does not change the embedding configuration of existing memory services. FreeLLMAPI also exposes /v1/embeddings; when configuring a vector store separately, pin a model family so existing vectors retain a compatible space.
Upstream references: client setup, API and embedding families.
Use the agents: map to override specific agents on top of the active preset. Only agents you list are affected; the rest follow the vendor settings in auto mode or the selected fixed preset defaults.
# .agents/oma-config.yaml
language: en
model_preset: auto
agents:
backend: { model: openai/gpt-5.5, effort: high }
qa: { model: anthropic/claude-sonnet-4-6 }Each entry is an AgentSpec object:
| Field | Type | Required | Description |
|---|---|---|---|
model |
string | Yes | Model slug (built-in or user-defined) |
effort |
none | low | medium | high | xhigh |
No | Reasoning effort (ignored on models that do not support it) |
thinking |
boolean | No | Enable extended thinking (model-specific) |
memory |
user | project | local |
No | Memory scope for the agent |
Valid agent IDs: orchestrator, architecture, qa, pm, backend, frontend, mobile, db, debug, refactor, docs, tf-infra, explore.
The merge is shallow: each field in your override replaces the preset value for that field. Fields you omit keep their preset value.
Register model slugs that are not yet in the built-in registry under models:. Once registered, reference the slug from agents: or custom_presets:.
# .agents/oma-config.yaml
models:
google/gemini-3-flash-fast:
cli: gemini
cli_model: gemini-3-flash
auth_hint: "Google AI Pro"
supports:
effort: null
apply_patch: false
task_budget: false
prompt_cache: false
computer_use: false
native_dispatch_from: [gemini]
api_only: falseTwo rules apply to a registered slug you reference from agents::
- The key must be in
owner/modelform.agents.<id>.modelvalidates against anowner/modelpattern, so a bare key likemy-fast-modelis rejected — use a slashed key such asgoogle/gemini-3-flash-fast(or the vendor's ownprovider/modelslug). - The spec must be complete.
cli,cli_model,auth_hint, and everysupportsboolean are required at resolution time. An incomplete spec is accepted by the config parser but fails model-registry validation and silently falls back to the core registry.
If a user-defined slug collides with a built-in slug, the user definition wins and a warning is emitted.
Define additional presets in custom_presets:. Use extends: to inherit all agent defaults from a built-in preset and override only the agents you care about.
# .agents/oma-config.yaml
language: en
model_preset: my-team
custom_presets:
my-team:
extends: claude # base preset — partial merge
description: "Team A — sonnet base, codex for implementation"
agent_defaults:
backend: { model: openai/gpt-5.5, effort: high }
db: { model: openai/gpt-5.5, effort: high }
# all other agents inherited from claudeWithout extends:, provide defaults for the canonical agent roles used by the preset. With extends:, only the entries you list are overridden; the rest are inherited from the base preset.
Run oma doctor --profile to inspect the fully resolved model matrix after preset defaults, custom_presets, and agents: overrides are merged.
oma doctor --profileSample output:
oh-my-agent — Profile Health (preset=mixed)
┌──────────────┬──────────────────────────────┬──────────┬──────────────────┬──────────┐
│ Role │ Model │ CLI │ Auth Status │ Source │
├──────────────┼──────────────────────────────┼──────────┼──────────────────┼──────────┤
│ orchestrator │ anthropic/claude-sonnet-4-6 │ claude │ ✓ logged in │ (preset) │
│ architecture │ anthropic/claude-opus-4-7 │ claude │ ✓ logged in │ (preset) │
│ qa │ anthropic/claude-sonnet-4-6 │ claude │ ✓ logged in │ (preset) │
│ backend │ openai/gpt-5.5 │ codex │ ✗ not logged in │ (override)│
│ explore │ google/gemini-3.1-flash-lite │ gemini │ ✗ not logged in │ (preset) │
└──────────────┴──────────────────────────────┴──────────┴──────────────────┴──────────┘
Each row shows the resolved model slug and which source applied it ((preset) or (override)). Use this whenever a subagent picks an unexpected vendor.
Migration 008 runs automatically on oma install and oma update. It converts legacy projects in place:
| Legacy config | Result after migration 008 |
|---|---|
All entries same vendor (e.g. all gemini) |
model_preset: gemini, no agents: |
| Mixed vendors | Most-frequent vendor → model_preset; others → agents: overrides |
AgentSpec object values |
Moved to agents: as-is |
models.yaml content |
Inlined into oma-config.yaml.models |
Customized defaults.yaml |
Preserved as custom_presets.user-customized with a warning |
Originals are backed up to .agents/.backup-pre-008-{timestamp}/ before any changes. The migration is idempotent. If model_preset is already present, it skips.
After migration, .agents/config/defaults.yaml, .agents/config/models.yaml, and the .agents/config/ directory are removed.
session.quota_cap is unchanged. Add it to oma-config.yaml to bound runaway subagent spawning:
session:
quota_cap:
tokens: 2_000_000
spawn_count: 40
per_vendor:
claude: 1_200_000
openai: 600_000
google: 200_000When a cap is reached, the orchestrator refuses further spawns and surfaces a QUOTA_EXCEEDED status.
# .agents/oma-config.yaml
language: en
model_preset: my-team
agents:
frontend: { model: anthropic/claude-sonnet-4-6 }
models:
google/gemini-3-flash-fast:
cli: gemini
cli_model: gemini-3-flash
auth_hint: "Google AI Pro"
supports:
effort: null
apply_patch: false
task_budget: false
prompt_cache: false
computer_use: false
native_dispatch_from: [gemini]
api_only: false
custom_presets:
my-team:
extends: claude
description: "Sonnet base, Codex for backend/db"
agent_defaults:
backend: { model: openai/gpt-5.5, effort: high }
db: { model: openai/gpt-5.5, effort: high }
session:
quota_cap:
tokens: 2_000_000
spawn_count: 40Run oma doctor --profile to confirm resolution, then start a workflow as usual.
pi (Earendil) is a multi-provider proxy
runtime rather than a model owner — it can run any real-provider model
(Anthropic, OpenAI, Google) under one CLI. oma treats pi as a transport
overlay: your model_preset and agents: overrides stay exactly as they are,
and pi becomes the executing CLI for a given agent.
Dispatch any agent through pi with the --vendor pi override:
oma agent spawn backend "Implement the export endpoint" <session> --vendor piWhat happens:
- The per-agent model resolved from your preset/overrides (e.g.
openai/gpt-5.5) is translated to pi's--model <provider/id>form, andeffortis translated to pi's--thinkinglevel. Per-subagent models work on pi exactly as they do natively — different agents can run different models. - The agent's persona (system prompt) is inlined from
.agents/agents/<id>.md, since pi has no vendor-side agent file to reference. - Auth is whatever pi itself is configured for (
~/.pi/agent/auth.jsonor a provider API key in the environment).oma doctorreports pi install + auth status alongside the other CLIs.
Constraint: pi only runs real-provider models. CLI-proprietary presets
(cursor, kiro, qwen, antigravity) name models that exist only inside
their own CLIs, so dispatching them through pi is rejected with a clear error.
Use a real-provider preset (claude, codex, gemini, or mixed) when routing
agents through pi.
pi's model catalog is release-tracked and auth-gated. If a resolved slug does not match what your pi install exposes, check
pi --list-models— pi's--modelmatching is fuzzy, so most provider slugs resolve as-is.
pi resolves --model against its built-in model registry, and its
defaultProvider setting is only consulted when no model is passed at all. For
Z.ai, pi ships only a subset of GLM ids (glm-4.7, glm-4.5-air,
glm-5-turbo, glm-5.1, glm-5v-turbo as of pi 0.80.x) — a preset that names
any other GLM id will fail to resolve.
Two ways to handle this:
- Registry ids — constrain your preset to registry model ids. Use the
provider/idform (e.g.zai/glm-4.7) to pin the provider explicitly; oma passes it through to pi's--modelas-is. - Unregistered ids — register them with a pi extension. The
apifield must name one of pi's api adapter ids (openai-completions,anthropic-messages, …) — not the provider name. Provider names like"zai"or shorthands like"openai"are not adapter ids and fail at dispatch withNo API provider registered for api: ….
// ~/.pi/agent/extensions/zai-glm-models/index.ts (or <project>/.pi/extensions/)
export default function (pi: ExtensionAPI) {
pi.registerProvider("zai", {
baseUrl: "https://api.z.ai/api/coding/paas/v4",
api: "openai-completions", // adapter id, NOT "zai"
apiKey: "$ZAI_API_KEY",
models: [
{ id: "glm-4.7-flash", api: "openai-completions", /* … */ },
// NOTE: `models` replaces ALL existing models for the provider —
// re-declare the built-in ids here if you still want them.
],
});
}Verify with pi --list-models before wiring the ids into a preset.
OpenCode is an extension-class vendor: like pi, it is not
a model owner but a CLI that runs models from its own catalog — the free
opencode provider, the low-cost opencode-go subscription plan, and the
opencode-zen gateway. oma integrates it as an in-process plugin vendor:
opencode auto-loads .opencode/plugins/oma/ instead of registering settings-file
hooks, and resolves each agent's persona from generated .opencode/agents/<id>.md
files.
Route any agent through opencode with the --vendor opencode override:
oma agent spawn pm "Draft the rollout plan" <session> --vendor opencodeThis runs opencode run --agent pm --dir <workspace> "<prompt>". The prompt is a
trailing positional argument — opencode's -p flag means --password, not
the prompt.
To route specific agents to an opencode model, register the model under models:
and reference it from agents:. Two requirements apply (see
Inlining model slugs):
- Slug must be in
owner/modelform. Use the opencodeprovider/modelslug as the registry key — bare names are rejected by theagents.<id>.modelschema. - The spec must be complete —
cli,cli_model,auth_hint, and everysupportsboolean. An incomplete spec fails validation and silently falls back to the core registry (so the agent would not route to opencode).
# .agents/oma-config.yaml
language: en
model_preset: claude # heavier impl roles stay on Claude
models:
opencode-go/deepseek-v4-flash:
cli: opencode
cli_model: opencode-go/deepseek-v4-flash
auth_hint: "OpenCode Go subscription — run: opencode auth login"
supports:
effort: null
apply_patch: false
task_budget: false
prompt_cache: false
computer_use: false
native_dispatch_from: [opencode]
api_only: false
agents:
pm: { model: opencode-go/deepseek-v4-flash }
qa: { model: opencode-go/deepseek-v4-flash }
docs: { model: opencode-go/deepseek-v4-flash }
explore: { model: opencode-go/deepseek-v4-flash }Each routed agent dispatches opencode run -m opencode-go/deepseek-v4-flash --agent <id> --dir <workspace> "<prompt>". This is a good fit for lightweight,
fast roles (pm, qa, docs, explore) while heavier implementation agents stay on
Codex/Claude/etc.
opencode's catalog is subscription- and login-gated, so oma does not hardcode opencode model slugs. Validate one against your installed catalog:
oma model probe opencode-go/deepseek-v4-flash --json # accepted | rejected | auth_required
opencode models opencode-go # list everything your plan exposesoma model probe reports accepted when the slug is listed by
opencode models, rejected when it is not, and auth_required when the
provider needs login or a subscription.
- Auth:
opencode auth loginstores credentials in~/.local/share/opencode/auth.json, one entry per provider.oma auth status/oma doctorreport opencode as authenticated when any provider has a credential.oma doctor --profileis provider-aware instead: each row is checked against the provider prefix of its registeredcli_model, so a model withcli_model: zai-coding-plan/glm-5.3is checked against thezai-coding-plancredential. A row whose model has no registeredprovider/modelcli_modelreports? unknownrather than a definite auth failure. - Generated files:
oma link(oroma link opencode) writes one.opencode/agents/<id>.mdpersona per agent plus the.opencode/plugins/oma/bridge. These are generated from the.agents/SSOT — do not edit them directly; re-runoma linkto regenerate.
Persistent-workflow note: opencode's
session.idleevent (its nearest analog to the ClaudeStophook) is notification-only and cannot block the session from ending. Persistent workflows (orchestrate / work / ultrawork) therefore run with degraded Stop semantics under opencode — workflow reinforcement happens on the next message rather than by holding the session open.
Kimi Code CLI reads hooks only from a global
config (~/.kimi-code/config.toml, KIMI_CODE_HOME), so oma install/oma link
write the Kimi hook chain and its skill symlinks into HOME under explicit consent
(like Antigravity). Kimi also scans oma's SSOT .agents/skills/ directly, so
skills resolve project-wide regardless. MCP needs no HOME write and is
project-scoped — written mode-aware to <cwd>/.kimi-code/mcp.json (project) or
~/.kimi-code/mcp.json (global).
Route any agent through Kimi with the --vendor kimi override:
oma agent spawn pm "Draft the rollout plan" <session> --vendor kimiThis runs kimi -p "<prompt>". Kimi's -p (non-interactive) mode auto-approves
regular tool calls under its auto permission policy, so oma does not append
--yolo/--auto (they are mutually exclusive with -p).
Like opencode, oma does not hardcode a Kimi model catalog (Kimi's lineup is
provider/subscription-dependent). To route specific agents to a Kimi model,
register a complete spec under models: with cli: kimi and reference it from
agents::
The registry key must be in owner/model form (bare names are rejected by the
agents.<id>.model schema), and cli_model is the exact alias passed to
kimi --model — Kimi's documented coding alias is kimi-code/kimi-for-coding.
Confirm the alias your subscription exposes with kimi --model <alias> before
committing it.
# .agents/oma-config.yaml
models:
kimi-code/kimi-for-coding:
cli: kimi
cli_model: kimi-code/kimi-for-coding
auth_hint: "Kimi subscription — run: kimi login"
supports:
effort: null
apply_patch: false
task_budget: false
prompt_cache: false
computer_use: false
native_dispatch_from: []
api_only: false
agents:
pm: { model: kimi-code/kimi-for-coding }
docs: { model: kimi-code/kimi-for-coding }Each routed agent dispatches kimi --model kimi-code/kimi-for-coding -p "<prompt>".
Persistent-workflow note: Kimi's documented Stop-blocking path is exit-code 2 / stderr, but the
oma hook runrouter always exits 0 and emits a stdout dialect. oma emits a best-effortpermissionDecision: "deny"(plus Claude-styledecision: "block") so persistent workflows degrade gracefully under Kimi.