OMW
OMW = OpenAI + MCP + WASM.
omw is an agent runtime. You declare agents in a TOML configuration, each
wiring a provider (an OpenAI-family chat service), tooling (MCP tool
servers), and a brain — either a compiled WASM component, a rhai script, or
a JavaScript script. omw then drives each agent through an actor model: every
agent owns a single inbox, and chat streams, tool results, timers, and messages
from other agents all arrive there as tagged events the brain consumes.
Who it is for
omw is for people who want a small, local agent runtime that is genuine about
its inputs and outputs: the brain is real WASM, tooling speaks MCP, and the
configuration is plain TOML. It is not a framework — there is no DSL to learn
and no orchestration layer. You bring a provider key, a couple of MCP servers,
and a brain, and omw runs it for one iteration (run) or keeps it going
(loop).
How it works
- Providers are OpenAI-family chat services.
provider.chat-streamopens a streaming response whose deltas arrive as events in the agent’s inbox; the in-bandprovider.chatblocks for the full result. - Tooling is MCP tool servers. They expose callable tools and readable resources; resource subscriptions deliver change events.
- Brains are runtimes. The
wasmruntime loads an agent as a compiled component; therhairuntime evaluates a script on an interpreter that ships as an opt-in variant, as does thejsruntime. The defaultomwpackage/binary ships with theruntime-wasm,provider-openai,tooling-mcp, andendpoint-openaiback ends but without either script runtime, whereas theomw-rhaipackage /omw-rhai-<arch>.tar.gzbinary (--features runtime-rhai) includes the rhai interpreter and theomw-jspackage /omw-js-<arch>.tar.gzbinary (--features runtime-js) includes the js interpreter. All three see the sameomwhost interface (rhai in snake_case, js in camelCase). To write a pure Rust brain, depend on theomw-wasm-rustguest SDK crate instead of runningwit-bindgenyourself; see Rust brains. - Agents are actors. They subscribe to each other explicitly, so a message only ever reaches an agent that chose to listen.
- Endpoint is an optional OpenAI-compatible HTTP server. Set
[endpoint]withkind = "openai"plus alistenaddress and agents can subscribe themselves under model names: inbound chat requests arrive in the agent’s inbox as events, and the agent streams its reply back (SSE or buffered JSON). Any OpenAI-compatible client can then drive an agent. - Hot reload is
--watchonrun/loop. When a brain script changes, the agent’s run restarts on the new script while inboxes, subscriptions, and sessions survive. The new script is validated before the live run ends, so a bad edit never kills a good run — and a broken script never starts.
Installation
omw is packaged as a Nix flake. Run it directly without installing:
nix run github:haras-unicorn/omw
or build the omw binary with:
nix build github:haras-unicorn/omw
Releases
Prebuilt binaries for x86_64-linux and aarch64-linux are attached to each
GitHub release as tarballs containing the omw binary. The default
omw-<arch>.tar.gz ships no rhai runtime; grab the omw-rhai-<arch>.tar.gz
tarball (or the rhai Nix package) when your brains are rhai scripts, or the
omw-js-<arch>.tar.gz tarball (or the js Nix package) when your brains are
JavaScript scripts:
curl -L -o omw.tar.gz \
https://github.com/haras-unicorn/omw/releases/latest/download/omw-x86_64-linux.tar.gz
tar -xzf omw.tar.gz
./omw-x86_64-linux
The rhai variant is the same shape, with the -rhai name:
curl -L -o omw-rhai.tar.gz \
https://github.com/haras-unicorn/omw/releases/latest/download/omw-rhai-x86_64-linux.tar.gz
tar -xzf omw-rhai.tar.gz
./omw-rhai-x86_64-linux
The js variant is the same shape, with the -js name:
curl -L -o omw-js.tar.gz \
https://github.com/haras-unicorn/omw/releases/latest/download/omw-js-x86_64-linux.tar.gz
tar -xzf omw-js.tar.gz
./omw-js-x86_64-linux
Usage
Configuration lives in a TOML file (default omw.toml in the current directory,
overridable with --config). It declares named providers, tooling, and runtimes
plus a list of agents:
[providers.openai]
kind = "openai"
api_key = "sk-…"
model = "gpt-4o"
[tooling.mcp]
kind = "mcp"
transport = "stdio"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-everything"]
[runtime.rhai]
kind = "rhai"
[[agents]]
name = "alice"
runtime = "rhai"
script = "brain.rhai"
Then drive it:
omw run # run every agent once
omw loop # keep every agent running, restarting on failure
Serve agents over HTTP with the optional endpoint: set kind plus a listen
address, have a brain subscribe itself under a model name, then any
OpenAI-compatible client can call it:
[endpoint]
kind = "openai"
listen = "127.0.0.1:8080"
let sub = omw::host::subscribe_endpoint("gpt-4o");
Inbound requests arrive in the agent’s inbox as endpoint-message events; the
brain streams its reply back with stream_endpoint (SSE for stream: true, one
buffered JSON completion otherwise). GET /v1/models lists subscribed models.
Edit brains live with --watch on either mode: when a brain file changes, the
agent’s current run ends and restarts on the new script, while inboxes,
subscriptions, and sessions survive on the shared bus. The new script is
validated before the live run ends, so a bad edit keeps the good run alive
(plus an error event if the brain subscribed to lifecycle events) — and a
broken script never starts (parks under --watch, fails fast without it).
Brains opt in to reload / shutdown notices with subscribe_lifecycle.
See the endpoint and hot reload pages for the full reference.
To write a pure Rust brain, depend on the omw-wasm-rust guest SDK crate
instead of running wit-bindgen yourself; see Rust brains.
Configuration can also be layered from the environment (OMW__ prefix) or
generated as a JSON schema:
omw schema --output config.schema.json
See the docs for the full reference.
NixOS
The flake ships a NixOS module exposing services.omw — a hardened systemd unit
that runs omw from a config file, layering OMW__-prefixed environment
variables over it so secrets never live in the Nix store. Plain-systemd and
Docker deployments are covered too — see Deployment in the documentation:
{
inputs = {
nixpkgs.url = "github:nixos/nixpkgs/nixos-26.05";
omw.url = "github:haras-unicorn/omw";
};
nixosConfigurations.my-machine = nixpkgs.lib.nixosSystem {
modules = [
omw.nixosModules.default
{
services.omw = {
enable = true;
settingsFile = "/etc/omw.toml";
environmentFile = "/var/lib/omw/env";
};
}
];
};
}
See The NixOS module in the documentation for the full option set, including
settings vs settingsFile, mode, user/group, and stateDir.
Library
omw can be used as a library in your own crate by adding omw to dependencies
and enabling the runtime features you want. See the library page for details.
Examples and testing
The examples are runnable agents that exercise the whole stack — provider,
tooling, endpoint, agents — with no keys, no network, and no external services,
each in rhai, js, and wasm. The omw-test binary runs them deterministically
against in-process scripted doubles and checks what every agent saw and did
against an [assertions] section; the testing pages cover the binary, the
assertion language, and each mock. See examples for the tour.
Binary cache
Builds are cached on the haras cachix cache. When the flake is used directly
(for example with nix run github:haras-unicorn/omw), the cache is configured
automatically through the flake’s nixConfig. To use it when the package comes
from an overlay, add the following to your nix configuration:
{
nix.settings = {
substituters = [ "https://haras.cachix.org" ];
trusted-public-keys = [
"haras.cachix.org-1:/HIo1JYqOIH1Nwk1EGXhuPPvDW0WekxIbY5CiXUZbYw="
];
};
}
References
The machine-readable configuration schema, alongside the human-readable pages in this book.
Configuration schema
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"title": "Config",
"description": "OMW configuration.",
"type": "object",
"properties": {
"agents": {
"description": "Agents that OMW is going to run.",
"type": "array",
"default": [],
"items": {
"$ref": "#/$defs/AgentConfig"
}
},
"endpoint": {
"description": "Optional endpoint implementation.",
"anyOf": [
{
"$ref": "#/$defs/ImplConfig"
},
{
"type": "null"
}
],
"default": null
},
"memory": {
"description": "Per-agent seeded memory, keyed by agent name then key. Seeded into the\nagent's memory before its brain first runs, so a test (or a deployment)\ncan fast-forward an agent to a state. Seeded values persist like any\nother memory, including across hot reloads.",
"type": "object",
"additionalProperties": {
"type": "object",
"additionalProperties": {
"type": "string"
}
},
"default": {}
},
"providers": {
"description": "Named provider implementations.",
"type": "object",
"additionalProperties": {
"$ref": "#/$defs/ImplConfig"
},
"default": {}
},
"runtime": {
"description": "Named runtime implementations.",
"type": "object",
"additionalProperties": {
"$ref": "#/$defs/ImplConfig"
},
"default": {}
},
"tooling": {
"description": "Named tooling implementations.",
"type": "object",
"additionalProperties": {
"$ref": "#/$defs/ImplConfig"
},
"default": {}
},
"tunables": {
"description": "Global runtime tunables.",
"$ref": "#/$defs/Tunables",
"default": {
"allow_unlocked_secrets": false,
"cancel_pumps_on_reload": true,
"inbox_bound": 1024,
"interrupt_budget_ms": 100,
"loop_backoff_cap_secs": 30,
"loop_backoff_start_ms": 100,
"recv_slice_ms": 200,
"recv_timeout_secs": 60,
"reload_grace_secs": 5,
"reload_poll_ms": 200,
"session_buffer": 8192,
"tooling_connect_backoff_cap_secs": 30,
"tooling_connect_backoff_start_ms": 100,
"trace_buffer": 4096,
"watch_debounce_ms": 200
}
}
},
"$defs": {
"AgentConfig": {
"description": "A single agent wiring itself to the globals above.",
"type": "object",
"properties": {
"name": {
"type": "string"
},
"runtime": {
"description": "Which named runtime implementation this agent's brain uses.",
"type": "string"
},
"script": {
"description": "The agent's brain script.",
"type": "string"
}
},
"required": ["name", "runtime", "script"]
},
"ImplConfig": {
"description": "A single configured implementation: which kind plus opaque params.",
"type": "object",
"properties": {
"kind": {
"description": "Which implementation this is.",
"type": "string"
}
},
"additionalProperties": true,
"required": ["kind"]
},
"Tunables": {
"description": "Runtime tunables: channel sizes, timeouts, and backoffs. All optional;\nomitted values fall back to the built-in defaults.",
"type": "object",
"properties": {
"allow_unlocked_secrets": {
"description": "Permit secrets to stay unlocked (pageable) when `mlock` fails, e.g.\ninside containers where the outer `RLIMIT_MEMLOCK` cannot be raised.\nDefault `false` (fail-closed). Only enable where the weaker guarantee\nis acceptable.",
"type": "boolean",
"default": false
},
"cancel_pumps_on_reload": {
"description": "Cancel open pumps (streams, timers, resources, tool calls) on reload.\n`false` keeps them across reload.",
"type": "boolean",
"default": true
},
"inbox_bound": {
"description": "How many events a single agent inbox buffers before sends fail.",
"type": "integer",
"format": "uint",
"default": 1024,
"minimum": 0
},
"interrupt_budget_ms": {
"description": "Uninterrupted runtime execution allowed after grace expires, in ms.",
"type": "integer",
"format": "uint64",
"default": 100,
"minimum": 0
},
"loop_backoff_cap_secs": {
"description": "Backoff cap for `loop` restarts on failure, in seconds.",
"type": "integer",
"format": "uint64",
"default": 30,
"minimum": 0
},
"loop_backoff_start_ms": {
"description": "Backoff start for `loop` restarts on failure, in ms.",
"type": "integer",
"format": "uint64",
"default": 100,
"minimum": 0
},
"recv_slice_ms": {
"description": "How long `recv_while` parks between early-abort checks, in ms.",
"type": "integer",
"format": "uint64",
"default": 200,
"minimum": 0
},
"recv_timeout_secs": {
"description": "How long a blocking `recv` waits before timing out, in seconds.",
"type": "integer",
"format": "uint64",
"default": 60,
"minimum": 0
},
"reload_grace_secs": {
"description": "How long the supervisor waits for a cooperative exit, in seconds.",
"type": "integer",
"format": "uint64",
"default": 5,
"minimum": 0
},
"reload_poll_ms": {
"description": "How long the blocking-call helper waits between reload checks, in ms.",
"type": "integer",
"format": "uint64",
"default": 200,
"minimum": 0
},
"session_buffer": {
"description": "How many deltas a single endpoint session buffers before drops.",
"type": "integer",
"format": "uint",
"default": 8192,
"minimum": 0
},
"tooling_connect_backoff_cap_secs": {
"description": "Backoff cap for tooling reconnects on failure, in seconds.",
"type": "integer",
"format": "uint64",
"default": 30,
"minimum": 0
},
"tooling_connect_backoff_start_ms": {
"description": "Backoff start for tooling reconnects on failure, in ms.",
"type": "integer",
"format": "uint64",
"default": 100,
"minimum": 0
},
"trace_buffer": {
"description": "How many trace events the `omw-test` broadcast channel buffers.",
"type": "integer",
"format": "uint",
"default": 4096,
"minimum": 0
},
"watch_debounce_ms": {
"description": "How long to coalesce the burst of file events a single save produces,\nin ms.",
"type": "integer",
"format": "uint64",
"default": 200,
"minimum": 0
}
}
}
}
}
The actor model
omw is built around a small actor model. Each configured agent is an actor: it
owns a single inbox on a shared event bus, and runs its “brain” — the agent
runtime — for repeated iterations against that inbox. Everything the agent
touches (chat streams, tooling resources, timers, other agents) arrives at that
one inbox as a tagged event, so the brain is a pure, mostly-synchronous event
consumer.
This page explains why the model is shaped that way and how the pieces fit together. The concrete interfaces are documented in host.
The single inbox
Every agent gets exactly one inbox: a bounded channel (see inbox_bound in
tunables) on a shared message bus. Nothing is routed to the
brain directly — chat deltas, tool results, timers, and messages from other
agents all land in the same inbox as an EventEnvelope carrying:
id— the UUID of the subscribed source the event came from.event— the strongly-typed payload (message,error,timer,chat-delta,chat-end,resource-list-updated,resource-updated,endpoint-message,endpoint-session-end).
The brain consumes events with host.recv (a blocking receive with a host-side
timeout, see recv_timeout_secs in tunables) or
host.try-recv (a non-blocking poll). Because every event is tagged with a
UUID, a single inbox is enough to multiplex many concurrent sources — the brain
correlates a delta or timer to the specific handle that opened it by matching
the envelope id against the UUID returned by the call that created it.
Nothing is shared by default
Inter-agent messaging is subscription-based. An agent does not receive anything from another agent unless it explicitly subscribed:
host.subscribe-agent(agent)returns a new UUID handle for that source.host.unsubscribe-agent(uuid)removes a subscription by handle.host.send-agent(agent, payload)delivers the message only if the recipient subscribed to the sender. Each recipient’s message is tagged with the UUID of its own subscription to the sender, not a global topic, so a sender fanning out to many subscribers reaches each one through a distinct handle.
This keeps the coupling between actors explicit and auditable: an agent can only ever be contacted by the actors it chose to listen to.
I/O sources are pump tasks
Synchronous wasm brains cannot await async I/O directly. Instead, every
long-lived I/O source is driven by a pump task spawned onto the agent’s bridge
runtime (the AgentContext.rt tokio runtime), which pushes events into the
inbox:
- a
provider.chat-streamcall spawns a chat-stream pump that deliverschat-deltaevents as chunks arrive and a terminalchat-end(orerror) event when the stream closes. - a
tooling.call-toolqueues a tool-call pump that delivers atool-result(orerroron failure) event. - a
tooling.subscribe-resource-list/subscribe-resourcecall spawns a resource pump that deliversresource-list-updated/resource-updatedevents. - a
host.wait-until/wait-for/wait-croncall schedules a timer that pushes aTimerevent at the deadline.
Every pump holds a cancel signal keyed by its UUID handle: provider.cancel,
host.cancel-timer, and the tooling.unsubscribe-* calls can drop it early,
stopping further deliveries before the source naturally ends.
Pull-style calls (list-models, list-tools, list-resources) are far
shorter, so the host runs them to completion with rt.block_on instead of
spawning a pump. Both approaches run off the wasm thread — pump tasks on the
tokio runtime, blocking calls on the spawn_blocking thread the engine runs on
— so the synchronous engine never blocks a tokio worker.
Endpoint requests
Endpoint requests are different: the endpoint server
routes an inbound chat completion straight into the owning agent’s inbox as an
endpoint-message event, and the agent streams deltas back through the session
registry. Each session ends exactly once with an endpoint-session-end event.
Why it is shaped this way
Keeping a single inbox per agent means the brain’s scheduling does not live in the host. The agent decides, iteration by iteration, which events to handle and in what order — the host just guarantees that everything relevant eventually shows up, tagged, in order on one queue. That is what lets the brain (whether a hand-written wasm component, a rhai script, or a js script) be written as a plain sequential program over a stream of facts.
The host interface
The host interface (host in src/lib/omw/wit/omw.wit) exposes the static,
baked-in capabilities of the runtime to an agent brain: logging, agent identity,
timer helpers, inter-agent messaging, event receipt, and UUID generation. It is
imported by every brain (wasm components and the bundled rhai / js interpreters
alike) and implemented 1:1 by the runtime’s runtime::host module.
This page describes the WIT surface from the guest’s point of view. The actor mechanics that back it are covered in actor.
Events
The unit of everything the guest can observe is an event-envelope, a record
with two fields:
id— the UUID handle of the subscribed source the event came from, andevent— one of the following variant payloads:
| variant | payload | meaning |
|---|---|---|
message(string) | the text | a message from a subscribed agent |
error(string) | the error text | a failed I/O surfaced to the guest |
timer | — | a timestamp / duration / cron timer fired |
reload | — | the brain script changed; exit so the run restarts |
shutdown | — | the process is shutting down; exit terminally |
chat-delta(chat-delta) | a stream chunk | a chat-stream delta |
chat-end | — | an open chat stream finished |
tool-result(tool-result) | { name, arguments, value } | a queued tool invocation returned |
resource-list-updated | list<resource-info> | a subscribed resource list changed, with the new list |
resource-updated | resource-content | a subscribed resource updated in place, with freshly read content |
endpoint-message(endpoint-message) | { session, messages, tools, params? } | an inbound endpoint chat request routed to a subscribed agent |
endpoint-session-end(endpoint-session-end) | { session, error? } | an endpoint session ended: normal or abrupt |
A chat-delta carries content, reasoning, a tool-call, a finish-reason,
and a usage block, all optional, so a chunk may carry text, reasoning, a
partial tool call, token counts, or a terminal reason.
A tool-result event’s payload carries the tool’s name, its arguments, and
its value — the text result queued call-tool returned.
A resource-updated event’s resource-content carries the resource’s uri, an
optional mime-type, and the content itself — actual text for textual
formats, base64 for anything else (match on mime-type to tell which).
An endpoint-message event payload carries the endpoint session’s session id,
the chat history as messages (a chat-message per entry), tools the tools
the client advertised, and params — the opaque JSON string of any extra
generation settings the client submitted. An endpoint-session-end payload
carries the session id and an optional error when the session was
interrupted.
The guest correlates an envelope with a specific source by matching id against
the UUID the opening call returned — for example the UUID from
provider.chat-stream, a host.wait-* call, tooling.call-tool, or
tooling.subscribe-*.
Message flow
subscribe-agent(agent)— subscribe to messages from another agent.unsubscribe-agent(uuid)— cancel a subscription by itssubscribe-agentUUID.subscribe-lifecycle()— subscribe to lifecycle events (reload,shutdown, and reload-failureerror); returns a UUID handle they arrive tagged with. Errors on a second subscribe (one per run).unsubscribe-lifecycle(uuid)— drop the lifecycle subscription; a foreign UUID is a no-op.send-agent(agent, payload)— send text to another agent. The message only lands in the recipient’s inbox if it subscribed to the sender, tagged with that subscription’s UUID.recv()— blocking receive of the next event from this agent’s single inbox, with a host-side timeout (seerecv_timeout_secsin tunables). Returns anevent-envelopeor an error.try-recv()— non-blocking poll of the next event; returnsnonewhen the inbox is empty.
Correlate a lifecycle event by matching id against the UUID
subscribe-lifecycle returned, and kind for reload / shutdown / error
(a reload-failure error means the edit was invalid and the live run kept
going).
The endpoint
The optional endpoint server lets each agent address itself as an OpenAI-compatible model.
The guest side is three calls:
subscribe-endpoint(model)— subscribe this agent to the endpoint under the model namemodel; returns a UUID handle. Inbound requests for that model arrive asendpoint-messageevents tagged with it, and the model is listed on/v1/modelswhile subscribed. Errors if the model is already taken.unsubscribe-endpoint(uuid)— drop the model from/v1/models, stop routing, and abruptly end every in-flight session of that subscription (each fires anendpoint-session-endevent with an error).stream-endpoint(session, delta)— stream onechat-deltato an endpoint session. Non-blocking: it buffers into the session’s local queue and returns immediately. The reply ends when a delta carries afinish-reason; a session ends exactly once (delivering anendpoint-session-endevent).
Timers
omw uses unsigned 64-bit ticks (milliseconds since the Unix epoch) as its
timestamp type. The guest gets a set of pure helpers plus three scheduling
calls:
time-now()— current time in ticks.time-format(ts, format)— format a tick with a strftime-style format.wait-until(ts)— wait until a future timestamp fires; errors iftsis not in the future.wait-for(ms)— wait formsmilliseconds.wait-cron(spec)— wait until the next fire of a cron spec.cancel-timer(uuid)— cancel a pending wait by the UUID itswait-*call returned.sleep-for(ms)— blocking wait formsmilliseconds; returns once the delay elapses. Unlikewait-for, notimerevent is scheduled.sleep-until(ts)— blocking wait until a future timestamp fires; errors iftsis not in the future. Unlikewait-until, notimerevent is scheduled.sleep-cron(spec)— blocking wait until the next fire of a cron spec. Unlikewait-cron, notimerevent is scheduled.
Each wait-* call returns a UUID immediately; when the deadline passes, a
timer event tagged with that UUID is delivered to the inbox. The brain reads
it back with recv/try-recv and matches id to know which timer fired. A
pending wait can be cancelled at any time with cancel-timer(uuid). The
sleep-* variants are the blocking mirror — they hold the brain until the wait
finishes and return directly (no timer event, no cancel handle, and an error
is reported in-band).
Identity
whoami()— the calling agent’s configured name. It lets a brain shared by several agents tell which one it is running as, for example to subscribe to its own name or pick a role.
Logging
log(level, message)— write a structured log line.levelis one oftrace,debug,info,warn, orerror, and unknown levels default toinfo. The calling agent’s name is attached as a structured field.
UUIDs
new-uuid()— a fresh v4 UUID string. Every handle used across the host (subscriptions, streams, timers) is one of these. The guest can also use it for its own purposes.
Base64
base64-encode(bytes)— encode raw bytes as standard padded base64 (RFC 4648 §4), matching the encoding of MCPblobresource contents.base64-decode(data)— decode standard padded base64 back to raw bytes. Errors on invalid input.
Memory
Per-agent string store that survives hot reloads:
memory-get(key)— read a value; none when absent.memory-set(key, value)— store a value, overwriting.memory-remove(key)— delete; true when a value was present.
Scoped to the calling agent, so agents cannot race each other. Treat entries like variables: subscription handles, state-machine state, small checkpoints. Not a database — keep values small.
The endpoint interface
An endpoint is an abstraction over an inbound server: something agents
subscribe to under model names, and which routes outside requests into their
inboxes. Unlike providers, tooling, and runtimes there is at most one endpoint
per process, and it is optional: the [endpoint] config block holds a kind
string plus opaque params, exactly like a single [providers.<name>] entry.
When [endpoint] is absent no server is started, and subscribe-endpoint
simply errors.
The abstraction
An endpoint exposes, through the endpoint::Endpoint trait:
kind()— which implementation this is (e.g.openai).serve(bus, registry)— own the serve loop: accept outside requests, open a session in the sharedEndpointRegistry, and route each request as anendpoint-messageinbox event.
build(kind, params) dispatches on kind to openai (behind the
endpoint-openai feature) and bails on anything else. The transport state
itself stays outside the trait: host/bus.rs (endpoint_subscribe,
endpoint_route, endpoint_models) owns model-to-agent routing, and
host/endpoint.rs (EndpointRegistry open / push / abort) owns the
per-session buffers — so a new transport only implements serve, never the
inbox protocol.
Configuration
[endpoint]
kind = "openai"
listen = "127.0.0.1:8080"
The endpoint is optional and disabled unless [endpoint] is set. The remaining
keys are opaque to the config layer and validated by the implementation at
construction time; see the OpenAI endpoint for the openai keys.
A build without the matching feature fails at startup when [endpoint] names
its kind.
The agent side
An agent becomes a model by subscribing: host.subscribe-endpoint(model)
returns a UUID handle, lists the model on the endpoint, and routes inbound
requests as endpoint-message events tagged with it; an agent may subscribe
many names, and host.unsubscribe-endpoint(uuid) drops a model and abruptly
ends every in-flight session of that subscription. The agent streams deltas back
with host.stream-endpoint(session, delta), non-blocking; a delta carrying a
finish-reason ends the session. Sessions end exactly once: normally when the
reply completes, or abruptly on client disconnect or unsubscribe. Each end fires
exactly one endpoint-session-end event into the owning agent’s inbox. Errors
on unknown, foreign-agent, or ended sessions are reported in-band.
Session lifecycle
The session buffer (see session_buffer in tunables) holds
deltas before further chunks are dropped with a tracing::warn — an emergency
lane, not a throttle. A terminal finish-reason delta is queued first, then the
session entry is removed, then a Close marker, and exactly one
endpoint-session-end fires. The host fns and events are documented in
host.
The OpenAI endpoint
The openai endpoint (endpoint::openai, kind openai) is a small axum server
exposing /v1/models and POST /v1/chat/completions. It is compiled behind the
endpoint-openai cargo feature, which is on by default; a build without it
fails at startup when [endpoint] names kind = "openai".
Configuration
| key | type | default | meaning |
|---|---|---|---|
listen | string | — | socket address to listen on, required |
[endpoint]
kind = "openai"
listen = "127.0.0.1:8080"
The listen address is parsed as a plain socket address (numeric host:port),
so hostnames are rejected at startup. Set it to "0.0.0.0:8080" to serve all
interfaces. Only one listener is supported.
The HTTP surface
GET /v1/models lists every currently subscribed model name, in sorted order.
POST /v1/chat/completions accepts an OpenAI-style request. stream: true
delivers SSE deltas ending with data: [DONE]; stream: false buffers deltas
into one JSON chat.completion. tools (OpenAI tool schema) is optional.
Errors — unknown roles, malformed tools, unknown models — are reported as OpenAI
error responses with matching status codes.
Any other request fields the client sends (temperature, max_tokens,
reasoning_effort, …) are collected verbatim and forwarded to the owning agent
as the endpoint-message’s opaque params JSON string, so a brain can pass
them straight to its own provider. Reasoning the agent streams back is relayed
as reasoning_content in both the SSE deltas and the buffered message.
The provider interface
A provider is an abstraction over an OpenAI-family chat service: something you
hand a model, a conversation, and (optionally) the tool signatures the model may
call, and which streams back deltas. Named providers live in the global
[providers.<name>] config map and are looked up at runtime by name.
The abstraction
A provider exposes, through the WIT provider interface:
kind()— which implementation this is (e.g.openai), letting a guest break the abstraction when it chooses to.name()— the configured name of the instance.list-models()— the model names this provider exposes; errors if they cannot be enumerated.chat(model, messages, tools, params)— run a chat conversation to completion, and return the full [chat-result] in-band: the concatenated content, the reassembled tool calls, the terminal finish reason, the concatenated reasoning/thinking content, and tokenusage. No events are delivered; the call blocks the brain until it finishes or errors, and cannot be cancelled.chat-stream(model, messages, tools, params)— open a streaming chat response. Returns a UUID handle; deltas flow into the agent’s inbox aschat-deltaevents until a terminalchat-end(orerror) event closes the stream.is-open(uuid)— whether a chat stream identified byuuidis still open.cancel(uuid)— cancel an open stream byuuid.
params is an optional opaque JSON object of generation settings
(temperature, max_tokens, reasoning_effort, response_format, …). It is
forwarded to the provider and merged over the provider’s configured defaults, so
a brain can set per-call settings without a config change. An implementation
ignores keys it does not understand, and the mandatory request fields (model,
stream, messages, tools) always win.
The handle is obtained once with provider.get(name), which returns a
provider resource; all further calls go through that handle so the guest never
repeats the name.
Reasoning and usage
Reasoning models stream a second channel of output alongside the answer. omw
carries it explicitly:
chat-messagehas areasoningfield. A brain can echo an assistant message’s reasoning back to the provider on a later turn (some models require this for multi-turn tool use).chat-deltahas areasoningfield for each streamed reasoning chunk, plus ausagefield (many providers report token counts only on the final chunk).chat-resultconcatenatesreasoningand carries the response’susage.
usage reports prompt-tokens, completion-tokens, and total-tokens, each
optional because not every provider reports every count.
The streaming contract
chat-stream is the only long-running call, and it drives the two shapes of the
host bridge at once:
- Open and return. The call starts a chat-stream pump on the bridge runtime and returns the stream’s UUID immediately; it does not block the brain.
- Consume in the inbox. The pump delivers each chunk as a
chat-deltaevent, then achat-endevent, all tagged with the returned UUID. The brain collects them withrecv/try-recv.
Two contracts matter when writing or using a provider implementation:
- Implementations must return an error before the first delta on transport
or authentication failure, rather than a silent empty stream — the guest sees
the failure as an
errorevent instead of a misleadingchat-end. - Dropping the returned stream aborts the in-flight request, unsubscribing any pump reading it. This is how cancellation is granted for free.
The blocking call
chat is the in-band counterpart of chat-stream: it runs the same
conversation, but collects every delta to completion itself, and returns the
accumulated result instead of delivering events. A brain that wants a simple
round-trip without stream bookkeeping can use chat and read
content/reasoning/tool_calls/finish_reason/usage off the returned
[chat-result].
The OpenAI provider
The openai provider (provider::openai, kind openai) talks to any
OpenAI-family HTTPS endpoint that exposes the chat completions API, streaming
server-sent events (SSE). It is built on reqwest and requires no extra
services.
Configuration
| key | type | default | meaning |
|---|---|---|---|
base_url | string | https://api.openai.com/v1 | API base, before /chat/completions and /models |
api_key | string | unset (no auth) | sent as a Bearer token |
model | string | unset | model returned when the endpoint’s list-models() is empty |
params | table | unset | default generation params merged into every request body |
All keys are optional. The api_key is never logged: it is a Secret (locked
with mlock, zeroized on drop) that redacts on Debug and serialize, and omw
fails at startup if the lock cannot be taken.
params is an opaque table of generation settings — temperature,
max_tokens, reasoning_effort, top_p, response_format, and anything else
the endpoint accepts — merged into every request body. Per-call params from
the brain override these defaults, and the mandatory fields (model, stream,
messages, tools) always win over both.
[providers.openai]
kind = "openai"
api_key = "sk-…" # usually sourced from your environment at runtime
model = "gpt-4o"
params = { temperature = 0.2, max_tokens = 1024, reasoning_effort = "high" }
list-models() asks GET {base_url}/models and errors if the request fails
(unreachable endpoint or non-2xx), so a brain calling it sees the failure. When
the endpoint answers with an empty list it falls back to the configured model.
The omw scaffold command treats a failure as an empty list.
Reasoning and usage
The provider reads reasoning output from both spellings endpoints use —
reasoning and reasoning_content — surfacing it on each chat-delta’s
reasoning field and concatenated on chat-result.reasoning. A reasoning
field on an outgoing assistant message is sent back as reasoning. Token counts
are read from any chunk’s usage (many endpoints only send it on the final
chunk, or after a stream_options.include_usage request); each delta and the
result carry a usage block with prompt_tokens, completion_tokens, and
total_tokens when reported.
How a chat stream works
chat sends a POST {base_url}/chat/completions with stream: true, the
model, the conversation, the tools (when non-empty), and any configured/call
params. Non-2xx responses are returned as an error before any delta —
satisfying the interface’s streaming contract. On success the response body is
decoded line by line:
- lines are split on newlines and stripped of their
data:prefix; - a
[DONE]marker ends the stream with a finalchat-end, - each JSON chunk contributes one
deltaevent.
Tool-call reassembly
OpenAI streams tool-call arguments in fragments. The provider accumulates
arguments per tool-index and only surfaces a tool-call once both its id
and name are known; when all fragments have arrived it emits the fully
reassembled call. The lowest tool index is surfaced first, keeping order
deterministic.
OpenAI vs. anything else
Because the interface is just “an OpenAI-family chat stream”, openai is the
only provider compiled into the binary by default. Alternative endpoints with
the same wire shape work by pointing base_url at them; anything genuinely
different would be a new provider kind.
The tooling interface
A tooling is an abstraction over an MCP-style tool server: it exposes callable
tools and readable resources. Named toolings live in the global
[tooling.<name>] config map and are looked up at runtime by name.
The abstraction
A tooling exposes, through the WIT tooling interface:
kind()— which implementation this is (e.g.mcp).name()— the configured name of the instance.list-tools()— every tool visible on the instance.call-tool(name, arguments)— queue a tool invocation, returning a UUID handle. The result arrives as atool-resultevent (or anerrorevent on failure) tagged with that UUID.is-open(uuid)— whether a queued tool call is still open.cancel(uuid)— cancel a queued tool call by UUID, dropping its pending result delivery.call-tool-blocking(name, arguments)— invoke a tool by name with opaque JSON arguments, blocking until the result is ready. Returns the tool’s result as atool-result(name,arguments,value) in-band (errors are surfaced as theerr).list-resources()— every URI-addressed resource the tooling exposes.read-resource(uri)— block and read one resource’s current content — it returns aresource-content(uri, optionalmime-type, andcontentwhich is text for textual formats and base64 for anything else).subscribe-resource-list()— subscribe to the resource list changing; returns a UUID handle tagged on eachresource-list-updatedevent, which carries the freshly fetched resource list.unsubscribe-resource-list(uuid)— cancel a resource-list subscription by that UUID.subscribe-resource(uri)— subscribe to one resource’s updates; returns a UUID handle tagged on eachresource-updatedevent, which carries the freshly read resource content ({ uri, mime_type?, content }).unsubscribe-resource(uuid)— cancel a single-resource subscription by that UUID.
The handle is obtained once with tooling.get(name), which returns a tooling
resource; all further calls go through that handle.
Tools
A tool has a name, an optional description, and an input-schema — a JSON
Schema describing the arguments the model must supply. The guest hands the
signature to a provider so the model can emit a tool-call for it, then invokes
it with call-tool, or synchronously with call-tool-blocking.
Resources
A resource-info is a single URI-addressed, readable value: it carries a uri,
a programmatic name, an optional description, and an optional mime-type.
Resources are how a tooling surfaces read-only data (a file, a database row, a
metric) without it being an instruction-following tool.
A resource-content is the content of one such resource read at the moment its
update fired. It carries the uri, an optional mime-type, and the content
itself: the content holds actual text for textual formats and base64-encoded
bytes for anything else. Match on mime-type to tell the two apart — a text/*
(or absent) type means plain text, anything else means base64.
Subscriptions
Resource list subscriptions and single-resource subscriptions are both
long-lived streams. As with chat streams, the host spawns a resource pump on the
bridge runtime that pushes resource-list-updated / resource-updated events
into the agent inbox, tagged with the UUID the subscribe call returned. A
resource-list-updated event carries the resource list fetched at change time
(an entire list<resource-info>), so the guest never has to call back to
list-resources; a resource-updated event carries the resource content read
at update time (a resource-content), so the guest never has to read the
resource back itself. Each subscription can be cancelled early with
unsubscribe-resource-list / unsubscribe-resource by that UUID. Delivery
contract: dropping the returned stream cancels the subscription.
The streaming contract
subscribe-resource-list and subscribe-resource mirror the provider’s
chat-stream contract: they return a UUID immediately, and events arrive later
through the inbox. The guest matches the envelope id against the UUID of the
subscription it wants to hear about. call-tool follows the same shape: it
returns a UUID held for the queued invocation, and the tool-result event
arrives later under that UUID (call-tool-blocking is instead the synchronous
mirror, returning the result in-band).
The MCP tooling
The mcp tooling (tooling::mcp, kind mcp) is a client for the Model Context
Protocol, built on the official rmcp Rust SDK. It owns the wire protocol and
the initialize lifecycle itself, mapping rmcp’s typed results onto omw’s
tool / resource-info types and joining text tool results.
Transports
| transport | config keys | speaks over |
|---|---|---|
stdio | command, args, env | a server subprocess, JSON-RPC |
http | url, auth_token | a streamable-HTTP MCP endpoint |
Configuration
| key | type | transport | meaning |
|---|---|---|---|
transport | string | both | stdio or http |
command | string | stdio | the server executable |
args | list | stdio | extra arguments |
env | attrs | stdio | extra server environment |
url | string | http | the endpoint URL |
auth_token | string | http | sent as a Bearer token |
The auth_token and env values are never logged: they are Secrets (locked
with mlock, zeroized on drop) that redact on Debug and serialize, and omw
fails at startup if the lock cannot be taken.
[tooling.mcp]
kind = "mcp"
transport = "stdio"
command = "npx"
args = ["-y", "@modelcontextprotocol/server-everything"]
Connecting
The build only parses and validates config; the server is dialed lazily on the
first tool or resource use. The first use retries with exponential backoff
(doubling from tooling_connect_backoff_start_ms up to
tooling_connect_backoff_cap_secs, see tunables). A failed
first use surfaces as Err for blocking callers and as the error event
variant for event-driven callers; dropping the call (pump cancel, reload, or
shutdown) cancels the wait. Tool results keep only their text content blocks,
joined with newlines; binary content is dropped.
Resources
mcp surfaces every resource the peer advertises via list-resources. The two
subscription modes map onto rmcp subscription filters:
subscribe-resource-list()subscribes to list-changed notifications and deliversresource-list-updatedevents carrying the freshly fetched list;subscribe-resource(uri)subscribes to a single resource and, at each update, the host reads the resource back (resources/read) and delivers aresource-updatedevent carrying the content. Textual resource content is passed through as-is; binary content arrives base64-encoded, so match on the content’smime-typeto decide whether to decode it.
Both return BoxStreams; the host’s resource pump drains them and pushes tagged
events into the agent inbox, and dropping the stream cancels the subscription.
The runtime interface
A runtime is how an agent’s “brain” is loaded and driven for one iteration.
Where providers and tooling are I/O, a runtime is the program the agent runs:
it is handed the agent’s brain (a wasm component, a rhai script, or a js script)
and asked to execute it against the AgentContext.
Named runtimes live in the global [runtime.<name>] config map. Each agent
wiring pins itself to one by name.
The abstraction
A runtime exposes, through the WIT runtime interface (exported by the guest
component and called by the host):
kind()— which brain implementation this is (wasm,rhai, orjs).run(script)— run one iteration. Returns the terminal message if the brain chose to exit, otherwise nothing.
The host’s Runtime::run(&AgentContext) drives this off the tokio worker on a
spawn_blocking thread (the wasm engine and the rhai / js interpreter
components built on it are synchronous). The script argument is the agent’s
brain. For a wasm brain the program is baked into the component and the script
is unused, while for Rhai it is the Rhai source text and for JS it is the
JavaScript source text.
How a run happens
build(kind) dispatches to wasm, rhai, or js. Each runtime loads its
engine, and pushes a blocking task that:
- builds a
Storewhose data is the runtimeHost(the agent context, a resource table, and a WASI context), - wires the imports — the WASI wasip2 imports plus the
provider,tooling, andhostinterfaces — into aLinker, - instantiates the component,
- calls the exported
runtime.run(script)and - surfaces the terminal message as
Exited(message)orCompletedif the brain ran to a finish.
The agent context
Every runtime call receives an AgentContext: the agent name, the brain
script path, the named providers and tooling registries, the shared
MessageBus, the agent’s own StreamRegistry (chat streams), the timer
CancelRegistry, the resource-subscription CancelRegistry, and the dedicated
tokio rt used to bridge synchronous host calls to the async provider/tooling.
Provider and tooling registries, and the message bus, are shared across all
agents in one process — the StreamRegistry, timer/resource registries, and
bridge runtime are per-agent.
The wasm runtime
The wasm runtime (runtime::wasm, kind wasm) loads the agent’s brain as a
compiled WebAssembly component implementing the exported runtime interface.
This is the general-purpose path: because the agent’s program is baked into the
component, the brain is fully portable and the host never sees its logic.
Configuration
The wasm runtime’s brain is a file named by the agent’s script, plus an
optional WASI sandbox (deny-by-default, like WasiCtxBuilder):
[runtime.wasm]
kind = "wasm"
inherit_env = true
env = { FOO = "bar" }
args = ["--flag"]
initial_cwd = "/work"
[[runtime.wasm.preopens]]
host_path = "./data"
guest_path = "/data"
perms = "read_write" # or "read_only" (default)
Available keys: inherit_stdio (shorthand for all three below), inherit_stdin
/ inherit_stdout / inherit_stderr, inherit_env, env, inherit_args,
args, initial_cwd, preopens, allow_blocking_current_thread,
insecure_random_seed, max_random_size, allow_tcp / allow_udp /
allow_ip_name_lookup, and inherit_network (enables all three network flags
with a permissive address check). The whole block is per named [runtime.*]
entry: to give another agent different sandboxing, declare another runtime and
point the agent at it.
[runtime.wasm]
kind = "wasm"
[[agents]]
name = "server"
runtime = "wasm"
script = "brain.wasm"
The file may be .wat (text), .wasm (binary), or .cwasm (AOT-compiled, also
the fastest to load). The component model is enabled, and the component must
export the omw.runtime interface.
Rust brains
A pure Rust brain is a library crate depending on the omw-wasm-rust guest SDK,
which re-exports the WIT bindings plus typed Provider/Tooling handles,
host helpers, and RAII guards. There is no single-file support: keep the brain
a real crate so rust-analyzer keeps working.
# Cargo.toml
[lib]
crate-type = ["cdylib"]
[dependencies]
omw-wasm-rust = "0.1"
#![allow(unused)]
fn main() {
// src/lib.rs
#![no_main]
omw_wasm_rust::brain!(|| {
omw_wasm_rust::host::info("hello from a rust brain");
Ok(())
});
}
Build it for wasm32-wasip2 and point this runtime at the component:
cargo build --target wasm32-wasip2
[runtime.wasm]
kind = "wasm"
[[agents]]
name = "server"
runtime = "wasm"
script = "target/wasm32-wasip2/debug/brain.wasm"
The engine
The shared WasmEngine (in runtime::engine) is deliberately generic: it has
no knowledge of any particular brain. It:
- loads the component from the file (WAT, WASM, or AOT-cached),
- builds a
Storeover the hostHost(agent context + resource table + WASI context), - wires the host imports into a
Linker— the WASI wasip2 imports and theprovider/tooling/hostinterfaces, - instantiates the component,
- calls
runtime.run(script)and returns the terminal message.
Because it is synchronous, the engine runs on a spawn_blocking thread rather
than a tokio worker, so the host imports’ use of rt.block_on stays legal.
WasmEngine is Clone, so one loaded component is reused across iterations.
The rhai runtime
The rhai runtime (runtime::rhai, kind rhai) lets an agent’s brain be a
rhai script instead of a hand-written wasm component. It evaluates the script
on the bundled rhai interpreter (a wasm component compiled into the host at
build time whose omw.* host imports route to the very same global provider /
tooling / bus as every other runtime).
The interpreter ships in the omw-rhai package/binary (nix run .#omw-rhai or
the omw-rhai-<arch>.tar.gz release tarball). The default omw binary doesn’t
include it. With the --features runtime-rhai build flag it is compiled into
the host at build time instead.
Configuration
The rhai runtime takes no required parameters beyond the shared WASI sandbox:
[runtime.rhai]
kind = "rhai"
[[agents]]
name = "alice"
runtime = "rhai"
script = "brain.rhai"
A custom interpreter component can be substituted via the runtime’s
interpreter parameter; otherwise the interpreter compiled into the binary is
used. The same WASI sandbox keys as the wasm runtime apply, flattened
alongside interpreter:
[runtime.rhai]
kind = "rhai"
inherit_env = true
env = { FOO = "bar" }
[[runtime.rhai.preopens]]
host_path = "./data"
guest_path = "/data"
perms = "read_only"
The interpreter and the WIT bindings
The bundled guest (omw-wasm-rhai-interpreter) exports the runtime interface
(kind returns rhai, run(script) evaluates the script) and imports the
omw world. On startup it registers an omw static module with three
sub-modules that expose the WIT interfaces to the script:
omw::provider::get(name)— returns a provider handle map whose blockingchat, streamingchat_stream,is-open,cancel,list_models, andkindentries are methods.chat/chat_streamtake an optional trailingparamsargument: either a map of generation settings (#{ temperature: 0.2, reasoning_effort: "high" }) or an already-JSON string (e.g. anendpoint-message’sparams), merged over the provider’s configured defaults.omw::tooling::get(name)— returns a tooling handle map whoselist-tools,call-tool,is-open,cancel,call-tool-blocking,list-resources,read-resource,subscribe-resource-list,subscribe-resource,unsubscribe-resource-list,unsubscribe-resource, andkindentries are methods.omw::host::*— the host helpers:log,time_now,time_format,wait_until,wait_for,wait_cron,cancel_timer,subscribe_agent,unsubscribe_agent,subscribe_lifecycle,unsubscribe_lifecycle,subscribe_endpoint,unsubscribe_endpoint,stream_endpoint,send_agent,recv,try_recv,new_uuid,base64_encode,base64_decode,memory_get,memory_set,memory_remove,sleep_for,sleep_until, andsleep_cron.
Handles are Rhai maps. Methods are FnPtrs stored on them, so scripts call them
method-style (provider.chat_stream(...), tooling.call-tool(...)). The time
functions take plain integer literals: the interpreter converts rhai’s i64
integers to the WIT u64 tick type at the boundary (rejecting negatives).
Values in rhai
Events come back as maps shaped #{ id, kind, payload }:
id— the envelope’s UUID;kind— one ofmessage,error,timer,chat-delta,chat-end,tool-result,resource-list-updated,resource-updated,endpoint-message,endpoint-session-end;payload— the text formessage/error, a map forchat-delta(withcontent,reasoning,tool_call{ id, name, arguments },finish_reason, andusage{ prompt_tokens?, completion_tokens?, total_tokens? }), a map fortool-result({ name, arguments, value }), a list of resource maps ({ uri, name, description?, mime_type? }) forresource-list-updated, a resource-content map ({ uri, mime_type?, content }) forresource-updated, a map forendpoint-message({ session, messages, tools, params? }, withmessagesa list of{ role, content?, reasoning?, tool_call? }maps,toolsa list of{ name, description?, input_schema }, andparamsthe opaque JSON string the client submitted), a map forendpoint-session-end({ session, error? }), and unit otherwise. Thecontentfield holds actual text for textual formats and base64 for anything else — match onmime_typeto tell which. Decode binary payloads withomw::host::base64_decode(which returns a blob) and encode back withomw::host::base64_encode.
Example brain
let p = omw::provider::get("openai");
p.chat_stream("gpt-4o", [
#{ role: "user", content: "say hi" },
], []);
let out = "";
loop {
let ev = omw::host::recv();
if ev.kind == "chat-delta" { out += ev.payload.content }
if ev.kind == "chat-end" { break }
if ev.kind == "error" { throw ev.payload }
}
out
The script’s final value becomes its terminal message when it is not unit.
Memory
memory_get returns the value or unit when absent; memory_set stores;
memory_remove returns true when a value was present:
omw::host::memory_set("timer", omw::host::wait_for(1000));
// ... after a reload, the same context still has it:
let timer = omw::host::memory_get("timer");
The js runtime
The js runtime (runtime::js, kind js) lets an agent’s brain be a
JavaScript script instead of a hand-written wasm component. It evaluates the
script on the bundled js interpreter (Boa, a wasm component compiled into the
host at build time whose omw.* host imports route to the very same global
provider / tooling / bus as every other runtime).
The interpreter ships in the omw-js package/binary (nix run .#omw-js or the
omw-js-<arch>.tar.gz release tarball). The default omw binary doesn’t
include it. With the --features runtime-js build flag it is compiled into the
host at build time instead.
Configuration
The js runtime takes no required parameters beyond the shared WASI sandbox:
[runtime.js]
kind = "js"
[[agents]]
name = "alice"
runtime = "js"
script = "brain.js"
A custom interpreter component can be substituted via the runtime’s
interpreter parameter; otherwise the interpreter compiled into the binary is
used. The same WASI sandbox keys as the wasm runtime apply, flattened
alongside interpreter:
[runtime.js]
kind = "js"
inherit_env = true
env = { FOO = "bar" }
[[runtime.js.preopens]]
host_path = "./data"
guest_path = "/data"
perms = "read_only"
The interpreter and the WIT bindings
The bundled guest (omw-wasm-js-interpreter) exports the runtime interface
(kind returns js, run(script) evaluates the script) and imports the omw
world. On startup it registers an omw global with three namespaces that expose
the WIT interfaces to the script:
omw.provider.get(name)— returns a provider handle object whose blockingchat, streamingchatStream,isOpen,cancel,listModels, andkindentries are methods.chat/chatStreamtake an optional trailingparamsargument: either an object of generation settings ({ temperature: 0.2, reasoning_effort: "high" }) or an already-JSON string (e.g. anendpoint-message’sparams), merged over the provider’s configured defaults.omw.tooling.get(name)— returns a tooling handle object whoselistTools,callTool,isOpen,cancel,callToolBlocking,listResources,readResource,subscribeResourceList,subscribeResource,unsubscribeResourceList,unsubscribeResource, andkindentries are methods.omw.host.*— the host helpers:log,timeNow,timeFormat,waitUntil,waitFor,waitCron,cancelTimer,subscribeAgent,unsubscribeAgent,subscribeLifecycle,unsubscribeLifecycle,subscribeEndpoint,unsubscribeEndpoint,streamEndpoint,sendAgent,recv,tryRecv,newUuid,base64Encode,base64Decode,memoryGet,memorySet,memoryRemove,sleepFor,sleepUntil, andsleepCron.
Handles are plain objects holding the configured name plus native methods, so
scripts call them method-style (p.chatStream(...), t.callTool(...)). Method
names are camelCase only. The time functions take plain numbers: the interpreter
converts JavaScript’s f64 numbers to the WIT u64 tick type at the boundary
(rejecting negatives, fractions truncate).
TypeScript declarations
The interpreter crate ships an omw.d.ts declaration describing the omw
global. TypeScript only resolves local files, so vendor a copy next to your
brain.js and reference it:
curl -L -o omw.d.ts \
https://raw.githubusercontent.com/haras-unicorn/omw/main/src/wasm/omw-wasm-js-interpreter/omw.d.ts
/// <reference path="./omw.d.ts" />
That gives brain.js completion and checking for every provider / tooling /
host method and the { id, kind, payload } event union. The declarations are
ambient (declare const omw), so they need no import.
Values in js
Events come back as objects shaped { id, kind, payload }:
id— the envelope’s UUID;kind— one ofmessage,error,timer,chat-delta,chat-end,tool-result,resource-list-updated,resource-updated,endpoint-message,endpoint-session-end;payload— the text formessage/error, an object forchat-delta(withcontent,reasoning,tool_call{ id, name, arguments },finish_reason, andusage{ prompt_tokens?, completion_tokens?, total_tokens? }), an object fortool-result({ name, arguments, value }), a list of resource objects ({ uri, name, description?, mime_type? }) forresource-list-updated, a resource-content object ({ uri, mime_type?, content }) forresource-updated, an object forendpoint-message({ session, messages, tools, params? }, withmessagesa list of{ role, content?, reasoning?, tool_call? }objects,toolsa list of{ name, description?, input_schema }, andparamsthe opaque JSON string the client submitted), an object forendpoint-session-end({ session, error? }), andnullotherwise. Thecontentfield holds actual text for textual formats and base64 for anything else — match onmime_typeto tell which. Decode binary payloads withomw.host.base64Decode(which returns an array of bytes) and encode back withomw.host.base64Encode.
Tool and resource shapes additionally carry camelCase aliases (inputSchema,
mimeType) alongside the snake_case keys.
The script’s completion value becomes its terminal message, stringified as JSON:
a string result is returned as-is, while arrays and objects surface as [...].
undefined and null complete with no message.
Example brain
let p = omw.provider.get("openai");
let id = p.chatStream("gpt-4o", [{ role: "user", content: "say hi" }], []);
let out = "";
while (true) {
let ev = omw.host.recv();
if (ev.id === id && ev.kind === "chat-delta") {
out += ev.payload.content;
}
if (ev.id === id && ev.kind === "chat-end") {
break;
}
if (ev.kind === "error") {
throw ev.payload;
}
}
out;
The script’s final value becomes its terminal message when it is not
undefined/null.
Memory
memoryGet returns the value or undefined when absent; memorySet stores;
memoryRemove returns true when a value was present:
omw.host.memorySet("timer", omw.host.waitFor(1000));
// ... after a reload, the same context still has it:
let timer = omw.host.memoryGet("timer");
Hot reload
omw run --watch / omw loop --watch restarts an agent when its brain script
changes, without losing bus state. Only the run restarts; everything addressable
survives.
This matches the actor model in actor: the inbox is the agent, the
brain is just its current reader. Lifecycle notifications (reload, shutdown,
reload-failure error) follow the same model: explicit subscribe, UUID
correlation, no magic inbox traffic for agents that did not opt in.
Using --watch
Pass --watch to either mode:
omw run --config omw.toml --watch
omw loop --config omw.toml --watch
The watcher tracks each agent’s script path. It watches parent directories
non-recursively (editors that save via write temp + rename still trigger) and
debounces (see watch_debounce_ms in tunables), so one save
restarts the agent once.
What survives and what dies
Preserved across reload:
- Inboxes (queued events are never drained or dropped).
- Agent subscriptions and endpoint subscriptions.
- Providers, tooling, and endpoint sessions.
- Per-agent memory (
host.memory-get/memory-set/memory-remove): the sameAgentContextis reused, so stored handles and state-machine state carry over. Treat entries like variables, not a database.
Discarded on reload:
- Wasm memory/stack (a fresh
Storeper run). - Open pump tasks: chat streams, timers, resource subscriptions, tool calls.
They are cancelled so stale events cannot leak into the next run (unless
cancel_pumps_on_reload = falsein tunables, which keeps them across reload).
Rhai brains re-read the script file and wasm brains reload the component on every iteration, so the next run picks up the edit.
Validation and startup
A bad edit never kills a good run. The supervisor validates the new script
before aborting the live one, while the old run keeps executing. Only a valid
script ends the current run. An invalid script keeps the old run alive, logs a
host-side warning, and delivers an error event on the lifecycle handle (if the
brain subscribed) so it can react however it wants.
Startup is gated the same way: a broken script never produces a first iteration.
Without --watch this fails fast. With --watch the agent parks until an edit
fixing the script validates, then starts normally.
How a restart happens
Three tiers, mirroring SIGTERM/SIGKILL:
- Cooperative:
recvaborts on the next poll slice (seerecv_slice_msin tunables) without draining the inbox, blocking calls abort through a helper, and areload/shutdownevent coverstry-recvpollers. - Grace timeout (see
reload_grace_secsin tunables): the supervisor stops waiting and reports the abort. Bounds restart latency. - Preemptive (interrupt,
interrupt_budget_msafter grace — see tunables): one increment traps unyielding loops that never yield to the host.
Shutdown reuses the same three tiers, but is terminal: run and loop both
exit Ok (0) once every in-flight iteration aborted, logging the shutdown; only
a genuine failure without a shutdown request errors. The signal subscription is
process-wide (one SIGTERM/SIGINT latch awaited by every iteration, the endpoint
server, and the loop backoff), so a signal arriving between iterations still
aborts the next one. The endpoint drains gracefully: in-flight requests get an
error session-end and the server stops accepting, and the watcher pump is
aborted with the agent tasks.
Every host call parks somewhere different, so each needs its own interrupt:
| Brain is stuck in | Where it parks | Interrupt | Upstream sees |
|---|---|---|---|
host.recv | inbox slices (recv_slice_ms) | flag poll, Err("agent reloaded"/"agent shutting down"), inbox kept; plus a reload/shutdown event for pollers | nothing (pure local wait) |
host.try-recv / send / etc | returns immediately | nothing needed; pollers observe the queued reload/shutdown event | nothing |
provider.chat-stream pump | biased select! on cancel | open streams are cancelled, pump breaks | TCP close, OpenAI sees client disconnect |
host.wait-* / timers | select! on cancel | open timers are cancelled | nothing |
| resource subscription pumps | stream + cancel select | open subscriptions are cancelled | depends on transport |
tooling.call-tool pump | select! on the future | abort drops the in-flight call | MCP call dropped client-side |
blocking provider.chat | abort slices (reload_poll_ms) | abort handle, Err("agent reloaded") | normal completed request, result dropped |
blocking call-tool-blocking, list-*, read-resource, subscribe-* | same helper | same as above | server runs to completion, result dropped |
blocking sleep-for/until/cron | same helper | same as above = true cancel | nothing |
pure wasm while true {} | executing wasm, never yields to host | interrupt after grace + interrupt_budget_ms | nothing |
All values are tunables with the defaults listed there.
Writing a reload-safe brain
Subscribe to lifecycle events explicitly and correlate by UUID, like every other subscription in the actor model (see host):
let lc = omw::host::subscribe_lifecycle();
let sub = omw::host::subscribe_agent("other");
loop {
let e = omw::host::recv();
if e.id == lc {
if e.kind == "reload" { break; }
if e.kind == "shutdown" { break; }
if e.kind == "error" {
omw::host::log("warn", "reload failed: " + e.payload);
continue;
}
}
// ... handle e ...
}
omw::host::unsubscribe_lifecycle(lc);
Rules:
- Lifecycle is opt-in:
subscribe_lifecyclereturns a UUID;reload,shutdown, and reload-failureerrorevents arrive tagged with it. A second subscribe errors (one lifecycle subscription per run). Unsubscribed brains still get aborted on a valid reload (recverrors), but get no events and no invalid-edit notice. reload-failedis the plainerrorvariant, not a new event: distinguish bye.id == lifecycle_uuid && e.kind == "error"and read the validation message from the payload. The host always logs regardless, so the event is the brain’s chance to notify itself, not the only signal.- Keep subscription handles in memory so the next run can reuse them instead of re-subscribing blindly:
let sub = omw::host::memory_get("other-sub");
if sub == () {
sub = omw::host::subscribe_agent("other");
omw::host::memory_set("other-sub", sub);
}
- The live run never exits on an invalid edit. If you see
erroron the lifecycle handle, keep running. - Exit on
reload(break, do cleanup); exit terminally onshutdown.recvmay also abort with"agent reloaded"/"agent shutting down"when blocked between polls. - Prefer evented calls (
chat-stream,call-tool,wait-*) over blocking ones (chat,call-tool-blocking,sleep-*): both cancel promptly, but evented handles keep delivering while blocking ones abort the handle. - Re-subscribe at the top of the script. Handles are one-shot UUIDs that die with the run (lifecycle included). Read them back from memory when they can outlive the run; otherwise rebuild any cached handles from scratch on each run.
- Never
while true {}without a host yield; an unyielding loop can only die by interrupt. - Pollers should use
try-recv+ smallwait-for, so thereload/errorevents are observed promptly. - A broken script never starts, so the first thing a fresh brain can assume is that it compiled.
Troubleshooting
- “Reload takes a while”: the brain is stuck in a blocking call or unyielding
loop and the grace (
reload_grace_secs) expired. Prefer evented calls, or yield to the host regularly. - “Agent restarts twice”: two saves in quick succession, or a stale
reloadevent surviving into the next run. The drain at iteration start drops queued system events; check the watcher debounce vs your editor’s save burst. - “Old stream still delivers”: a pump from the previous run outlived the reload.
Reload cancels all open pumps; if you held the UUID, re-check
is-openafterreloadinstead of assuming it is alive. - “CPU spins after save”: a
while true {}without a host yield. Only the interrupt can kill it, after the grace. Add arecv,wait-for, orsleep-forto the loop. - “Edit did nothing”: the script was invalid. Check host logs for the warning
and, if subscribed, the lifecycle
errorpayload. The live run kept going; fix the script and save again. - “Agent won’t start”: the startup gate rejected the script. Same signals as
above. With
--watchthe agent parks until a fixing edit validates; without it the process fails fast.
Tunables
All timing, buffering, and backoff knobs live in one global [tunables]
section. Every field is optional; omitted fields fall back to the defaults
below.
[tunables]
inbox_bound = 1024
recv_slice_ms = 200
recv_timeout_secs = 60
reload_poll_ms = 200
interrupt_budget_ms = 100
reload_grace_secs = 5
loop_backoff_start_ms = 100
loop_backoff_cap_secs = 30
tooling_connect_backoff_start_ms = 100
tooling_connect_backoff_cap_secs = 30
watch_debounce_ms = 200
trace_buffer = 4096
session_buffer = 8192
cancel_pumps_on_reload = true
allow_unlocked_secrets = false
They can also be layered from the environment (OMW__ prefix, __ separator),
e.g. OMW__TUNABLES__RECV_TIMEOUT_SECS=30.
Reference
| Knob | Units | Default | What it does |
|---|---|---|---|
inbox_bound | events | 1024 | Per-agent inbox capacity; sends fail past it. |
recv_slice_ms | ms | 200 | How often a blocking recv checks for reload. |
recv_timeout_secs | secs | 60 | How long a blocking recv waits before timeout. |
reload_poll_ms | ms | 200 | How often blocking calls check for reload. |
interrupt_budget_ms | ms | 100 | Unwind budget after the interrupt traps a loop. |
reload_grace_secs | secs | 5 | How long the supervisor waits for a coop exit. |
loop_backoff_start_ms | ms | 100 | Backoff start for loop restarts on failure. |
loop_backoff_cap_secs | secs | 30 | Backoff cap for loop restarts (doubling). |
tooling_connect_backoff_start_ms | ms | 100 | Backoff start for MCP tooling reconnects. |
tooling_connect_backoff_cap_secs | secs | 30 | Backoff cap for MCP tooling reconnects (doubling). |
watch_debounce_ms | ms | 200 | How long the watcher coalesces one save’s events. |
trace_buffer | events | 4096 | Capacity of the omw-test trace broadcast channel. |
session_buffer | deltas | 8192 | Per-session endpoint reply buffer before drops. |
cancel_pumps_on_reload | bool | true | Cancel open pumps on reload; false keeps them. |
allow_unlocked_secrets | bool | false | Permit secrets to stay unlocked when mlock fails. |
See hot reload for how the reload knobs interact.
Library
omw can be embedded in your own crate. Parse a Config, build a Registries
value, register any custom back ends, then call run_agents or loop_agents.
Add the dependency with the back ends you want:
[dependencies]
omw = "*"
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }
The default features are runtime-wasm, provider-openai, tooling-mcp, and
endpoint-openai. The runtime-rhai and runtime-js script runtimes are
opt-in features; a --no-default-features build yields empty registries. The
CLI-only stack (clap, config, tracing-subscriber) lives in the separate
omw-cli crate, so library consumers never pull it in.
TLS setup
omw follows the standard Rust contract: features select where TLS comes from,
the final binary installs it. The provider-openai and tooling-mcp features
imply rustls, which links the ring crypto backend; reqwest is built on
rustls-no-provider, so the embedding binary must install exactly one
process-global crypto provider once before running agents:
#![allow(unused)]
fn main() {
if rustls::crypto::CryptoProvider::get_default().is_none() {
rustls::crypto::ring::default_provider()
.install_default()
.expect("another crate installed a crypto provider");
}
}
Install your own provider instead (for example aws-lc-rs) when you prefer a
different backend; omw never installs or overwrites one itself — the ring
setup lives in the omw-cli binary only. A --no-default-features build
without the provider/tooling features is TLS-free.
Certificate trust is orthogonal to the crypto backend: verification uses
rustls-platform-verifier, which on Linux loads the system CA bundle once at
startup (honoring SSL_CERT_FILE). Static binaries therefore still trust
whatever the host distribution trusts, with no Mozilla bundle baked in. The
bundled zstd (via wasmtime) needs no action either: downstream builds can
set ZSTD_SYS_USE_PKG_CONFIG=1 or unify zstd-sys features when they want the
system library instead.
Embed with defaults
use omw::prelude::*;
#[tokio::main]
async fn main() -> anyhow::Result<()> {
let raw: String = std::fs::read_to_string("omw.toml")?;
let cfg: Config = toml::from_str(&raw)?;
let registries = Registries::default();
run_agents(&cfg, false, ®istries).await
}
Registries::default() carries the feature-gated built-ins, one per family.
Registries::new() is the same shape with no built-ins. Use loop_agents to
restart agents on failure instead of running once. Config is impl-agnostic
(kind plus opaque params per entry), so unknown kind values fail at build
time with the list of registered kinds.
Running
run_agents(&cfg, watch, ®istries) runs every agent once and returns when
they all stop. loop_agents keeps every agent running forever, restarting on
success immediately and on failure with exponential backoff (the
loop_backoff_* tunables). The watch flag enables hot reload: with it, a
brain-script change restarts just the affected agent (without backoff) while the
shared bus, inboxes, and subscriptions survive.
Both have _traced variants that take a TraceSender. run_agents_traced
returns the collected Vec<TraceEvent> once the run ends (and errors if the
receiver lagged); loop_agents_traced streams events as they happen and never
returns while agents keep looping. The trace is the same stream omw-test
asserts on. Subscribe to it yourself to build a live view (a logger, a UI): the
observability library example streams every event and then verifies the same
run with check.
Testing
omw::testing is the deterministic brain-testing substrate. Assertions is the
parsed [assertions] model, with parse reading it from a config string and
collect gathering a recorded trace into something check can verify.
Harness drives a Config through the controlled run path against the
in-config kind = "mock" doubles, consuming the trace live, stopping
outcome = "asserted" agents as their assertions settle, and returning a
Report of per-agent verdicts. The omw-test binary is a thin CLI over this.
See testing for the assertion language and each mock.
Watching
omw::watch exposes the filesystem watching that powers hot reload.
Watcher::watch(path, RecursiveMode, debounce) is the general primitive: point
it at one or more paths (add more with Watcher::add) and await the next
debounced batch with next_change(). scope turns a file or directory into the
directory Watcher::watch should watch. Scripts builds on Watcher to map
each agent’s brain script to the agents running it, so a supervisor can restart
just the affected agents (next_reload() yields sorted agent names). All are
re-exported from the prelude.
The debounce window is the watch_debounce_ms tunable (see
tunables). omw-test reads its tunables from the OMW_TEST__
environment overlay and holds one Watcher across reruns.
Custom back ends
Each family (provider, tooling, runtime, endpoint) has a back-end trait,
a Factory trait with one build method, an opaque *Entry handle, a
Registry, and a register_* macro. Implement the back-end trait plus
Factory, add one macro line, reference the kind from config. Built-in impl
Config structs stay private: build deserializes the opaque
serde_json::Value itself.
Custom provider registered by type:
#![allow(unused)]
fn main() {
use omw::prelude::*;
struct MyProvider { model: String }
impl Provider for MyProvider {
fn kind() -> &'static str { "my-llm" }
async fn list_models(&self) -> Vec<String> {
vec![self.model.clone()]
}
async fn chat_stream(
&self,
model: &str,
messages: Vec<ChatMessage>,
tools: Vec<Tool>,
) -> anyhow::Result<
futures_util::stream::BoxStream<
'static,
Result<ChatDelta, String>,
>,
> {
// Call any API, map chunks to `ChatDelta`, return the stream.
// Return `Err` before the first delta on auth or transport
// failure, per the provider contract.
todo_stream()
}
}
impl omw::provider::Factory for MyProvider {
fn build(
name: &str,
params: &serde_json::Value,
) -> anyhow::Result<std::sync::Arc<Self>> {
let model: String = params
.get("model")
.and_then(|v| v.as_str())
.unwrap_or("my-model")
.to_owned();
Ok(std::sync::Arc::new(Self { model }))
}
}
// In `main`, before running agents:
let mut registries = Registries::default();
omw::register_providers!(registries.providers, MyProvider);
}
[providers mine]
kind = "my-llm"
model = "my-model"
The other families follow the same pattern with their own build shape:
tooling::Factory::build takes (name, params, tunables),
runtime::Factory::build takes (name, params), and endpoint::Factory::build
takes just (params). A test double without config plumbing uses the closure
path:
#![allow(unused)]
fn main() {
let echo: Arc<MyFakeProvider> = Arc::new(MyFakeProvider::new());
registries.providers.register_factory("fake", move |name, _params| {
Ok(echo.clone() as Arc<dyn Provider>)
});
}
Both register::<T>() and register_factory reject a duplicate kind with an
error, never overwrite.
Lazy tooling
Tooling connects lazily on first use instead of at build time, so every family
shares one sync factory shape. Connection errors surface on first tool or
resource use: blocking callers get Err, event-driven callers get the error
event variant in the agent inbox. Retries use the
tooling_connect_backoff_start_ms (default 100) and
tooling_connect_backoff_cap_secs (default 30) tunables, doubling up to the
cap. See tunables.
Prelude
use omw::prelude::*; re-exports the embedding subset: Config, AgentConfig,
ImplConfig, Tunables, Registries, the four back-end traits plus their
Factory traits aliased as ProviderFactory / ToolingFactory /
RuntimeFactory / EndpointFactory, the DTOs (Role, ChatMessage,
ChatDelta, ChatResult, ToolCall, Tool, ResourceInfo,
ResourceContent, ResourceNotification), the *Entry handles, RunOutcome,
Event / EventEnvelope, AgentContext (name() only), Secret, Shutdown,
run_agents / loop_agents, the Watcher / Scripts watcher types, plus the
register_* macros. The macros are also #[macro_export] at the crate root
(omw::register_providers!, …).
Testing
omw-test is the deterministic brain-testing binary. It runs each discovered
omw.test.toml through the same traced path the omw binary uses
(run_agents_traced), against in-process scripted doubles (kind = "mock"),
records what every agent saw and did, and checks that recording against an
[assertions] section. No keys, no network, no external services.
A test config is a normal omw.toml-shaped file named omw.test.toml: the same
config omw consumes, plus an [assertions] table that stock omw ignores the
same way it ignores unknown keys. Values can be layered from the environment
with the OMW_TEST__ prefix, exactly like OMW__ for omw itself.
Running
omw-test run examples/01-hello # one config or directory
omw-test run examples # every discovered config
<path>is a file → run that config once.<path>is a directory → recursively find every test config and run each, printingPASS/FAIL <root-relative path>and a tally. It exits non-zero if any failed.--include <glob>/--exclude <glob>(repeatable, OR within each) match the test’s root-relative path including its file name (*does not cross/,**does); include is applied first, then exclude. A missingscriptis always a failure, never a skip; narrow the set with the globs instead.--watchre-runs on change instead of exiting: after each pass it waits for a debounced filesystem event and runs again (file mode watches the config’s parent directory; directory mode watches the root recursively). The library hot-reload watch is always off.
Discovery skips hidden directories and collects any file whose name is
omw.test.toml or ends with .omw.test.toml, so several test configs can live
side by side in one directory. The shared templates (omw.test.template.toml)
never match, since they end in template.toml.
Scaffolding
omw scaffold (in the omw binary) turns a deployment config into a starter
test config: it introspects the real back ends and writes an omw.test.toml
whose provider, tooling and endpoint are the in-config mocks, pre-populated
where possible.
omw scaffold omw.toml # writes ./omw.test.toml
omw scaffold omw.toml --output t.toml # explicit output
omw scaffold omw.toml --no-resources # skip listing/reading tooling resources
omw scaffold omw.toml --force # overwrite an existing output
- the provider mock gets the models the endpoint reported (
GET /models, empty if the request fails), with an emptyturnsscript; - the tooling mock gets the server’s
tools(with their input schemas),initial_resource_list, andinitial_resource_contents(unless--no-resources), with an emptytool_callsscript; - the endpoint mock gets an empty
requestslist; runtime,agents,[memory]and[tunables]are copied through verbatim.
Everything is best-effort: a back end that cannot be built or enumerated yields
an empty mock and a warning instead of failing the conversion. The original
params are never carried over, so secrets do not end up in the output. Fill in
the turns, tool_calls and requests to script the run.
Assertions
[assertions.<agent>] is compared against the agent’s recorded trace: an
optional terminal outcome plus an ordered events list.
[assertions.alice]
outcome = "completed"
events = [
{ kind = "call", op = "chat", detail = { model = "^gpt-" } },
{ kind = "inbound", event = "chat-delta" },
]
outcomeis"completed",{ exited = "<msg>" }, or"asserted"(below).- each
eventsentry is one of:{ kind = "call", op = "...", detail = { ... } }— an outbound host call.opis exact;detailis a partial pattern over the call’s JSON detail.{ kind = "inbound", event = "...", payload = { ... } }— an inbox event.eventis the kebab-case kind (chat-delta,chat-end,tool-result,endpoint-message,message,timer,reload,shutdown,error, …);payloadis a partial pattern over the serialized event.{ "$while" = { kind = "call", ... } }— greedily consume a run of matching trace events, stopping at the first non-match.{ "$until" = { kind = "call", ... } }— skip trace events until one matches, consuming it.
The list is matched as an ordered subsequence over partial patterns:
- Only the events you write are checked, in order; anything between them is
ignored, and trailing events are fine. So
events = [{ kind = "call", op = "chat" }]passes as soon as the firstchatis seen, even if the brain does a hundred things afterwards. - A pattern object matches when every key it names is present and matches in the
candidate; extra candidate keys are ignored. String leaves are regular
expressions matched against the candidate string, so
"^gpt-a.*"is a regex. Numbers, booleans and null are compared for equality. - Arrays match as ordered subsequences too, with the same rules as
events. Unlisted elements between matches are skipped and leading/trailing elements are ignored, so[ "a", "b" ]matches[ "x", "a", "b", "y" ]. An empty pattern array matches any array. - Use
detail/payloadto pin only the fields you care about, and the$while/$untilsentinels to step over look-alikes or force the cursor forward.
$while / $until sentinels
The events list and every array inside a pattern share one vocabulary.
{ "$while" = P } greedily consumes a run of consecutive elements matching the
inner P, stopping at the first non-match (zero-or-more). { "$until" = P }
skips ahead to the first element matching P and consumes it. Under events,
P is a call/inbound assertion; inside a pattern array it is an ordinary
pattern. $-prefixed keys are reserved and never mean a partial-match field.
[assertions.alice]
events = [
{ kind = "call", op = "chat" },
{ "$while" = { kind = "inbound", event = "chat-delta" } },
{ "$until" = { kind = "call", op = "chat" } },
]
An empty object under either sentinel is an ordinary pattern that matches any
object, so { "$until" = {} } consumes the next object and { "$while" = {} }
consumes a run of objects (objects only: {} does not match a primitive). An
invalid regex fails at parse time with the offending pattern.
Chat detail
chat and chat_stream calls trace { provider, model, messages, tools },
serializing the conversation the brain sent. Together with array subsequence
matching this expresses “many messages, only the tail matters”:
detail = { messages = [
{ role = "system" },
{ "$while" = { role = "assistant" } },
{ role = "user", content = "final" },
] }
detail stays partial, so naming only model (or nothing at all) still
matches.
Seeded memory
Top-level [memory.<agent>] seeds the named agent’s memory before its brain
first runs, so a test can fast-forward an agent to an interesting state instead
of walking it there:
[memory.alice]
handle = "seed-42"
The brain reads it with memory_get like any other value, and seeded entries
persist exactly like memory written at runtime (including across hot reloads).
This is a first-class omw feature, not a test-only one.
outcome = "asserted"
A test usually cares about a prefix of a run, not its terminal outcome. If the
brain loops or waits forever, making it exit on its own is awkward and easy to
hang. outcome = "asserted" means: stop this agent as soon as its events
settle, and fail it on the first mismatch. The assertion verdict is the
result.
Stopping is per-agent, so one agent can be checked in isolation while others run normally. When every asserted agent has a verdict the harness forces the whole run down, so a brain that loops or blocks can never hang a test.
Because a $while is zero-or-more, a trailing $while settles as soon as
its prefix does (immediately for an assertion that is only a $while, exactly
like an empty events list), so it asserts nothing on its own. Bound a run you
care about with a following anchor — $until or a plain assertion — which is
also what makes the run meaningful under outcome = "asserted":
[assertions.alice]
outcome = "asserted"
events = [
{ "$while" = { kind = "inbound", event = "chat-delta" } },
{ kind = "inbound", event = "chat-end" },
]
Mock back ends
The mock cargo feature backs the deterministic doubles. omw-test is built
with just the mocks, so a test config wires kind = "mock" for its provider,
tooling and endpoint:
- Provider mock — sequenced chat turns.
- Tooling mock — canned tool results and resources.
- Endpoint mock — scripted client requests, optionally gated on the trace.
Tracing
omw-test asserts on what agents actually saw and did, through the library’s
trace channel (host/trace.rs, exported via the prelude):
#![allow(unused)]
fn main() {
pub enum TraceEvent {
Inbound { agent: String, id: String, event: Event },
Call { agent: String, op: String, detail: serde_json::Value },
Outcome { agent: String, outcome: RunOutcome },
}
pub type TraceSender = tokio::sync::broadcast::Sender<TraceEvent>;
}
run_agents_traced(cfg, watch, registries, tx) (and the loop_ twin) spawns a
receiver-drain task, emits one Outcome per agent, and returns the flattened
Vec<TraceEvent>; the omw::testing harness consumes the stream live and
groups it per agent with host::trace::group. The channel is None for
omw-cli and embedders, so it is zero-overhead when unset.
Embedders can drive the same machinery in-process through omw::testing
(Harness, Assertions, parse, check, watch) instead of shelling out to
the binary.
Provider mock
kind = "mock" is an in-process scripted chat provider. It records every chat
/ chat-stream invocation (model, messages, tools, and the opaque params the
brain passed) and returns scripted turns, so a brain’s provider traffic is fully
deterministic and inspectable.
[providers.openai]
kind = "mock"
turns = [
{ content = "first reply" },
{ tool_call = { id = "call-1", name = "echo", arguments = '{"input":"hi"}' } },
{ reasoning = "let me think", content = "the answer", usage = { prompt_tokens = 12, completion_tokens = 5 } },
]
models = ["gpt-test"]
Keys
turns— the scripted turns, popped one perchat. A turn is a table with any of:content = "..."— text; the mock emits it and a terminalstopfinish reason.reasoning = "..."— reasoning/thinking text; emitted as its own delta before the content/tool-call delta.tool_call = { id, name, arguments }— a tool call; the mock emits the call and a terminaltool_callsfinish reason.argumentsis the raw JSON string the model would have produced. Write it as a TOML string (arguments = '{"input":"hi"}') or, more readably, as the inline JSON value (arguments = { input = "hi" }); an inline value is stringified when the mock is built, so the guest always sees the wire string.usage = { prompt_tokens = 12, completion_tokens = 5, total_tokens = 17 }— token counts attached to the turn’s terminal delta (every key optional).- Once the script is exhausted the last turn repeats for every further chat, so a looping brain keeps working without re-listing the script.
models— the model nameslist-modelsreturns. Defaults to["mock-model"].
A bare kind = "mock" with no turns emits an empty stream, which is useful
for tests that only assert the call itself. Because the recorded call includes
params, an assertion can pin a brain’s per-call generation settings:
[assertions.alice]
events = [
{ kind = "call", op = "chat", detail = { params = { temperature = 0.2 } } },
]
Tooling mock
kind = "mock" is an in-process scripted MCP-style tooling. Its config maps
one-to-one onto the WIT tooling interface, with resources scriptable in order
and every step optionally gated on the trace. It also records every tool call.
[tooling.mcp]
kind = "mock"
delay_ms = 10
tools = [{ name = "echo", description = "echo back", input_schema = {} }]
tool_calls = [
{ name = "echo", result = "hi" },
{ name = "add", result = "3", after = { kind = "call", op = "call_tool" } },
]
initial_resource_list = [
{ uri = "mem://notes", name = "notes", mime_type = "text/plain" },
]
initial_resource_contents = { "mem://notes" = "v1" }
resource_list_updates = [
{
resources = [
{ uri = "mem://notes", name = "notes", mime_type = "text/plain" },
],
after = { kind = "call", op = "subscribe_resource_list" },
},
]
resource_content_updates = [
{
uri = "mem://notes",
content = "v2",
after = { kind = "call", op = "read_resource" },
},
]
Keys
tools— the toolslist-toolsreturns, asToolvalues.tool_calls— an ordered list ofcall-toolresults. Each call consumes the next entry and verifies the invoked name matches; a mismatch, running past the end, or a call with no script at all is a tool-call error (delivered to the brain, which can react to it), so the mock stays honest about call order.initial_resource_list— the resourceslist-resourcesstarts from.initial_resource_contents—read-resourcecontent keyed by URI. Reading a URI with no content (initial or applied) errors.resource_list_updates— ordered full replacement lists replayed bysubscribe-resource-list; each step replaces the current list and fires aresource-list-updated.resource_content_updates— ordered one-by-one updates replayed bysubscribe-resourcefor the matching URI; each step sets the content and fires aresource-updated.delay_ms— how long every scripted step waits before firing (default10).
There is no fallback: a subscription with no configured updates emits nothing.
Ordering with after
Each tool_calls / resource_*_updates step takes an optional after, the
shared gate the endpoint mock uses and the same patterns as
assertions:
- absent or
"start"— fire as soon as the step is reached. - a
call/inboundpattern — wait until a matching trace event has been observed at any point in the run, then fire.
The mock logs the trace when the run is built, so a gate can observe events that precede the step that waits on it; because the log is append-only, gates never consume each other’s events (so this is safe across subscriptions and across agents sharing one tooling). Without a trace channel a gate warns and fires immediately rather than hanging.
Endpoint mock
kind = "mock" is an in-process scripted client for the endpoint. It plays
requests into a subscribed agent and drains the reply, with no socket and no
listen address.
[endpoint]
kind = "mock"
poll_ms = 10
requests = [
{ model = "gpt-4o", messages = [{ content = "hi" }], stream = true },
{
model = "gpt-4o",
messages = [{ content = "bye" }],
stream = false,
after = { kind = "call", op = "chat" },
},
]
Keys
poll_ms — how long to wait between checks for a model subscription (default
10); the request fires as soon as the agent subscribes, so this only caps how
often the mock re-checks.
Each requests entry is:
model— the model the request targets; the mock waits until an agent subscribes to it under this name.messages— the inbound chat messages, as{ role, content }(role defaults touser).tools— any tools the caller offered.params— opaque generation params (temperature, …) to forward with the request, surfaced as theendpoint-message’sparamsJSON string.stream— whether the caller asked for SSE. Informational only: the mock drains the same session either way.after— the ordering gate (below).
Ordering with after
after is the shared step gate: the same "start" / pattern vocabulary the
tooling mock uses and the same patterns as
assertions. Timing-based endpoint scripting is
flaky; in an actor model what matters is event order. after gates a request
on the recorded trace:
- absent or
"start"— fire as soon as the model is subscribed (the default). - a
call/inboundpattern — wait until an event matching the pattern has been observed at any point in the run, then fire. The mock logs the trace before any agent runs, so it can gate on an event that precedes its own subscription.
So after = { kind = "call", op = "chat" } routes the request only once the
agent has made its first chat, letting a test assert that the brain handles
the request at a particular point relative to its other work.
A $while / $until gate resolves against its inner call / inbound
condition; an embedder that runs without a trace channel fires immediately with
a warning.
Examples
The examples/ directory holds runnable agents that exercise the whole omw
stack with no keys, no network and no external services. Each case is a shared
omw.test.template.toml plus a brain per variant under rhai/, js/ and
wasm/, and a committed <variant>/omw.test.toml generated from the template.
They double as documentation: read a case’s omw.test.toml to see what a brain
does, and its brain.rhai / brain.js / brain.rs to see how to write it.
Running
Run one case (every variant) or all of them:
omw-test run examples/01-hello
omw-test run examples
Each case is checked against the assertions in its config, so a green run means the brain’s provider calls, inbox events and terminal outcome match what the case claims. See Testing for the binary and the assertion language.
The wasm cells need their brain.wasm, which the repo’s dev shell builds for
you through dev test; to build one cell by hand, see
Rust brains. The dev wrapper runs every cell
(dev test brain examples) and gates the wasm ones on the
OMW_TEST_WASM_RUNTIME_NON_NATIVE environment variable.
Variants and templates
A case is one config shared across three languages. The template uses two
placeholders — {{RUNTIME}} (the runtime kind) and {{SCRIPT}} (the sibling
brain file) — and names its runtime runtime:
[runtime.runtime]
kind = "{{RUNTIME}}"
[[agents]]
name = "alice"
runtime = "runtime"
script = "{{SCRIPT}}"
dev format expands each template into its per-variant omw.test.toml (only
for variants whose brain file exists) and commits the result; dev lint
regenerates and compares, failing when a committed config is stale. You can
regenerate by hand with omw generate test config <case> <variant>.
Assertions
Each case carries an [assertions.<agent>] block describing the calls and inbox
events that agent should see, in order, plus its outcome. Assertions are an
ordered subsequence over partial patterns, so a case pins only the interesting
slice of the trace. See Testing for the full
language.
The cases
01-hello— provider wiring: one blockingchat.02-tool-agent— the ReAct tool round-trip:chat→call_tool_blocking→chat.03-endpoint— an OpenAI-compatible endpoint session:subscribe_endpoint, an inboundendpoint-message, a reply streamed back withstream_endpoint.04-ping-pong— two agents on one shared brain,subscribe_agent/send_agentplus per-agent memory.05-patterns— assertion patterns: regexdetailleaves,$while,$until, and a subsequence over a multi-turn provider script.06-memory—[memory.alice]seeds a value the brain reads and branches on.07-asserted—outcome = "asserted"stops a brain that would otherwise loop forever.08-endpoint-order— the endpoint mock’saftergate asserts a request arrives at a specific point relative to the brain’s calls.09-resources— the tooling mock’s resource list and content: subscribe to both, read a resource, and react to scripted,after-gated updates.
Deployment
omw runs anywhere its single static binary runs. The release tarballs
(omw-<arch>.tar.gz, plus -rhai / -js variants) contain a statically linked
musl binary — no runtime dependencies besides CA certificates for TLS — so the
same artifact drops onto NixOS (via the NixOS module), onto
any systemd host (via the plain unit), or into a minimal
container (see Docker).
Configuration everywhere
Every deployment reads the same TOML file (--config, default omw.toml).
assets/omw.example.toml in the repo is the shared starting point; the
Docker page walks through it. Secrets layer over the file
from OMW__-prefixed environment variables (__ separator), e.g.
OMW__PROVIDERS__OPENAI__API_KEY overrides providers.openai.api_key — keep
keys out of the file and supply them from the environment.
Workspace convention
Agents that touch the filesystem (filesystem MCP tooling, brain scripts on disk) expect a persistent workspace directory:
| deployment | workspace | how it is provided |
|---|---|---|
| NixOS | /var/lib/<stateDir> | services.omw.stateDir (StateDirectory=) |
| systemd | /var/lib/omw | StateDirectory=omw in the unit |
| Docker | /var/lib/omw | named volume or bind mount |
Point brain script paths and filesystem tooling roots at the workspace so all
three deployments share one config shape.
Static builds
The flake cross-compiles x86_64-unknown-linux-musl /
aarch64-unknown-linux-musl and checks the result with file + ldd
(checks.static* in src/nix/dev.nix). That is what makes the Alpine image
(assets/Dockerfile) a plain COPY with no toolchain.
Systemd (non-NixOS)
For hosts without the NixOS module, the unit below is a
drop-in that mirrors what the module generates with
services.omw.hardening = true (in the repo: assets/omw.service):
[Unit]
Description=OMW agent runtime
After=network-online.target
Wants=network-online.target
[Service]
Type=simple
User=omw
Group=omw
WorkingDirectory=/var/lib/omw
EnvironmentFile=/etc/omw/env
ExecStart=/usr/local/bin/omw loop --config /etc/omw/omw.toml
Restart=on-failure
RestartSec=5s
StateDirectory=omw
# Systemd sandboxing. Mirrors `services.omw.hardening = true` from the NixOS
# module. `MemoryDenyWriteExecute` is deliberately absent: it would break the
# wasmtime JIT and nodejs MCP servers. `PrivateDevices=false` keeps /dev
# usable for MCP servers and bwrap wrappers. `LimitMEMLOCK=infinity` is
# required: secrets are mlock()ed and the empty `CapabilityBoundingSet`
# drops CAP_IPC_LOCK, so a zero memlock limit would fail every secret with
# EPERM. Add `BindReadOnlyPaths=` entries if brains or the config live outside
# /var/lib/omw, and `ReadWritePaths=` entries for filesystem MCP workspaces
# outside the state directory.
NoNewPrivileges=true
PrivateDevices=false
PrivateIPC=true
ProtectClock=true
ProtectControlGroups=true
ProtectHome=true
ProtectKernelModules=true
ProtectProc=invisible
ProtectSystem=strict
RemoveIPC=true
RestrictAddressFamilies=AF_UNIX AF_NETLINK AF_INET AF_INET6
RestrictRealtime=true
RestrictSUIDSGID=true
LockPersonality=true
SystemCallArchitectures=native
UMask=0077
LimitMEMLOCK=infinity
CapabilityBoundingSet=
AmbientCapabilities=
[Install]
WantedBy=multi-user.target
Install it (the assets/… paths are repo-relative):
sudo useradd -r -d /var/lib/omw omw
sudo install -m 644 assets/omw.service /etc/systemd/system/omw.service
sudo mkdir -p /etc/omw
sudo install -m 600 assets/omw.example.toml /etc/omw/omw.toml
sudo install -m 600 assets/omw.example.env /etc/omw/env
sudo systemctl daemon-reload
sudo systemctl enable --now omw
Secrets live in /etc/omw/env as OMW__-prefixed variables (see the
deployment overview); they layer over /etc/omw/omw.toml at
runtime. Put brains and the filesystem tooling workspace under /var/lib/omw
(the unit’s StateDirectory + WorkingDirectory).
Hardening
The unit carries the same sandbox as the NixOS module (see the NixOS module for the rationale):
NoNewPrivileges,RestrictSUIDSGID,RestrictRealtime,LockPersonality, empty capability sets plusLimitMEMLOCK=infinity(for allowing secret protection against swapping) — safe, and compatible withbwrap-wrapped MCP servers.ProtectSystem=strict+ProtectHome+ProtectProc=invisible+PrivateIPC— the filesystem is read-only outside API mounts and the state directory; allow-list extras withBindReadOnlyPaths=(brains or config outside/var/lib/omw) andReadWritePaths=(filesystem MCP workspaces elsewhere).RestrictAddressFamilies=AF_UNIX AF_NETLINK AF_INET AF_INET6— provider egress plus local stdio MCP sockets.MemoryDenyWriteExecuteis deliberately absent: it would break the wasmtime JIT and nodejs MCP servers.PrivateDevices=falsekeeps/devusable for MCP servers andbwrapwrappers.
Docker
All paths below are repo-relative (assets/Dockerfile, assets/compose.yaml,
assets/omw.example.toml, assets/omw.example.env). The Dockerfile builds a
minimal Alpine image straight from the GitHub release tarballs (the binaries are
static musl, so there is nothing to compile):
# omw on Alpine. The release binaries are statically linked (musl), so they
# drop straight onto Alpine with no extra runtime dependencies besides
# ca-certificates and tzdata.
#
# docker build \
# --build-arg OMW_VERSION=0.1.0 \
# --build-arg OMW_VARIANT=rhai \
# --build-arg OMW_ARCH=x86_64-linux \
# -t omw .
#
# Variants: `default` (wasm brains only), `rhai`, `js`. The tarball names are
# `omw-<arch>.tar.gz`, `omw-rhai-<arch>.tar.gz`, `omw-js-<arch>.tar.gz`.
ARG OMW_VERSION=0.1.0
ARG OMW_VARIANT=rhai
ARG OMW_ARCH=x86_64-linux
FROM alpine:3.21 AS fetch
ARG OMW_VERSION
ARG OMW_VARIANT
ARG OMW_ARCH
RUN apk add --no-cache ca-certificates \
&& if [ "${OMW_VARIANT}" = "default" ]; then TARBALL="omw-${OMW_ARCH}.tar.gz"; BIN="omw-${OMW_ARCH}"; \
else TARBALL="omw-${OMW_VARIANT}-${OMW_ARCH}.tar.gz"; BIN="omw-${OMW_VARIANT}-${OMW_ARCH}"; fi \
&& wget -O "/tmp/${TARBALL}" \
"https://github.com/haras-unicorn/omw/releases/download/v${OMW_VERSION}/${TARBALL}" \
&& tar -xzf "/tmp/${TARBALL}" -C /tmp \
&& mv "/tmp/${BIN}" /tmp/omw \
&& chmod +x /tmp/omw
FROM alpine:3.21
RUN apk add --no-cache ca-certificates tzdata \
&& adduser -D -H omw \
&& mkdir -p /var/lib/omw /etc/omw \
&& chown omw:omw /var/lib/omw
COPY --from=fetch /tmp/omw /usr/local/bin/omw
# Config, brains and the agent workspace. Mount your config at
# /etc/omw/omw.toml and the workspace at /var/lib/omw (the stateDir
# equivalent); see assets/omw.example.toml and assets/compose.yaml.
VOLUME /var/lib/omw
EXPOSE 8080
USER omw
ENTRYPOINT ["/usr/local/bin/omw", "loop", "--config", "/etc/omw/omw.toml"]
docker build \
--build-arg OMW_VERSION=0.1.0 \
--build-arg OMW_VARIANT=rhai \
-t omw .
docker run -d --name omw --restart unless-stopped \
--env-file omw.env \
-e OMW__TUNABLES__ALLOW_UNLOCKED_SECRETS=true \
-v ./omw.toml:/etc/omw/omw.toml:ro \
-v omw-workspace:/var/lib/omw \
-p 8080:8080 \
omw
Secrets cannot mlock() inside containers (the container’s own RLIMIT_MEMLOCK
is enforced regardless of the unit’s LimitMEMLOCK=), so the shipped example
docker-compose.yaml sets OMW__TUNABLES__ALLOW_UNLOCKED_SECRETS=true — see
tunables. Start from the example config and env.
Secrets travel as OMW__-prefixed variables (--env-file or -e), never baked
into the image. /var/lib/omw is the stateDir equivalent: mount a named
volume or bind mount there for brains and the filesystem MCP workspace. Publish
8080 only when [endpoint] is configured (listen = "0.0.0.0:8080" inside
containers).
Compose
The compose example wires it together, including an optional HTTP MCP server on the same network. The example config it mounts is:
# Example compose setup for omw. Builds the image straight from the
# release tarballs (see assets/Dockerfile); any git URL works as a build
# context, e.g. `context: https://github.com/haras-unicorn/omw.git`.
services:
omw:
build:
context: https://github.com/haras-unicorn/omw.git
dockerfile: assets/Dockerfile
args:
OMW_VERSION: "0.1.0"
# default | rhai | js
OMW_VARIANT: rhai
# x86_64-linux | aarch64-linux
OMW_ARCH: x86_64-linux
env_file:
- ./omw.env
environment:
# Secrets cannot mlock() inside containers (the container's own
# RLIMIT_MEMLOCK wins over the unit's LimitMEMLOCK=), so ensure
# that this doesn't make entire OMW fail.
- OMW__TUNABLES__ALLOW_UNLOCKED_SECRETS=true
# Layers over providers.openai.api_key in omw.toml; prefer env_file.
- OMW__PROVIDERS__OPENAI__API_KEY=${OPENAI_API_KEY}
volumes:
# Config + brains (read-only) and the agent workspace (read-write,
# the stateDir equivalent for filesystem MCP tooling).
- ./omw.toml:/etc/omw/omw.toml:ro
- omw-workspace:/var/lib/omw
# Only needed with [endpoint] (listen = "0.0.0.0:8080").
ports:
- "8080:8080"
restart: unless-stopped
# Example HTTP MCP server on the same network. Point a tooling at it with
# transport = "http" and url = "http://mcp:8000/mcp".
# mcp:
# image: ghcr.io/modelcontextprotocol/server-everything
# networks:
# - default
volumes:
omw-workspace:
Compose accepts a git URL as build.context, so you can build without cloning
the repo; point dockerfile at the in-repo path and keep your omw.toml /
omw.env beside your own compose file:
services:
omw:
build:
context: https://github.com/haras-unicorn/omw.git
dockerfile: assets/Dockerfile
MCP servers
- HTTP transport (easy): run the server as another compose service and point
the tooling at it (
transport = "http",url = "http://mcp:8000/…"). No image changes needed. - stdio transport (node, uvx, bwrap wrappers, …): the server command must exist inside the omw image — a sidecar cannot help, since stdio means a subprocess. Extend the image:
FROM omw AS with-mcp
RUN apk add --no-cache nodejs
RUN npm install -g @modelcontextprotocol/server-filesystem
Then reference command = "server-filesystem" (or npx …) in the tooling
config. The same applies to bwrap-wrapped commands: install bubblewrap in
the image; NoNewPrivileges-style restrictions do not apply inside containers,
and user namespaces work under the default Docker seccomp profile.
The NixOS module
The flake ships a NixOS module exposing a single services.omw option set that
runs omw as a systemd service. It is the recommended way to run an omw agent (or
several — a single service can run omw run / omw loop, which already
supports multiple agents from one config) on NixOS.
The full option reference is generated from the module by the flake’s
omw-options package; this page explains the design and how to use it.
Enabling the module
{
inputs = {
nixpkgs.url = "github:nixos/nixpkgs/nixos-26.05";
omw.url = "github:haras-unicorn/omw";
};
outputs =
{ nixpkgs, omw, ... }:
{
nixosConfigurations.my-machine = nixpkgs.lib.nixosSystem {
modules = [
omw.nixosModules.default
{
services.omw = {
enable = true;
mode = "loop";
settingsFile = "/etc/omw.toml";
environmentFile = "/var/lib/omw/env";
};
}
];
};
};
}
How the service runs
The unit runs omw <mode> --config <file> directly (ExecStart, no shell):
omw loop --config /etc/omw.toml
Two things follow from this:
settingsandsettingsFileare mutually exclusive.settingsis an attribute set rendered to TOML at build time;settingsFileis a path to a TOML file on the system. Choose whichever fits.- Secrets are layered from the environment.
OMW__-prefixed variables (__separator, e.g.OMW__PROVIDERS__OPENAI__API_KEY) override file values at runtime, so API keys never have to live in the Nix store. Set them with theenvironmentoption (systemdEnvironment=) or anenvironmentFile.
mode selects run (every agent once) or loop (keep agents running,
restarting on failure — the default, suited to a service).
extraArgs passes extra CLI flags after the mode; use [ "--watch" ] to
hot-reload agent scripts (a changed brain file restarts its agent while inboxes
and subscriptions survive).
variant selects which package variant runs: default (the
crates.io-equivalent build, no script runtime), rhai (the omw-rhai package,
which compiles the bundled rhai interpreter in) or js (the omw-js package,
which compiles the bundled js interpreter in). Overridable entirely with
package.
Users and state
By default the service runs under a systemd dynamic user (no user /
group). Set user and/or group to pin a specific identity. stateDir
declares a StateDirectory (created under /var/lib, also used as
WorkingDirectory), which is where a filesystem MCP tooling’s workspace would
live and where the service can persist state.
Example with an MCP filesystem tooling rooted at the state directory:
{
services.omw = {
enable = true;
variant = "rhai";
user = "omw";
group = "omw";
stateDir = "omw";
settings.providers.openai = {
kind = "openai";
model = "gpt-4o";
};
settings.tooling.fs = {
kind = "mcp";
transport = "stdio";
command = "npx";
args = [
"-y"
"@modelcontextprotocol/server-filesystem"
"/var/lib/omw"
];
};
settings.runtime.rhai.kind = "rhai";
settings.agents = [
{
name = "alice";
runtime = "rhai";
script = "/var/lib/omw/brain.rhai";
}
];
environment.OMW__PROVIDERS__OPENAI__API_KEY = "…";
};
}
Hardening
services.omw.hardening (default true) applies a systemd sandbox modeled on a
service that spawns nodejs MCP servers and bwrap wrappers:
- Identity/capabilities:
NoNewPrivileges=true(confirmed compatible withbwrap-wrapped MCP servers),RestrictSUIDSGID=true,RestrictRealtime=true,LockPersonality=true,SystemCallArchitectures=native, emptyCapabilityBoundingSet/AmbientCapabilities.LimitMEMLOCK=infinityraises the unit’s ceiling, but inside containers the outerRLIMIT_MEMLOCKstill wins — setOMW__TUNABLES__ALLOW_UNLOCKED_SECRETS=truein the environment there instead (see tunables), since secrets cannotmlock()under the container’s limit. Deliberately omitted:MemoryDenyWriteExecute=truewould break the wasmtime JIT and nodejs MCP servers;SystemCallFilter=is left open for the same reason (wasmtime/node need broad syscalls). - Filesystem/IPC:
ProtectSystem=strict,ProtectHome=true,ProtectClock=true,ProtectKernelModules=true,ProtectControlGroups=true,ProtectProc=invisible,PrivateIPC=true,RemoveIPC=true,UMask=0077.PrivateDevices=falsekeeps/devusable for MCP servers andbwrap. - Network:
RestrictAddressFamilies=AF_UNIX AF_NETLINK AF_INET AF_INET6— provider egress plus local stdio sockets and NSS lookups.
Under ProtectSystem=strict the unit only sees API mounts plus its state
directory. Use services.omw.readOnlyPaths (BindReadOnlyPaths=) when brains
or the config file live elsewhere (e.g. the /etc/brain.rhai used in tests),
and services.omw.readWritePaths (ReadWritePaths=) for filesystem MCP
workspaces outside the state directory. Set services.omw.hardening = false to
drop the sandbox entirely, or override individual keys with
services.omw.serviceConfig, which merges last.
NixOS module options
services.omw.enable
Whether to enable the omw agent runtime.
Type: boolean
Default:
false
Example:
true
services.omw.package
The omw package to run.
Type: package
Default:
<derivation omw>
services.omw.environment
Environment variables for the service (becomes systemd Environment=).
OMW__-prefixed variables layer over the file configuration, e.g.
OMW__PROVIDERS__OPENAI__API_KEY overrides providers.openai.api_key, which is
the intended way to supply API keys and other secrets.
Type: attribute set of string
Default:
{ }
services.omw.environmentFile
Path to a systemd EnvironmentFile for the service.
Type: null or absolute path
Default:
null
services.omw.extraArgs
Extra arguments passed to the omw command line after the mode.
Type: list of string
Default:
[ ]
services.omw.group
The group the service runs as. When both services.omw.user and
services.omw.group are null, a dynamic user is allocated.
Type: null or string
Default:
null
services.omw.hardening
Enable systemd hardening (NoNewPrivileges, ProtectSystem=strict, …). Safe
for stdio MCP servers (including ones wrapping commands in bwrap) and
node-based servers. Sets LimitMEMLOCK=infinity alongside the empty capability
set (inside containers the outer RLIMIT_MEMLOCK still wins — set
OMW__TUNABLES__ALLOW_UNLOCKED_SECRETS=true in the environment there instead).
Deliberately omits MemoryDenyWriteExecute, which would break the wasmtime JIT
and nodejs MCP servers. Set to false if the sandbox gets in the way;
services.omw.serviceConfig can override individual keys either way.
Type: boolean
Default:
true
services.omw.mode
Which mode to run omw in: run executes every agent once, loop keeps running
them, restarting agents that fail.
Type: one of “run”, “loop”
Default:
"loop"
services.omw.readOnlyPaths
Extra paths exposed read-only inside the sandbox (BindReadOnlyPaths=). Needed
with hardening when brains or the config file live outside the state directory
(e.g. /etc).
Type: list of string
Default:
[ ]
services.omw.readWritePaths
Extra paths exposed read-write inside the sandbox (ReadWritePaths=). Needed
with hardening when a filesystem MCP tooling works outside the state directory.
Type: list of string
Default:
[ ]
services.omw.serviceConfig
Extra systemd serviceConfig merged last, so it wins over the module defaults
(including the hardening set). Escape hatch for anything the module does not
model explicitly.
Type: attribute set
Default:
{ }
services.omw.settings
The omw configuration provided as an attribute set, rendered to TOML at build
time. Mutually exclusive with services.omw.settingsFile. Secrets are layered
at runtime from OMW__-prefixed environment variables (see
services.omw.environment), so API keys never have to live in the Nix store.
Type: null or TOML value
Default:
null
services.omw.settingsFile
Path to an omw configuration file (TOML). Mutually exclusive with
services.omw.settings.
Type: null or absolute path
Default:
null
services.omw.stateDir
Name of the state directory created for the service (StateDirectory=). When
set, the directory is created under /var/lib and the service can persist state
there, e.g. the workspace of a filesystem MCP tooling.
Type: null or string
Default:
null
services.omw.user
The user the service runs as. When both services.omw.user and
services.omw.group are null, a dynamic user is allocated.
Type: null or string
Default:
null
services.omw.variant
Which package variant to run: default (the crates.io-equivalent build,
without the rhai runtime), rhai (adds the bundled rhai interpreter via the
omw-rhai package) or js (adds the bundled js interpreter via the omw-js
package). Overridable with package.
Type: one of “default”, “rhai”, “js”
Default:
"default"