Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Testing

omw-test is the deterministic brain-testing binary. It runs each discovered omw.test.toml through the same traced path the omw binary uses (run_agents_traced), against in-process scripted doubles (kind = "mock"), records what every agent saw and did, and checks that recording against an [assertions] section. No keys, no network, no external services.

A test config is a normal omw.toml-shaped file named omw.test.toml: the same config omw consumes, plus an [assertions] table that stock omw ignores the same way it ignores unknown keys. Values can be layered from the environment with the OMW_TEST__ prefix, exactly like OMW__ for omw itself.

Running

omw-test run examples/01-hello   # one config or directory
omw-test run examples            # every discovered config
  • <path> is a file → run that config once.
  • <path> is a directory → recursively find every test config and run each, printing PASS / FAIL <root-relative path> and a tally. It exits non-zero if any failed.
  • --include <glob> / --exclude <glob> (repeatable, OR within each) match the test’s root-relative path including its file name (* does not cross /, ** does); include is applied first, then exclude. A missing script is always a failure, never a skip; narrow the set with the globs instead.
  • --watch re-runs on change instead of exiting: after each pass it waits for a debounced filesystem event and runs again (file mode watches the config’s parent directory; directory mode watches the root recursively). The library hot-reload watch is always off.

Discovery skips hidden directories and collects any file whose name is omw.test.toml or ends with .omw.test.toml, so several test configs can live side by side in one directory. The shared templates (omw.test.template.toml) never match, since they end in template.toml.

Scaffolding

omw scaffold (in the omw binary) turns a deployment config into a starter test config: it introspects the real back ends and writes an omw.test.toml whose provider, tooling and endpoint are the in-config mocks, pre-populated where possible.

omw scaffold omw.toml                 # writes ./omw.test.toml
omw scaffold omw.toml --output t.toml # explicit output
omw scaffold omw.toml --no-resources  # skip listing/reading tooling resources
omw scaffold omw.toml --force         # overwrite an existing output
  • the provider mock gets the models the endpoint reported (GET /models, empty if the request fails), with an empty turns script;
  • the tooling mock gets the server’s tools (with their input schemas), initial_resource_list, and initial_resource_contents (unless --no-resources), with an empty tool_calls script;
  • the endpoint mock gets an empty requests list;
  • runtime, agents, [memory] and [tunables] are copied through verbatim.

Everything is best-effort: a back end that cannot be built or enumerated yields an empty mock and a warning instead of failing the conversion. The original params are never carried over, so secrets do not end up in the output. Fill in the turns, tool_calls and requests to script the run.

Assertions

[assertions.<agent>] is compared against the agent’s recorded trace: an optional terminal outcome plus an ordered events list.

[assertions.alice]
outcome = "completed"
events = [
  { kind = "call", op = "chat", detail = { model = "^gpt-" } },
  { kind = "inbound", event = "chat-delta" },
]
  • outcome is "completed", { exited = "<msg>" }, or "asserted" (below).
  • each events entry is one of:
    • { kind = "call", op = "...", detail = { ... } } — an outbound host call. op is exact; detail is a partial pattern over the call’s JSON detail.
    • { kind = "inbound", event = "...", payload = { ... } } — an inbox event. event is the kebab-case kind (chat-delta, chat-end, tool-result, endpoint-message, message, timer, reload, shutdown, error, …); payload is a partial pattern over the serialized event.
    • { "$while" = { kind = "call", ... } } — greedily consume a run of matching trace events, stopping at the first non-match.
    • { "$until" = { kind = "call", ... } } — skip trace events until one matches, consuming it.

The list is matched as an ordered subsequence over partial patterns:

  • Only the events you write are checked, in order; anything between them is ignored, and trailing events are fine. So events = [{ kind = "call", op = "chat" }] passes as soon as the first chat is seen, even if the brain does a hundred things afterwards.
  • A pattern object matches when every key it names is present and matches in the candidate; extra candidate keys are ignored. String leaves are regular expressions matched against the candidate string, so "^gpt-a.*" is a regex. Numbers, booleans and null are compared for equality.
  • Arrays match as ordered subsequences too, with the same rules as events. Unlisted elements between matches are skipped and leading/trailing elements are ignored, so [ "a", "b" ] matches [ "x", "a", "b", "y" ]. An empty pattern array matches any array.
  • Use detail/payload to pin only the fields you care about, and the $while/$until sentinels to step over look-alikes or force the cursor forward.

$while / $until sentinels

The events list and every array inside a pattern share one vocabulary. { "$while" = P } greedily consumes a run of consecutive elements matching the inner P, stopping at the first non-match (zero-or-more). { "$until" = P } skips ahead to the first element matching P and consumes it. Under events, P is a call/inbound assertion; inside a pattern array it is an ordinary pattern. $-prefixed keys are reserved and never mean a partial-match field.

[assertions.alice]
events = [
  { kind = "call", op = "chat" },
  { "$while" = { kind = "inbound", event = "chat-delta" } },
  { "$until" = { kind = "call", op = "chat" } },
]

An empty object under either sentinel is an ordinary pattern that matches any object, so { "$until" = {} } consumes the next object and { "$while" = {} } consumes a run of objects (objects only: {} does not match a primitive). An invalid regex fails at parse time with the offending pattern.

Chat detail

chat and chat_stream calls trace { provider, model, messages, tools }, serializing the conversation the brain sent. Together with array subsequence matching this expresses “many messages, only the tail matters”:

detail = { messages = [
  { role = "system" },
  { "$while" = { role = "assistant" } },
  { role = "user", content = "final" },
] }

detail stays partial, so naming only model (or nothing at all) still matches.

Seeded memory

Top-level [memory.<agent>] seeds the named agent’s memory before its brain first runs, so a test can fast-forward an agent to an interesting state instead of walking it there:

[memory.alice]
handle = "seed-42"

The brain reads it with memory_get like any other value, and seeded entries persist exactly like memory written at runtime (including across hot reloads). This is a first-class omw feature, not a test-only one.

outcome = "asserted"

A test usually cares about a prefix of a run, not its terminal outcome. If the brain loops or waits forever, making it exit on its own is awkward and easy to hang. outcome = "asserted" means: stop this agent as soon as its events settle, and fail it on the first mismatch. The assertion verdict is the result.

Stopping is per-agent, so one agent can be checked in isolation while others run normally. When every asserted agent has a verdict the harness forces the whole run down, so a brain that loops or blocks can never hang a test.

Because a $while is zero-or-more, a trailing $while settles as soon as its prefix does (immediately for an assertion that is only a $while, exactly like an empty events list), so it asserts nothing on its own. Bound a run you care about with a following anchor — $until or a plain assertion — which is also what makes the run meaningful under outcome = "asserted":

[assertions.alice]
outcome = "asserted"
events = [
  { "$while" = { kind = "inbound", event = "chat-delta" } },
  { kind = "inbound", event = "chat-end" },
]

Mock back ends

The mock cargo feature backs the deterministic doubles. omw-test is built with just the mocks, so a test config wires kind = "mock" for its provider, tooling and endpoint:

Tracing

omw-test asserts on what agents actually saw and did, through the library’s trace channel (host/trace.rs, exported via the prelude):

#![allow(unused)]
fn main() {
pub enum TraceEvent {
  Inbound { agent: String, id: String, event: Event },
  Call { agent: String, op: String, detail: serde_json::Value },
  Outcome { agent: String, outcome: RunOutcome },
}
pub type TraceSender = tokio::sync::broadcast::Sender<TraceEvent>;
}

run_agents_traced(cfg, watch, registries, tx) (and the loop_ twin) spawns a receiver-drain task, emits one Outcome per agent, and returns the flattened Vec<TraceEvent>; the omw::testing harness consumes the stream live and groups it per agent with host::trace::group. The channel is None for omw-cli and embedders, so it is zero-overhead when unset.

Embedders can drive the same machinery in-process through omw::testing (Harness, Assertions, parse, check, watch) instead of shelling out to the binary.