Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Hot reload

omw run --watch / omw loop --watch restarts an agent when its brain script changes, without losing bus state. Only the run restarts; everything addressable survives.

This matches the actor model in actor: the inbox is the agent, the brain is just its current reader. Lifecycle notifications (reload, shutdown, reload-failure error) follow the same model: explicit subscribe, UUID correlation, no magic inbox traffic for agents that did not opt in.

Using --watch

Pass --watch to either mode:

omw run --config omw.toml --watch
omw loop --config omw.toml --watch

The watcher tracks each agent’s script path. It watches parent directories non-recursively (editors that save via write temp + rename still trigger) and debounces (see watch_debounce_ms in tunables), so one save restarts the agent once.

What survives and what dies

Preserved across reload:

  • Inboxes (queued events are never drained or dropped).
  • Agent subscriptions and endpoint subscriptions.
  • Providers, tooling, and endpoint sessions.
  • Per-agent memory (host.memory-get / memory-set / memory-remove): the same AgentContext is reused, so stored handles and state-machine state carry over. Treat entries like variables, not a database.

Discarded on reload:

  • Wasm memory/stack (a fresh Store per run).
  • Open pump tasks: chat streams, timers, resource subscriptions, tool calls. They are cancelled so stale events cannot leak into the next run (unless cancel_pumps_on_reload = false in tunables, which keeps them across reload).

Rhai brains re-read the script file and wasm brains reload the component on every iteration, so the next run picks up the edit.

Validation and startup

A bad edit never kills a good run. The supervisor validates the new script before aborting the live one, while the old run keeps executing. Only a valid script ends the current run. An invalid script keeps the old run alive, logs a host-side warning, and delivers an error event on the lifecycle handle (if the brain subscribed) so it can react however it wants.

Startup is gated the same way: a broken script never produces a first iteration. Without --watch this fails fast. With --watch the agent parks until an edit fixing the script validates, then starts normally.

How a restart happens

Three tiers, mirroring SIGTERM/SIGKILL:

  1. Cooperative: recv aborts on the next poll slice (see recv_slice_ms in tunables) without draining the inbox, blocking calls abort through a helper, and a reload/shutdown event covers try-recv pollers.
  2. Grace timeout (see reload_grace_secs in tunables): the supervisor stops waiting and reports the abort. Bounds restart latency.
  3. Preemptive (interrupt, interrupt_budget_ms after grace — see tunables): one increment traps unyielding loops that never yield to the host.

Shutdown reuses the same three tiers, but is terminal: run and loop both exit Ok (0) once every in-flight iteration aborted, logging the shutdown; only a genuine failure without a shutdown request errors. The signal subscription is process-wide (one SIGTERM/SIGINT latch awaited by every iteration, the endpoint server, and the loop backoff), so a signal arriving between iterations still aborts the next one. The endpoint drains gracefully: in-flight requests get an error session-end and the server stops accepting, and the watcher pump is aborted with the agent tasks.

Every host call parks somewhere different, so each needs its own interrupt:

Brain is stuck inWhere it parksInterruptUpstream sees
host.recvinbox slices (recv_slice_ms)flag poll, Err("agent reloaded"/"agent shutting down"), inbox kept; plus a reload/shutdown event for pollersnothing (pure local wait)
host.try-recv / send / etcreturns immediatelynothing needed; pollers observe the queued reload/shutdown eventnothing
provider.chat-stream pumpbiased select! on cancelopen streams are cancelled, pump breaksTCP close, OpenAI sees client disconnect
host.wait-* / timersselect! on cancelopen timers are cancellednothing
resource subscription pumpsstream + cancel selectopen subscriptions are cancelleddepends on transport
tooling.call-tool pumpselect! on the futureabort drops the in-flight callMCP call dropped client-side
blocking provider.chatabort slices (reload_poll_ms)abort handle, Err("agent reloaded")normal completed request, result dropped
blocking call-tool-blocking, list-*, read-resource, subscribe-*same helpersame as aboveserver runs to completion, result dropped
blocking sleep-for/until/cronsame helpersame as above = true cancelnothing
pure wasm while true {}executing wasm, never yields to hostinterrupt after grace + interrupt_budget_msnothing

All values are tunables with the defaults listed there.

Writing a reload-safe brain

Subscribe to lifecycle events explicitly and correlate by UUID, like every other subscription in the actor model (see host):

let lc = omw::host::subscribe_lifecycle();
let sub = omw::host::subscribe_agent("other");
loop {
  let e = omw::host::recv();
  if e.id == lc {
    if e.kind == "reload" { break; }
    if e.kind == "shutdown" { break; }
    if e.kind == "error" {
      omw::host::log("warn", "reload failed: " + e.payload);
      continue;
    }
  }
  // ... handle e ...
}
omw::host::unsubscribe_lifecycle(lc);

Rules:

  • Lifecycle is opt-in: subscribe_lifecycle returns a UUID; reload, shutdown, and reload-failure error events arrive tagged with it. A second subscribe errors (one lifecycle subscription per run). Unsubscribed brains still get aborted on a valid reload (recv errors), but get no events and no invalid-edit notice.
  • reload-failed is the plain error variant, not a new event: distinguish by e.id == lifecycle_uuid && e.kind == "error" and read the validation message from the payload. The host always logs regardless, so the event is the brain’s chance to notify itself, not the only signal.
  • Keep subscription handles in memory so the next run can reuse them instead of re-subscribing blindly:
let sub = omw::host::memory_get("other-sub");
if sub == () {
  sub = omw::host::subscribe_agent("other");
  omw::host::memory_set("other-sub", sub);
}
  • The live run never exits on an invalid edit. If you see error on the lifecycle handle, keep running.
  • Exit on reload (break, do cleanup); exit terminally on shutdown. recv may also abort with "agent reloaded" / "agent shutting down" when blocked between polls.
  • Prefer evented calls (chat-stream, call-tool, wait-*) over blocking ones (chat, call-tool-blocking, sleep-*): both cancel promptly, but evented handles keep delivering while blocking ones abort the handle.
  • Re-subscribe at the top of the script. Handles are one-shot UUIDs that die with the run (lifecycle included). Read them back from memory when they can outlive the run; otherwise rebuild any cached handles from scratch on each run.
  • Never while true {} without a host yield; an unyielding loop can only die by interrupt.
  • Pollers should use try-recv + small wait-for, so the reload / error events are observed promptly.
  • A broken script never starts, so the first thing a fresh brain can assume is that it compiled.

Troubleshooting

  • “Reload takes a while”: the brain is stuck in a blocking call or unyielding loop and the grace (reload_grace_secs) expired. Prefer evented calls, or yield to the host regularly.
  • “Agent restarts twice”: two saves in quick succession, or a stale reload event surviving into the next run. The drain at iteration start drops queued system events; check the watcher debounce vs your editor’s save burst.
  • “Old stream still delivers”: a pump from the previous run outlived the reload. Reload cancels all open pumps; if you held the UUID, re-check is-open after reload instead of assuming it is alive.
  • “CPU spins after save”: a while true {} without a host yield. Only the interrupt can kill it, after the grace. Add a recv, wait-for, or sleep-for to the loop.
  • “Edit did nothing”: the script was invalid. Check host logs for the warning and, if subscribed, the lifecycle error payload. The live run kept going; fix the script and save again.
  • “Agent won’t start”: the startup gate rejected the script. Same signals as above. With --watch the agent parks until a fixing edit validates; without it the process fails fast.