Agent OS · v0.1 · Phase 4 built · Offline acceptance

Let an AI agent change your code without trusting it.

Agent OS runs a coding agent against a repository under a contract you sign off on first. Every side effect goes through a journal, every permission is a revocable handle, and, with the Firecracker worker, the code it touches runs inside a locked-down microVM. A task only counts as done when an independent check passes on exactly the final files.

It survives crashesKill the controller at any point. A restart recovers the same task and never repeats a step that already finished.
It checks permissions every timeThe agent holds no authority itself. A broker checks a handle before each action and records the answer.
It contains the code it runsWith --worker firecracker, each step runs in its own microVM with no network, fixed-size disks and a jailed host process. The default host worker is not sandboxed.

This static walkthrough simulates the controls in your browser; it does not run code, contact a model or start a VM. The controller supports a deterministic patch agent and an Anthropic model agent with five tools. Model reliability and recovery are tested offline; live-provider and real KVM acceptance remain open for this build. The event sequence below comes from a fixture run. Sandbox figures are assertions in the gated KVM tier, not measurements from this review.

Use it for a real bug fix · Try the offline model workflow · Testing guide · Current progress and open gates

Before anything runs

You approve a contract, not a prompt

The task is a small JSON document. submit prints what it grants and waits for approval. Nothing runs until you approve it with --yes (or later with resume), and the capabilities it lists become the only authority the task ever has.

{
  "goal": "fix the parser",
  "repository": {
    "source": "fixtures/parser-repo",
    "revision": "recorded-at-submission"
  },
  "profile": "python-stdlib-v1",
  "editable_paths": ["src/**"],
  "verification_profile": "parser-checks-v1",
  "capabilities": [
    "snapshot.read",
    "workspace.apply_patch",
    "verification.run",
    "artifact.export"
  ],
  "limits": {
    "model_requests": 1,
    "max_output_tokens_per_request": 1000,
    "tool_actions": 10,
    "deadline_seconds": 600,
    "worker_vcpus": 1,
    "worker_memory_mib": 256
  }
}
$ agentos submit task.json --yes …
task 01a10214-… submitted; approve these
permissions before it runs:
  goal:         fix the parser
  repository:   fixtures/parser-repo
                at be77aa19c032…
  capabilities: snapshot.read,
                workspace.apply_patch,
                verification.run,
                artifact.export
  editable:     src/**
  acceptance:   protected verification
                profile parser-checks-v1
                (9ff584f31b7f…)
  limits:       model_requests=1
                max_output_tokens_per_request=1000
                tool_actions=10
                deadline_seconds=600
                worker_vcpus=1
                worker_memory_mib=256
  agent:        fake-agent (patch 5128fe0b…)
  worker:       firecracker, guest image
                python-stdlib-v1@f231e3ea…,
                jailed

The repository is pinned by digest, and the check that decides success is a registered profile pinned by digest too. The agent cannot edit the check, and a patch that touches a path outside src/** is refused.

Phase 4 · offline model workflow

A model can repair the fixture through five tools

The model agent can list files, read a file, apply a patch, run protected verification and finish. A scripted provider exercises the same loop without a key or network call: it reads the parser, tries a patch, sees the check fail, then fixes it. Only the passing check on the final workspace makes the task succeed.

Save this contract as model-task.json in the repository root. Run docker compose build test, then open the test container with docker compose run --rm test sh and paste these commands. This example uses the host worker on the trusted fixture. The exported bundle stays in a new directory under build/.

{
  "goal": "fix the parser",
  "repository": {
    "source": "/work/fixtures/parser-repo",
    "revision": "recorded-at-submission"
  },
  "profile": "python-stdlib-v1",
  "editable_paths": ["src/**"],
  "verification_profile": "parser-checks-v1",
  "capabilities": [
    "snapshot.read",
    "workspace.apply_patch",
    "verification.run",
    "artifact.export",
    "model.request"
  ],
  "limits": {
    "model_requests": 8,
    "max_output_tokens_per_request": 16000,
    "tool_actions": 8,
    "deadline_seconds": 600,
    "worker_vcpus": 1,
    "worker_memory_mib": 256
  }
}
cargo build --locked -q -p agentos-cli
mkdir -p build
demo_home=${DEMO_HOME:-$(mktemp -d /work/build/model-demo.XXXXXX)}
contract=${DEMO_CONTRACT:-model-task.json}
agentos() {
  target/debug/agentos --home "$demo_home" --worker host "$@"
}
agentos profile register fixtures/profiles/parser-checks-v1
agentos submit "$contract" --yes \
  --model fake:fixtures/transcripts/parser-fix.json
task_id=$(basename "$demo_home"/tasks/*)
agentos status "$task_id"
agentos export "$task_id" "$demo_home/bundle"
printf 'Export: %s/bundle\n' "$demo_home"

The fixture makes six model calls and settles five tool actions. Its final verification pins the workspace digest and the protected profile; the export includes the patch, requests, responses and verification evidence. The original crash simulation below still illustrates the deterministic patch-agent run.

Model failures have durable rules

Permanent errors stopNew policy-1 tasks stop on 400, 401, 403, 404, redirects or malformed complete responses. Only 408, 429 and 5xx responses retry.
Retries wait and survive a restartTransient failures and lost calls get journaled waits of 2, 4, 8, 16, 32, then 60 seconds. Retry-After can extend them. Every new attempt gets its own effect and reservation.
Interrupted calls stay countedA sent call whose response is lost remains uncertain and counts against the budget. Pause, cancel, revocation and deadline expiry stop waiting without another send.
Model I/O is boundedNew tasks cap requests at 8 MiB. Success bodies are capped at 4 MiB, error prefixes and regular key files at 4096 bytes. The task deadline bounds network waiting.

New submissions record the endpoint and policy versions. Resume refuses a different provider endpoint before reading credentials or changing the task. Older journals retain their recorded behavior. Version-2 recordings preserve ordered attempts, including transient rejections and retry metadata; legacy response-only fixtures still replay.

Test repair and replay on both worker paths

docker compose build test
sh scripts/acceptance.sh offline

This runs the host and fake-jail harnesses against a loopback fake API, including a 429 followed by repair, then replays without network. It compares request digests, call counts, final workspace, exported patch and protected profile. It needs no real API key or /dev/kvm. Fake-jail tests exercise the harness; they do not prove VM isolation.

Real host and jailed provider runs, real KVM tests and candidate guest-image acceptance remain pending. See the acceptance setup and evidence requirements before enabling those tiers.

From a bug report to a patch you can review

Your Python service reads settings such as PORT = 8080. A customer reports that spaces around the equals sign break the configuration. Give the agent this narrow bug, allow edits only under src/, and let a protected regression check decide whether the fix works. The result is a reviewable export; applying it to your project is a separate step you control.

1. Reproduce the bug

Start with a clean source snapshot and a check that fails for the reported input. Register that check separately from the editable code. The practice run below uses the included parser as the service's configuration reader.

2. Review, then approve

Submit without --yes. Inspect the goal, editable paths, protected profile, worker and budgets. The task stays READY; resume approves this recorded task and starts it.

3. Inspect the evidence

Check status and export the result. Require SUCCEEDED, matching final and verified workspace digests, and a passing verification accepted for the final workspace. Read the diff and the verification results before accepting the change.

4. Apply it to a clean checkout

Run git apply --check, apply the reviewed export, and rerun your checks. The practice uses a disposable review copy. The original source snapshot stays unchanged throughout the agent run.

Practice the whole handoff without an API key

Save the contract above as model-task.json, then paste this in docker compose run --rm test sh. The scripted answers are specific to this parser example; an arbitrary project needs a live provider.

cargo build --locked -q -p agentos-cli
mkdir -p build
practice=${PRACTICE_DIR:-$(mktemp -d /work/build/usage.XXXXXX)}
mkdir -p "$practice"
cp -R fixtures/parser-repo "$practice/service"
cp -R "$practice/service" "$practice/review"
git init -q "$practice/review"
check=fixtures/profiles/parser-checks-v1/check_parser.py

# The protected check must expose the bug before we ask for a fix.
if pyenv exec python "$check" "$practice/service"; then
  echo 'Expected the original parser to fail' >&2; exit 1
else
  [ "$?" -eq 1 ]
fi

pyenv exec python - "${MODEL_CONTRACT:-model-task.json}" "$practice" <<'PY'
import json, sys
from pathlib import Path
root = Path(sys.argv[2]).resolve()
task = json.loads(Path(sys.argv[1]).read_text())
task["goal"] = ("Fix config parsing: trim keys and values, "
                "skip indented comments, preserve equals signs in values.")
task["repository"]["source"] = str(root / "service")
(root / "task.json").write_text(json.dumps(task))
PY
agentos() {
  target/debug/agentos --home "$practice/home" --worker host "$@"
}
agentos profile register fixtures/profiles/parser-checks-v1
agentos submit "$practice/task.json" \
  --model fake:fixtures/transcripts/parser-fix.json > "$practice/approval.json"
task_id=$(pyenv exec python -c \
  'import json,sys; print(json.load(open(sys.argv[1]))["task_id"])' \
  "$practice/approval.json")
agentos status "$task_id" > "$practice/ready-status.json"
cat "$practice/ready-status.json"

Stop here and review the printed permissions. The task is READY and has run no effects. When you approve it, continue in the same shell:

agentos resume "$task_id"
agentos events "$task_id" > "$practice/events.ndjson"
agentos export "$task_id" "$practice/bundle"
cat "$practice/bundle/patch.diff"

# Apply only to this disposable review copy after inspecting the diff.
git -C "$practice/review" apply --check "$practice/bundle/patch.diff"
git -C "$practice/review" apply "$practice/bundle/patch.diff"
pyenv exec python "$check" "$practice/review"
printf 'Task, evidence and review copy: %s\n' "$practice"

Use the same process on your own repository

Prepare a clean snapshot of the committed revision, replace the source path and goal in the contract, and register your project's regression checks as a protected profile. Keep credentials outside the snapshot: Agent OS copies regular files, including untracked files, and does not use .gitignore as a filter. The current Python profile supports standard-library checks; projects needing extra runtimes or packages need a matching execution image.

For a live task, mount the snapshot read-only and a regular key file read-only into the container, then submit with --model anthropic:MODEL_ID --api-key-file /run/agentos/key. Choose an available ID from Anthropic's model list. Submit first, approve with resume after reviewing permissions, and use the same export/review flow. Repository content included in model requests goes to the selected provider; live calls consume your API account's budget.

Use the host worker for a trusted first test; it is not sandboxed. For untrusted checks, prepare the jailed Firecracker worker on a Linux KVM host. Live-provider and real KVM acceptance remain pending. The hands-on testing guide includes setup, the own-repository commands, expected results, recovery tests and the later KVM/live gates.

Station 1 · durability

Kill the controller. Watch it recover.

Every effect (read the snapshot, apply the patch, run the check) is written to the journal as separate steps: intended, dispatched, executed, published, completed. Each effect runs as a job under its own supervisor process, so the job keeps going when the controller dies. A job ends by writing a receipt (its result) into its own directory, and a lease is the time the supervisor allows it before killing it. Pick a moment and kill the controller there.

ControllerNOT STARTED
One short-lived process per command. It drives the task and writes the journal.
Supervisor and jobNO JOB
Launched detached, so it outlives the controller. It enforces the lease and writes a receipt.
Patch applied to the workspace
0times

Choose a moment, then press Run and kill here. The default is the exact moment recorded in the real demo: the patch job is dispatched and the controller dies.

Journal (SQLite, append-only)

The sequence matches the real run of scripts/demo.sh. Crash points other than the recorded one are modelled from the recovery rules.

How the pieces fit

What the sandboxed worker keeps away from the host

host controllerrun loop journalSQLite, blobs brokerhandles, scopes supervisorone per effect jail: uid 61000, chroot, cgroup Firecracker guest agentgit, python authorize append launch job vsock what the guest never gets no network device (only lo) no host secrets, no env no capability handle no writable root filesystem no say over the verdict outcomes are built on the host
The controller asks the broker before it launches each job. The job's supervisor boots one microVM, talks to the guest only over vsock, and writes a receipt. The journal records every step on the host.
Station 2 · authority

Revoke a permission. Try to use it.

Capabilities are opaque handles issued when you approve the task. The broker checks one at intent, again at dispatch, and again at export. Every decision, allowed or denied, lands in the journal. Revoking a handle also stops any job that needs it.

Handles for this task
Try an action
Journal and broker decisions
Nothing yet. Try an action.

Same behaviour as agentos revoke <id> --capability artifact.export, which the README demo runs: the next export fails with export denied: capability artifact.export is not usable (revoked).

Station 3 · containment

Run hostile code. See what stops it.

The check that verifies a task is code from the registered verification profile, and it runs as untrusted. The gated KVM suite includes hostile profiles that try to burn CPU, eat memory, fork without limit, fill the disk, reach the network and find a host secret. These controls illustrate its expected assertions; this build has not rerun them on real KVM. Pick one.

One honest boundary: a bug in Firecracker or KVM could let code escape the VM. It would land in the jail (an unprivileged uid inside a chroot and cgroup, under Firecracker's own seccomp filter), and a second kernel-level bug would be needed to reach the host. With --allow-unjailed the escape lands as the controller's own user instead.

Where it stands

Built, tested, and what is next

DONE1 ContractsTask state machine, effects, CLI, fixture repo.
DONE2 DurabilitySQLite journal, receipts, recovery after a kill.
DONE3 AuthorityBroker, leases, supervisors, cancel and revoke.
BUILT3b-1 IsolationJailed Firecracker worker; real KVM revalidation pending.
BUILT4 Model workflowTools, repair loop, deadlines and durable retries; host/fake-jail tests pass. Live acceptance pending.
TESTED OFFLINEReliability and CIBounded model I/O, Python pins, formatting, lint and Compose test gates.
LATERResource lifecycleSafe collection, publication failure tests and measured VM resource limits.
LATER5 Wasm ABIComponent interfaces and a repository analyzer.
LATER6 ReleaseSource-built kernel, fresh-host installer and evidence across two snapshots.

All four required offline gates pass: formatting, strict Clippy, the host suite and the fake-jail suite. Offline and skipped gated tests do not establish real provider or KVM acceptance. The new Python guest is a candidate; the legacy image remains the default.

Completion plan · Verification and review decisions

Known limits in v0.1

Live acceptance pendingThe Anthropic adapter exists and passes offline tests; no live evidence is claimed.
Host worker is not sandboxedThe VM sandbox is opt-in with --worker firecracker.
Jailing needs rootIt also needs a delegated cgroup v2 tree and /dev/kvm. x86_64 only.
Nothing is garbage collectedJob directories and disk images grow with every attempt.
One driver per homeOnly one process drives a home at a time.
Export is not a journaled effectIt is authorized and audited, but outside the effect model.