Agent OS runs a coding agent against a repository under a contract you sign off on first. Every side effect goes through a journal, every permission is a revocable handle, and, with the Firecracker worker, the code it touches runs inside a locked-down microVM. A task only counts as done when an independent check passes on exactly the final files.
--worker firecracker, each step runs in its own microVM with no network, fixed-size disks and a jailed host process. The default host worker is not sandboxed.This static walkthrough simulates the controls in your browser; it does not run code, contact a model or start a VM. The controller supports a deterministic patch agent and an Anthropic model agent with five tools. Model reliability and recovery are tested offline; live-provider and real KVM acceptance remain open for this build. The event sequence below comes from a fixture run. Sandbox figures are assertions in the gated KVM tier, not measurements from this review.
Use it for a real bug fix · Try the offline model workflow · Testing guide · Current progress and open gates
The task is a small JSON document. submit prints what it grants and waits for approval. Nothing runs until you approve it with --yes (or later with resume), and the capabilities it lists become the only authority the task ever has.
{ "goal": "fix the parser", "repository": { "source": "fixtures/parser-repo", "revision": "recorded-at-submission" }, "profile": "python-stdlib-v1", "editable_paths": ["src/**"], "verification_profile": "parser-checks-v1", "capabilities": [ "snapshot.read", "workspace.apply_patch", "verification.run", "artifact.export" ], "limits": { "model_requests": 1, "max_output_tokens_per_request": 1000, "tool_actions": 10, "deadline_seconds": 600, "worker_vcpus": 1, "worker_memory_mib": 256 } }
$ agentos submit task.json --yes …
task 01a10214-… submitted; approve these
permissions before it runs:
goal: fix the parser
repository: fixtures/parser-repo
at be77aa19c032…
capabilities: snapshot.read,
workspace.apply_patch,
verification.run,
artifact.export
editable: src/**
acceptance: protected verification
profile parser-checks-v1
(9ff584f31b7f…)
limits: model_requests=1
max_output_tokens_per_request=1000
tool_actions=10
deadline_seconds=600
worker_vcpus=1
worker_memory_mib=256
agent: fake-agent (patch 5128fe0b…)
worker: firecracker, guest image
python-stdlib-v1@f231e3ea…,
jailed
The repository is pinned by digest, and the check that decides success is a registered profile pinned by digest too. The agent cannot edit the check, and a patch that touches a path outside src/** is refused.
The model agent can list files, read a file, apply a patch, run protected verification and finish. A scripted provider exercises the same loop without a key or network call: it reads the parser, tries a patch, sees the check fail, then fixes it. Only the passing check on the final workspace makes the task succeed.
Save this contract as model-task.json in the repository root. Run docker compose build test, then open the test container with docker compose run --rm test sh and paste these commands. This example uses the host worker on the trusted fixture. The exported bundle stays in a new directory under build/.
{
"goal": "fix the parser",
"repository": {
"source": "/work/fixtures/parser-repo",
"revision": "recorded-at-submission"
},
"profile": "python-stdlib-v1",
"editable_paths": ["src/**"],
"verification_profile": "parser-checks-v1",
"capabilities": [
"snapshot.read",
"workspace.apply_patch",
"verification.run",
"artifact.export",
"model.request"
],
"limits": {
"model_requests": 8,
"max_output_tokens_per_request": 16000,
"tool_actions": 8,
"deadline_seconds": 600,
"worker_vcpus": 1,
"worker_memory_mib": 256
}
}
cargo build --locked -q -p agentos-cli
mkdir -p build
demo_home=${DEMO_HOME:-$(mktemp -d /work/build/model-demo.XXXXXX)}
contract=${DEMO_CONTRACT:-model-task.json}
agentos() {
target/debug/agentos --home "$demo_home" --worker host "$@"
}
agentos profile register fixtures/profiles/parser-checks-v1
agentos submit "$contract" --yes \
--model fake:fixtures/transcripts/parser-fix.json
task_id=$(basename "$demo_home"/tasks/*)
agentos status "$task_id"
agentos export "$task_id" "$demo_home/bundle"
printf 'Export: %s/bundle\n' "$demo_home"
The fixture makes six model calls and settles five tool actions. Its final verification pins the workspace digest and the protected profile; the export includes the patch, requests, responses and verification evidence. The original crash simulation below still illustrates the deterministic patch-agent run.
Retry-After can extend them. Every new attempt gets its own effect and reservation.New submissions record the endpoint and policy versions. Resume refuses a different provider endpoint before reading credentials or changing the task. Older journals retain their recorded behavior. Version-2 recordings preserve ordered attempts, including transient rejections and retry metadata; legacy response-only fixtures still replay.
docker compose build test sh scripts/acceptance.sh offline
This runs the host and fake-jail harnesses against a loopback fake API, including a 429 followed by repair, then replays without network. It compares request digests, call counts, final workspace, exported patch and protected profile. It needs no real API key or /dev/kvm. Fake-jail tests exercise the harness; they do not prove VM isolation.
Real host and jailed provider runs, real KVM tests and candidate guest-image acceptance remain pending. See the acceptance setup and evidence requirements before enabling those tiers.
Your Python service reads settings such as PORT = 8080. A customer reports that spaces around the equals sign break the configuration. Give the agent this narrow bug, allow edits only under src/, and let a protected regression check decide whether the fix works. The result is a reviewable export; applying it to your project is a separate step you control.
Start with a clean source snapshot and a check that fails for the reported input. Register that check separately from the editable code. The practice run below uses the included parser as the service's configuration reader.
Submit without --yes. Inspect the goal, editable paths, protected profile, worker and budgets. The task stays READY; resume approves this recorded task and starts it.
Check status and export the result. Require SUCCEEDED, matching final and verified workspace digests, and a passing verification accepted for the final workspace. Read the diff and the verification results before accepting the change.
Run git apply --check, apply the reviewed export, and rerun your checks. The practice uses a disposable review copy. The original source snapshot stays unchanged throughout the agent run.
Save the contract above as model-task.json, then paste this in docker compose run --rm test sh. The scripted answers are specific to this parser example; an arbitrary project needs a live provider.
cargo build --locked -q -p agentos-cli
mkdir -p build
practice=${PRACTICE_DIR:-$(mktemp -d /work/build/usage.XXXXXX)}
mkdir -p "$practice"
cp -R fixtures/parser-repo "$practice/service"
cp -R "$practice/service" "$practice/review"
git init -q "$practice/review"
check=fixtures/profiles/parser-checks-v1/check_parser.py
# The protected check must expose the bug before we ask for a fix.
if pyenv exec python "$check" "$practice/service"; then
echo 'Expected the original parser to fail' >&2; exit 1
else
[ "$?" -eq 1 ]
fi
pyenv exec python - "${MODEL_CONTRACT:-model-task.json}" "$practice" <<'PY'
import json, sys
from pathlib import Path
root = Path(sys.argv[2]).resolve()
task = json.loads(Path(sys.argv[1]).read_text())
task["goal"] = ("Fix config parsing: trim keys and values, "
"skip indented comments, preserve equals signs in values.")
task["repository"]["source"] = str(root / "service")
(root / "task.json").write_text(json.dumps(task))
PY
agentos() {
target/debug/agentos --home "$practice/home" --worker host "$@"
}
agentos profile register fixtures/profiles/parser-checks-v1
agentos submit "$practice/task.json" \
--model fake:fixtures/transcripts/parser-fix.json > "$practice/approval.json"
task_id=$(pyenv exec python -c \
'import json,sys; print(json.load(open(sys.argv[1]))["task_id"])' \
"$practice/approval.json")
agentos status "$task_id" > "$practice/ready-status.json"
cat "$practice/ready-status.json"
Stop here and review the printed permissions. The task is READY and has run no effects. When you approve it, continue in the same shell:
agentos resume "$task_id" agentos events "$task_id" > "$practice/events.ndjson" agentos export "$task_id" "$practice/bundle" cat "$practice/bundle/patch.diff" # Apply only to this disposable review copy after inspecting the diff. git -C "$practice/review" apply --check "$practice/bundle/patch.diff" git -C "$practice/review" apply "$practice/bundle/patch.diff" pyenv exec python "$check" "$practice/review" printf 'Task, evidence and review copy: %s\n' "$practice"
Prepare a clean snapshot of the committed revision, replace the source path and goal in the contract, and register your project's regression checks as a protected profile. Keep credentials outside the snapshot: Agent OS copies regular files, including untracked files, and does not use .gitignore as a filter. The current Python profile supports standard-library checks; projects needing extra runtimes or packages need a matching execution image.
For a live task, mount the snapshot read-only and a regular key file read-only into the container, then submit with --model anthropic:MODEL_ID --api-key-file /run/agentos/key. Choose an available ID from Anthropic's model list. Submit first, approve with resume after reviewing permissions, and use the same export/review flow. Repository content included in model requests goes to the selected provider; live calls consume your API account's budget.
Use the host worker for a trusted first test; it is not sandboxed. For untrusted checks, prepare the jailed Firecracker worker on a Linux KVM host. Live-provider and real KVM acceptance remain pending. The hands-on testing guide includes setup, the own-repository commands, expected results, recovery tests and the later KVM/live gates.
Every effect (read the snapshot, apply the patch, run the check) is written to the journal as separate steps: intended, dispatched, executed, published, completed. Each effect runs as a job under its own supervisor process, so the job keeps going when the controller dies. A job ends by writing a receipt (its result) into its own directory, and a lease is the time the supervisor allows it before killing it. Pick a moment and kill the controller there.
Choose a moment, then press Run and kill here. The default is the exact moment recorded in the real demo: the patch job is dispatched and the controller dies.
The sequence matches the real run of scripts/demo.sh. Crash points other than the recorded one are modelled from the recovery rules.
The check that verifies a task is code from the registered verification profile, and it runs as untrusted. The gated KVM suite includes hostile profiles that try to burn CPU, eat memory, fork without limit, fill the disk, reach the network and find a host secret. These controls illustrate its expected assertions; this build has not rerun them on real KVM. Pick one.
One honest boundary: a bug in Firecracker or KVM could let code escape the VM. It would land in the jail (an unprivileged uid inside a chroot and cgroup, under Firecracker's own seccomp filter), and a second kernel-level bug would be needed to reach the host. With --allow-unjailed the escape lands as the controller's own user instead.
All four required offline gates pass: formatting, strict Clippy, the host suite and the fake-jail suite. Offline and skipped gated tests do not establish real provider or KVM acceptance. The new Python guest is a candidate; the legacy image remains the default.
Completion plan · Verification and review decisions
--worker firecracker./dev/kvm. x86_64 only.