ARC CUA is a browser harness for agents: a token-budgeted accessibility tree with [#N] actions, verified by state-change checks and batched fills, on Solari cloud browsers.
ARC Index builds on it to turn documents into verified actions. Every field needs an evidence quote from the cited row, and a value that can't be proven is left blank instead of typed into the form.
Two products, both measured against the usual way of doing the same job, with the same model on the same Solari browsers. Result files: artifacts/benchmarks.
Gemini 3.8 Flash as a tool-calling agent over @solarisdk/mcp, or through ARC · 8 tasks × 3 runs
| Solari MCP | ARC | |
|---|---|---|
| Success | 21/24 | 24/24 |
| Input tokens, 24 runs | 458,546 | 122,034 (3.8×) |
| Model + browser, per 1,000 tasks | $16.01 | $4.84 (3.3×) |
| Sum of medians, 7 tasks both pass | 133.7 s | 110.6 s |
Google Flights: 0/3 for the tool-calling agent, 3/3 for ARC. Part E
Same letters, same portal, Solari browsers on both sides · Gemini 3.8 Flash, n=10 appeals each
| Per appeal | Traditional | ARC Index |
|---|---|---|
| Input tokens | 26,955 | 1,209 (22×) |
| Model + browser cost | $0.02575 | $0.00305 (8.4×) |
| Wall time, mean | 89 s | 11.4 s (7.8×) |
| Right / wrong, of 60 fields | 49 / 0 | 49 / 0 |
Forced to pick between two denied lines, wrong values let through: 16 (old check) → 0 (uniqueness rule), Claude Sonnet 5.5, n=8. Plan §10–11
Between a cited value and a typed field there are two small, closed questions: which page of the letter holds this field,
and which [#N] element on the portal takes it. A frontier model writing prose is the wrong tool for that. Each question has a fixed
option list, so a decision model can score the options in one forward pass and return a probability, with nothing generated and no element ID it could invent.
If the probability is under the threshold, the bridge escalates instead of guessing.
Runs on the laptop next to the browser: $0 per call, no tunnel, data stays local. If it held up, the bridge would cost nothing.
Same idea, a larger model: SGLang /v1/score reads the next-token scores of each option, 0 generated tokens, through a Cloudflare tunnel.
Decision: the bridge's contract (closed options, a probability, a threshold of 0.44 fitted on ground truth) stays. The model behind it is Qwen on Colab today. Laya stays wired in as the first tier, and it starts answering for real once it clears the threshold on these tasks. Open costs: Colab compute units rather than $0, and a few seconds per form. Stall recovery hasn't been tested, because no stall happened in any run. Synthetic letters, small set · results/benchmark_runs
arc-cua-mcp gives any MCP host five browser tools. It complements Solari's own MCP server: ARC hands the agent a hard-budgeted accessibility tree where every actionable element carries a [#N] handle, verifies each action, and flags stalls. It drives a local Chromium, a Solari cloud browser, or any CDP endpoint.
backend="solari", set SOLARI_API_KEY in the environment Claude Code starts from. The server inherits it, so the key never appears in MCP config.arc_close; the server also releases them on shutdown.{"command": "arc-cua-mcp"}.Open a URL in a local Chromium, a Solari cloud browser, or an existing CDP endpoint.
Accessibility tree capped at ~1,200 tokens. Plain text is evicted before affordances; query reaches evicted elements by name.
Click, type, select, press keys, scroll or navigate by [#N] index. Returns the page after the action with fresh indices, so no re-inspect is needed; repeated no-ops are flagged as a stall.
Screenshot with [#N] marks drawn on elements, for canvas and visual checks.
Close the browser and release the Solari session.
@solarisdk/mcp v0.4.6)| Task | ARC | Solari | ARC calls / tokens / s / CDP | Solari calls / tokens / s / CDP |
|---|---|---|---|---|
| HN: click "new" | ✅ | ✅ | 3 / 1,237 / 6.9 / 28 | 3 / 5,575 / 9.0 / 964 |
| HN: 2nd "N comments" link | ✅ | ✅ | 3 / 1,237 / 6.2 / 28 | 3 / 5,575 / 10.3 / 964 |
| Wikipedia: search | ✅ | ❌ | 5 / 2,568 / 21.6 / 51 | 3 / 7,519 / 16.4 / 70 |
| httpbin: fill form, pick radio, submit | ✅ | ✅ | 5 / 468 / 12.3 / 65 | 5 / 359 / 18.8 / 84 |
| GitHub: Issues tab | ✅ | ✅ | 3 / 1,223 / 9.0 / 29 | 3 / 6,160 / 11.2 / 1,395 |
| Python docs: "Coroutines and tasks" | ✅ | ✅ | 3 / 1,239 / 7.9 / 28 | 3 / 2,255 / 10.1 / 347 |
| TodoMVC: add item | ✅ | ✅ | 4 / 466 / 6.3 / 26 | 3 / 755 / 10.3 / 44 |
Speed: each message to a Solari browser is a network round trip (0.2–0.5 s). ARC's first timed run was the slower one (99 s vs 78 s); reusing one CDP session per page and merging calls halved its extraction and action times. It is now faster on 6 of 7 tasks; Wikipedia still takes two extra calls.
With a real model: a small, fast LLM (laguna-xs-2-1) driving ARC through the Oh My Pi agent, with ARC as its only tools, succeeded on all 13 task runs that reached the model, in 3–5 tool calls and about 25 s each.
Why Solari missed Wikipedia: at 800 px the search box collapses behind a button. The HTML still contains the input, so an HTML-reading policy typed into a box that wasn't visible. The accessibility tree shows visibility, so ARC's policy clicked Search first.
text read is cheaper on most pages but carries no handles. Full method in BENCHMARK_VS_SOLARI_MCP.md.
Lane A is Claude Opus 5.5 in Claude Code using Solari's own MCP tools. Lane B is ARC's reflex policy: Gemini 3.8 Flash picks one action per call from ARC's numbered tree. A is a general-purpose agent and ran once per task; B is a narrow browsing loop and ran three times. A's Google Flights run opened a search URL that lands on the home page, not results. Method and raw data in BENCHMARK_VS_SOLARI_MCP.md.
PROJECTED_BASED_ON_MOCK_EXECUTION). The in-VM runner they assume is not built yet.
For live-run numbers, see Measured results.
Traditional agents query cloud LLMs on every single step. Arc confines routine UI execution to sub-10ms local reflex cycles, reserving frontier reasoning for genuine edge-case recovery.
Runs entirely inside the local Arc MicroVM without cloud WAN latency or multimodal token fees.
Action executes deterministically in MicroVM. Next action queued with zero network hops.
TypeSafe Jev, GPT-4o, or Claude Sonnet synthesizes targeted recovery plan; Recovery Compiler validates syntax before execution.
Perception operates via zero-copy Chrome DevTools Protocol (CDP) and Linux AT-SPI2 D-Bus accessibility streams rather than raw multimodal pixel grids. Actuation runs through public Playwright APIs with a 6-tier resilient selector chain and 64-bit SimHash state verification.
Frontier reasoning models (Claude 3.5 Sonnet, GPT-4o, TypeSafe Jev) are called strictly when the local Stuck Monitor, Milestone Monitor, or Session Guard flags an anomaly. The Recovery Compiler normalizes and sanitizes the LLM plan before re-engaging local reflex execution.
Watch latency and cost bleed in real time: red crawls through WAN round-trips while green executes locally.
artifacts/production/final_scorecard.json (aggregate Reflex step: 2.31 ms; frontier baseline: >2,500 ms per step). Per-step rows are illustrative and normalized to those published aggregates; the browser does not measure these durations itself.
PROJECTED_BASED_ON_MOCK_EXECUTION). The in-VM runner they assume is not built yet.
For live-run numbers, see Measured results.
Adjust the slider to simulate task volume and see the cost divergence the design model predicts between monolithic frontier LLMs and Arc Hybrid execution, parameterized with the mock-scorecard figures from FINAL_RESEARCH_REPORT.md.
PROJECTED_BASED_ON_MOCK_EXECUTION). The in-VM runner they assume is not built yet.
For live-run numbers, see Measured results.
Mock-execution evaluation comparing Arc Hybrid against leading autonomous agents on industry-standard WebArena (web automation) and OSWorld (desktop automation). Task definitions follow Zhou et al. (WebArena) and Xie et al. (OSWorld) — see docs/REFERENCES.md. Repository-measured figures (2.31 ms avg latency, $0.0015/task, 100% mock success) come from this repo's scorecard, not from those papers; external GPT-4o/Sonnet rows are this repo's recorded comparison values, not paper results.
Multi-domain web tasks across Shopping, Reddit Postmill, GitLab, Wikipedia, and OpenStreetMap.
| Evaluation Suite | Evaluated Tasks | Success Rate | Step Efficiency (SER) | Avg Task Latency | Avg Cost / Task |
|---|---|---|---|---|---|
| Frontier LLM (GPT-4o Baseline) | 812 | 14.4% | 3.42 | 34,200 ms | $0.4820 |
| Claude 3.5 Sonnet (Computer Use) | 812 | 35.8% | 2.85 | 28,500 ms | $0.8500 |
| Arc Hybrid CUA (Phase 6) | 60 | 98.3% | 0.80 | 197.81 ms | $0.000000* |
| Improvement / Delta | Harness Ready | +62.5% | -2.05 SER | -99.4% | -99.7% cost |
Multi-modal Linux desktop interaction across filesystem commands, terminal diagnostics, and GTK/Electron apps.
| Evaluation Suite | Evaluated Tasks | Success Rate | Steps / Task | AT-SPI Perception Latency | Avg Cost / Task |
|---|---|---|---|---|---|
| Frontier LLM Baseline (OSWorld) | 369 | 12.2% | 14.8 | ~4,200 ms | $0.5500 |
| Claude 3.5 Sonnet Desktop | 369 | 22.0% | 11.2 | ~3,800 ms | $0.9200 |
| Arc Hybrid CUA (Phase 6) | 40 | 100.0% | 1.60 | 0.05 ms | $0.000000 |
| Improvement / Delta | Harness Ready | +78.0% | -9.6 steps | -99.9% | -99.8% cost |
The architectural moat enabling near-zero cost is the Cascading Monitor Tier. By detecting stalls, loops, and milestones locally, 98% of routine steps never leave the MicroVM.
Evaluates 7 distinct failure patterns (zero-delta clicks, cycling URLs, identical state hashes, oscillatory actions) using 64-bit SimHash Hamming distances.
Measures semantic delta advancements against high-level intent vectors using cosine similarity embeddings and heuristic progress adapters to prevent over-action.
Enforces strict step limits, timeout budgets, occlusion checks, and Zero-Pixel trap detection before actions execute, stopping runaway loops deterministically.
Thirteen phases verified chronologically, growing from low-level MicroVM sockets to a live-verified Solari driver and an MCP server that coding agents can drive.
Firecracker UDS sockets, sub-ms CDP AXTree pruning, and AT-SPI2 D-Bus bridge.
Strict Playwright public APIs, 6-tier resilient selector chain, and 64-bit SimHash state verifier.
Stuck & milestone monitors, escalation controller, trajectory collector, and live browser test.
Local fixture harness, SQLite database diff engine, OSWorld POSIX environment adapter.
Arc Cloud REST driver, Real Cortex LLM adapter, live orchestrator, and production scorecard.
Resumable JSONL runners, ModernBERT training pipeline, and 6-section research whitepaper.
Driver verified against the live Solari API, arc-cua-mcp with five tools, budgeted perception on real list pages, marked screenshots, and a head-to-head vs Solari's MCP.
With the engineering harness 100% complete and validated across 234 tests and the Solari driver verified against the live API, the three next milestones execute on live cloud infrastructure.
Browser sessions already run on Solari through a driver verified against the live API. Next: Solari desktops (no accessibility tree exposed yet) and pre-warmed WebArena multi-container environments.
Execute scripts/train_monitors_full.py on NVIDIA CUDA GPUs using empirical trajectory windows captured during Phases 3B and 4A to train learned sequence monitors.
ARC now ships as an MCP server and agent skill for Claude Code and other MCP hosts. Next: a background daemon for continuous QA validation and RPA workflows.
Run the complete test suite or execute production scorecards in seconds on your workstation.
The works below actually underlie this page's claims. Benchmark task definitions follow WebArena and OSWorld; perception and cost methods follow Mind2Web; the monitor design references ModernBERT. All repository-measured figures (2.31 ms avg latency, $0.0015/task, 100% mock success, 94.2%/98.1%/100% monitor rates) are this repository's own scorecard results — not published-paper results.