MEASURED ON SOLARI · 2026-09-29 | NOW AN MCP SERVER FOR CLAUDE CODE

ARC on Solari Verified actions, measured

ARC CUA is a browser harness for agents: a token-budgeted accessibility tree with [#N] actions, verified by state-change checks and batched fills, on Solari cloud browsers. ARC Index builds on it to turn documents into verified actions. Every field needs an evidence quote from the cited row, and a value that can't be proven is left blank instead of typed into the form.

8.4×
Cheaper per appeal
$0.00305 vs $0.02575 · ARC Index vs traditional agent
24 / 24
vs 21/24 on Solari MCP
Same model · 3.3× cheaper per task
16 → 0
Wrong values let through
Forced pick, two denied lines
arc-trace (illustrative)
● MOCK TRACE
// Illustrative trace from the mock design model (not timed on Solari)
[0.018ms] AT-SPI2 D-Bus: Accessible desktop tree serialized
[0.024ms] CDP AXTree: Sanitized DOM accessibility tree (41 nodes)
[1.050ms] ReflexRunner: Public Playwright click #checkout-btn
[1.053ms] StateVerifier: 64-bit SimHash Hamming dist = 14 (STATE_CHANGED)
[1.095ms] StuckMonitor: 7-pattern check = HEALTHY (p=0.012)
[1.110ms] MilestoneMonitor: Goal progress delta = +0.85
Step Routing Verdict:
Local Reflex handled step without cloud reasoning.
Marginal Cost: $0.000000 Escalate: NO (98% Local)
// Measured on Solari, 2026-09-29
→ fill_many: 7 fields in one call, 386 ms
→ appeal filed: 11.4 s (traditional agent: 89 s)
[PASS] wrong values let through: 0 · confirmation matched
Live runs on Solari cloud browsers

Measured results

Two products, both measured against the usual way of doing the same job, with the same model on the same Solari browsers. Result files: artifacts/benchmarks.

ARC CUA · browser harness

vs Solari's own MCP, same model

Gemini 3.8 Flash as a tool-calling agent over @solarisdk/mcp, or through ARC · 8 tasks × 3 runs

Solari MCPARC
Success21/2424/24
Input tokens, 24 runs458,546122,034 (3.8×)
Model + browser, per 1,000 tasks$16.01$4.84 (3.3×)
Sum of medians, 7 tasks both pass133.7 s110.6 s

Google Flights: 0/3 for the tool-calling agent, 3/3 for ARC. Part E

ARC Index · documents → verified actions

vs a traditional browser agent, same denial appeal

Same letters, same portal, Solari browsers on both sides · Gemini 3.8 Flash, n=10 appeals each

Per appealTraditionalARC Index
Input tokens26,9551,209 (22×)
Model + browser cost$0.02575$0.00305 (8.4×)
Wall time, mean89 s11.4 s (7.8×)
Right / wrong, of 60 fields49 / 049 / 0

Forced to pick between two denied lines, wrong values let through: 16 (old check) → 0 (uniqueness rule), Claude Sonnet 5.5, n=8. Plan §10–11

The bridge · System-1 / System-2 decisions

ARC Index knows the value. ARC knows the page. The bridge decides where it goes.

Between a cited value and a typed field there are two small, closed questions: which page of the letter holds this field, and which [#N] element on the portal takes it. A frontier model writing prose is the wrong tool for that. Each question has a fixed option list, so a decision model can score the options in one forward pass and return a probability, with nothing generated and no element ID it could invent. If the probability is under the threshold, the bridge escalates instead of guessing.

System-1 · Laya 421M, local · why we tested it

Runs on the laptop next to the browser: $0 per call, no tunnel, data stays local. If it held up, the bridge would cost nothing.

  • 38 ms on an L4 · 565 ms on an 8-thread CPU
  • 15% of page lookups right
  • kept 24% of benchmark field picks, 18/18 right
  • kept 0 decisions in the live browser run
System-2 · Qwen3.8-27B on Colab L4 · why it answers today

Same idea, a larger model: SGLang /v1/score reads the next-token scores of each option, 0 generated tokens, through a Cloudflare tunnel.

  • 100% of page lookups right · MRR 0.738 → 1.000
  • tasks fully right: 0% → 91.7% (keyword → hybrid)
  • reworded live form: 3/6 → 6/6 fields
  • 296 ms per lookup · 4.6–6.0 s to match one form

Decision: the bridge's contract (closed options, a probability, a threshold of 0.44 fitted on ground truth) stays. The model behind it is Qwen on Colab today. Laya stays wired in as the first tier, and it starts answering for real once it clears the threshold on these tasks. Open costs: Colab compute units rather than $0, and a few seconds per form. Stall recovery hasn't been tested, because no stall happened in any run. Synthetic letters, small set · results/benchmark_runs

The bridge, ARC Index side: which page holds the field
Local Laya vs Qwen on Colab, same question: 15% vs 100% of pages right
The bridge, ARC side: which [#N] field takes the value
Reworded live form: 3/6 → 6/6 fields, every decision escalated to Colab
Economics: traditional agent vs ARC Index
8.4× cheaper and 7.8× faster per appeal, same accuracy
Wrong values let through
A confident pick between two denied lines, and the rule that rejects it
ARC Index: one value, letter to receipt
Cited row, batched fill on Solari, SimHash check, receipt
Same model, two harnesses
Gemini 3.8 Flash on Solari's MCP vs on ARC
Phases 9–13 · Agent Integration

Use ARC from Claude Code

arc-cua-mcp gives any MCP host five browser tools. It complements Solari's own MCP server: ARC hands the agent a hard-budgeted accessibility tree where every actionable element carries a [#N] handle, verifies each action, and flags stalls. It drives a local Chromium, a Solari cloud browser, or any CDP endpoint.

Install
# Package with the MCP extra
pip install -e ".[mcp]"
playwright install chromium
# Register with Claude Code
claude mcp add --scope user arc -- arc-cua-mcp
  • For backend="solari", set SOLARI_API_KEY in the environment Claude Code starts from. The server inherits it, so the key never appears in MCP config.
  • Solari browsers bill hourly until arc_close; the server also releases them on shutdown.
  • Other MCP hosts: same command over stdio, {"command": "arc-cua-mcp"}.
arc_open

Open a URL in a local Chromium, a Solari cloud browser, or an existing CDP endpoint.

arc_inspect

Accessibility tree capped at ~1,200 tokens. Plain text is evicted before affordances; query reaches evicted elements by name.

arc_act

Click, type, select, press keys, scroll or navigate by [#N] index. Returns the page after the action with fresh indices, so no re-inspect is needed; repeated no-ops are flagged as a stall.

arc_screenshot

Screenshot with [#N] marks drawn on elements, for canvas and visual checks.

arc_close

Close the browser and release the Solari session.

Head-to-head vs Solari's MCP (@solarisdk/mcp v0.4.6)

Both on Solari cloud browsers · 800×600 · 2026-09-27
7 / 7 vs 6 / 7
Grounding tasks succeeded
3.3× fewer tokens
8,438 vs 28,198 perception tokens
15× fewer CDP
255 vs 3,868 protocol commands
70 s vs 86 s
Time in tool calls
TaskARCSolariARC calls / tokens / s / CDPSolari calls / tokens / s / CDP
HN: click "new"✅✅3 / 1,237 / 6.9 / 283 / 5,575 / 9.0 / 964
HN: 2nd "N comments" link✅✅3 / 1,237 / 6.2 / 283 / 5,575 / 10.3 / 964
Wikipedia: search✅❌5 / 2,568 / 21.6 / 513 / 7,519 / 16.4 / 70
httpbin: fill form, pick radio, submit✅✅5 / 468 / 12.3 / 655 / 359 / 18.8 / 84
GitHub: Issues tab✅✅3 / 1,223 / 9.0 / 293 / 6,160 / 11.2 / 1,395
Python docs: "Coroutines and tasks"✅✅3 / 1,239 / 7.9 / 283 / 2,255 / 10.1 / 347
TodoMVC: add item✅✅4 / 466 / 6.3 / 263 / 755 / 10.3 / 44

Speed: each message to a Solari browser is a network round trip (0.2–0.5 s). ARC's first timed run was the slower one (99 s vs 78 s); reusing one CDP session per page and merging calls halved its extraction and action times. It is now faster on 6 of 7 tasks; Wikipedia still takes two extra calls.

With a real model: a small, fast LLM (laguna-xs-2-1) driving ARC through the Oh My Pi agent, with ARC as its only tools, succeeded on all 13 task runs that reached the model, in 3–5 tool calls and about 25 s each.

Why Solari missed Wikipedia: at 800 px the search box collapses behind a button. The HTML still contains the input, so an HTML-reading policy typed into a box that wasn't visible. The accessibility tree shows visibility, so ARC's policy clicked Search first.

CAVEATS: scripted policies stand in for an agent (no LLM), one run each on public sites that change daily. Token counts are what each tool hands the agent; Solari's text read is cheaper on most pages but carries no handles. Full method in BENCHMARK_VS_SOLARI_MCP.md.

Opus + Solari MCP vs Gemini 3.8 Flash + ARC, in 15 seconds

Same Solari browser · 8 tasks · 2026-09-27
24/24 vs 7/8
Tasks passed
1.9×
Faster · 172 s vs 323 s
39×
Fewer input tokens
33×
Cheaper · $4.83 vs $160 / 1k tasks

Lane A is Claude Opus 5.5 in Claude Code using Solari's own MCP tools. Lane B is ARC's reflex policy: Gemini 3.8 Flash picks one action per call from ARC's numbered tree. A is a general-purpose agent and ran once per task; B is a narrow browsing loop and ran three times. A's Google Flights run opened a search URL that lands on the home page, not results. Method and raw data in BENCHMARK_VS_SOLARI_MCP.md.

Agent Skill Solari API Contract View Source
Design model, not live results. The figures in this section (routing percentages, per-step costs and latencies) come from the repo's mock-execution scorecard (PROJECTED_BASED_ON_MOCK_EXECUTION). The in-VM runner they assume is not built yet. For live-run numbers, see Measured results.
Hierarchical Execution

Reflex + Cortex Cascading Architecture

Traditional agents query cloud LLMs on every single step. Arc confines routine UI execution to sub-10ms local reflex cycles, reserving frontier reasoning for genuine edge-case recovery.

Tier 1: Local Fast Path (Reflex)
Tier 2: Cloud Cortex (Frontier LLM)
User Intent & Benchmark Task
Goal & Instruction Input
TIER 1: REFLEX ENGINE (LOCAL FAST PATH)
2.31 ms Avg Latency • $0.000000 Cost

Deterministic UI Perception & Actuation

Runs entirely inside the local Arc MicroVM without cloud WAN latency or multimodal token fees.

CDP AXTree
p50: 0.024ms
AT-SPI2 D-Bus
p50: 0.018ms
State Verifier
64-bit SimHash
Playwright API
0 AST Leaks
Cascading Perception & Health Gatekeepers
Stuck Monitor (94.2%) Milestone Monitor (98.1%) Session Guard (100%)
HEALTHY STATE (98.0% OF ALL STEPS)

Advance Next Action Locally

Action executes deterministically in MicroVM. Next action queued with zero network hops.

Execution Cost: $0.000000 Latency: 2.31 ms
ANOMALY DETECTED (2.0% OF STEPS)

Tier 2: Cloud Cortex Escalation

Frontier LLM

TypeSafe Jev, GPT-4o, or Claude Sonnet synthesizes targeted recovery plan; Recovery Compiler validates syntax before execution.

Recovery Cost: ~$0.001500 Resume Reflex Queue →
Tier 1: Deterministic Reflex (98% Traffic)

Perception operates via zero-copy Chrome DevTools Protocol (CDP) and Linux AT-SPI2 D-Bus accessibility streams rather than raw multimodal pixel grids. Actuation runs through public Playwright APIs with a 6-tier resilient selector chain and 64-bit SimHash state verification.

Tier 2: Escalation Cortex (2% Traffic)

Frontier reasoning models (Claude 3.5 Sonnet, GPT-4o, TypeSafe Jev) are called strictly when the local Stuck Monitor, Milestone Monitor, or Session Guard flags an anomaly. The Recovery Compiler normalizes and sanitizes the LLM plan before re-engaging local reflex execution.

02 — LIVE OBSERVATION–ACTION LOOP COMPARISON

Side-by-Side Execution Timing

Watch latency and cost bleed in real time: red crawls through WAN round-trips while green executes locally.

VISUALIZED BENCHMARK — PROJECTED_BASED_ON_MOCK_EXECUTION. Timing behavior is animated from the mock-scorecard parameters recorded in artifacts/production/final_scorecard.json (aggregate Reflex step: 2.31 ms; frontier baseline: >2,500 ms per step). Per-step rows are illustrative and normalized to those published aggregates; the browser does not measure these durations itself.
Traditional Frontier CUA
SCREENSHOT → WAN UPLOAD → VLM INFERENCE
STEP TIMER
0 ms
PIPELINE STAGE NORMALIZED SHARE · 2,850 ms CYCLE
01 Full display buffer raster capture
320 ms
02 PNG upload over WAN to cloud provider
410 ms
03 Multimodal 70B+ parameter VLM inference
1,840 ms · $0.048
04 Coordinate parse & actuation download
190 ms
05 Blind mouse click dispatch (drift prone)
90 ms
STEP CYCLE PROGRESS 0%
[LATENCY] Total step latency: ~2,850 ms · Cost per task: ~$0.4820
Arc Hybrid CUA
LOCAL CDP/DOM → SIMHASH VERIFY → ZERO WAN
COMPLETED LOOPS
0
PIPELINE STAGE MEASURED AVG · 2.31 ms AGGREGATE
01 Zero-copy CDP AXTree / AT-SPI extraction
0.024 ms
02 6-tier resilient selector LRU cache lookup
0.001 ms
03 Public Playwright UI actuation
1.050 ms
04 64-bit SimHash Hamming state verification
0.003 ms
05 Stuck & Milestone monitor health gating
0.047 ms
STEP EXECUTION LATENCY 2.31 ms (AVG)
[CONFIRMED] Local step cost: $0.000000 98.0% steps handled off-API
Design model, not live results. The figures in this section ($0.0015 vs $0.4820 per task, 2.31 ms vs 2,850 ms per step) come from the repo's mock-execution scorecard (PROJECTED_BASED_ON_MOCK_EXECUTION). The in-VM runner they assume is not built yet. For live-run numbers, see Measured results.
Live Economic Model

Interactive Cost & Latency Simulator

Adjust the slider to simulate task volume and see the cost divergence the design model predicts between monolithic frontier LLMs and Arc Hybrid execution, parameterized with the mock-scorecard figures from FINAL_RESEARCH_REPORT.md.

Simulates agent interaction steps across production workflows
1,000 steps
Quick Presets:
Frontier LLM Baseline
$48.20
Cost at ~$0.482 / 10 steps
Avg Latency: 2,850 ms / step
Arc Hybrid CUA
$0.15
Cost at ~$0.0015 / 10 steps
Avg Latency: 2.31 ms / step
Net Capital Savings
$48.05
Operational budget preserved
Efficiency Delta: 99.69% Margin
Cumulative Time Saved
47.4 min
Zero WAN round-trip lag
Throughput Speedup: 1,233× faster
Cumulative Execution Cost ($ USD) vs Step Volume
Frontier Baseline Arc Hybrid
Design model, not live results. The figures in this section (WebArena and OSWorld 100% success, per-task costs) come from the repo's mock-execution scorecard (PROJECTED_BASED_ON_MOCK_EXECUTION). The in-VM runner they assume is not built yet. For live-run numbers, see Measured results.
Standardized Evaluation

Autonomous Benchmark Results

Mock-execution evaluation comparing Arc Hybrid against leading autonomous agents on industry-standard WebArena (web automation) and OSWorld (desktop automation). Task definitions follow Zhou et al. (WebArena) and Xie et al. (OSWorld) — see docs/REFERENCES.md. Repository-measured figures (2.31 ms avg latency, $0.0015/task, 100% mock success) come from this repo's scorecard, not from those papers; external GPT-4o/Sonnet rows are this repo's recorded comparison values, not paper results.

STATUS: PROJECTED_BASED_ON_MOCK_EXECUTION (Workstation Host Evaluation)

WebArena-Verified Web Automation 812 Tasks Total

Multi-domain web tasks across Shopping, Reddit Postmill, GitLab, Wikipedia, and OpenStreetMap.

+62.5% Success Delta vs GPT-4o
Evaluation Suite Evaluated Tasks Success Rate Step Efficiency (SER) Avg Task Latency Avg Cost / Task
Frontier LLM (GPT-4o Baseline) 812 14.4% 3.42 34,200 ms $0.4820
Claude 3.5 Sonnet (Computer Use) 812 35.8% 2.85 28,500 ms $0.8500
Arc Hybrid CUA (Phase 6) 60 98.3% 0.80 197.81 ms $0.000000*
Improvement / Delta Harness Ready +62.5% -2.05 SER -99.4% -99.7% cost
* $0.000000 reflects deterministic local reflex actions with mock cortex. Full production cloud run with microVM & proxy storage is $0.001504/task.

OSWorld Desktop Environment Suite 369 Tasks Total

Multi-modal Linux desktop interaction across filesystem commands, terminal diagnostics, and GTK/Electron apps.

+78.0% Success Delta vs GPT-4o
Evaluation Suite Evaluated Tasks Success Rate Steps / Task AT-SPI Perception Latency Avg Cost / Task
Frontier LLM Baseline (OSWorld) 369 12.2% 14.8 ~4,200 ms $0.5500
Claude 3.5 Sonnet Desktop 369 22.0% 11.2 ~3,800 ms $0.9200
Arc Hybrid CUA (Phase 6) 40 100.0% 1.60 0.05 ms $0.000000
Improvement / Delta Harness Ready +78.0% -9.6 steps -99.9% -99.8% cost
Gatekeeper Intelligence

Monitor Efficacy & Escalation Prevention

The architectural moat enabling near-zero cost is the Cascading Monitor Tier. By detecting stalls, loops, and milestones locally, 98% of routine steps never leave the MicroVM.

Stuck Monitor p95 < 0.05ms

Mechanical Stall & Loop Prevention

Evaluates 7 distinct failure patterns (zero-delta clicks, cycling URLs, identical state hashes, oscillatory actions) using 64-bit SimHash Hamming distances.

Escalations Prevented 94.2%
Method: State Hashing + Hamming Active
Milestone Monitor p95 < 0.05ms

Goal Progress Advancement

Measures semantic delta advancements against high-level intent vectors using cosine similarity embeddings and heuristic progress adapters to prevent over-action.

Escalations Prevented 98.1%
Method: Semantic Progress Adapter Active
Session Guard Zero-Pixel Trap

Invariant & Boundary Enforcement

Enforces strict step limits, timeout budgets, occlusion checks, and Zero-Pixel trap detection before actions execute, stopping runaway loops deterministically.

Escalations Prevented 100.0%
Method: Rate & Invariant Guard Active
98%
98 Out of 100 Steps Executed Off-API
Routine web & desktop actions run completely free on local micro-runtimes.
Recovery Resolution: Avg 1 Step
Development History

Phase-by-Phase Engineering Timeline

Thirteen phases verified chronologically, growing from low-level MicroVM sockets to a live-verified Solari driver and an MCP server that coding agents can drive.

Phase 1 · Complete

MicroVM Foundation & Perceptual Extraction

Firecracker UDS sockets, sub-ms CDP AXTree pruning, and AT-SPI2 D-Bus bridge.

AT-SPI p50: 0.018ms · CDP p50: 0.024ms
1
25 Tests Passing
Phase 2 · Complete

Deterministic Reflex Engine & Public Actuation

Strict Playwright public APIs, 6-tier resilient selector chain, and 64-bit SimHash state verifier.

Reflex Step p50: 1.05ms · 0 Private Internals
2
52 Tests Passing
Phase 3 · Complete

Monitor Infrastructure & Recovery Compiler

Stuck & milestone monitors, escalation controller, trajectory collector, and live browser test.

Stuck p95: 0.0475ms · Compiler p95: 0.009ms
3
85 Tests Passing
Phase 4 · Complete

Benchmark Integration (WebArena & OSWorld)

Local fixture harness, SQLite database diff engine, OSWorld POSIX environment adapter.

Hybrid Success: 92.3% · Assertion p95: 0.005ms
4
125 Tests Passing
Phase 5 · Complete

Production Deployment & Live Cloud Driver

Arc Cloud REST driver, Real Cortex LLM adapter, live orchestrator, and production scorecard.

Production Success: 100% · SER: 0.80
5
143 Tests Passing
Phase 6 · Complete

Full-Scale Benchmark Scaling & Architecture Report

Resumable JSONL runners, ModernBERT training pipeline, and 6-section research whitepaper.

Cost Reduction: 99.69% · Latency: -99.90% (mock scorecard; measured: 8.4× cheaper per appeal)
152 Tests (100% PASS)
Phases 7–13 · Complete

Live Solari Driver, MCP Server & Agent Skill

Driver verified against the live Solari API, arc-cua-mcp with five tools, budgeted perception on real list pages, marked screenshots, and a head-to-head vs Solari's MCP.

7/7 vs 6/7 tasks · 3.3× fewer tokens
13
234 Tests (100% PASS)
Production Horizons

Roadmap for Live Deployment

With the engineering harness 100% complete and validated across 234 tests and the Solari driver verified against the live API, the three next milestones execute on live cloud infrastructure.

01

Live Cloud Infrastructure

Browser sessions already run on Solari through a driver verified against the live API. Next: Solari desktops (no accessibility tree exposed yet) and pre-warmed WebArena multi-container environments.

Target: Solari desktops Browsers Live
02

ModernBERT Fine-Tuning

Execute scripts/train_monitors_full.py on NVIDIA CUDA GPUs using empirical trajectory windows captured during Phases 3B and 4A to train learned sequence monitors.

Target: F1 > 0.90 Classifier Pipeline Built
03

Enterprise Daemon Integration

ARC now ships as an MCP server and agent skill for Claude Code and other MCP hosts. Next: a background daemon for continuous QA validation and RPA workflows.

Target: arc-cua-mcp MCP Shipped

Deploy Arc Hybrid CUA Today

Run the complete test suite or execute production scorecards in seconds on your workstation.

# Install editable package
pip install -e ".[mcp]"
# Verify all 234 unit and integration tests
pytest tests/
# Run the mock scorecard (live benchmarks: see README)
python scripts/report_production.py
# Use it from Claude Code
claude mcp add --scope user arc -- arc-cua-mcp
Evidence Provenance

Sourced References

The works below actually underlie this page's claims. Benchmark task definitions follow WebArena and OSWorld; perception and cost methods follow Mind2Web; the monitor design references ModernBERT. All repository-measured figures (2.31 ms avg latency, $0.0015/task, 100% mock success, 94.2%/98.1%/100% monitor rates) are this repository's own scorecard results — not published-paper results.

STATUS: PROJECTED_BASED_ON_MOCK_EXECUTION. External GPT-4o/Sonnet comparison rows are this repo's recorded baselines, not paper-measured values. Full provenance with verification status in docs/REFERENCES.md.
Benchmark task definitions
WebArena — Zhou et al., 2024
A Realistic Web Environment for Building Autonomous Agents · ICLR 2024
arxiv.org/abs/2307.13854
Benchmark task definitions
OSWorld — Xie et al., 2024
Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments · NeurIPS 2024
arxiv.org/abs/2404.07972
Perception / cost method
Mind2Web — Deng et al., 2023
Towards a Generalist Agent for the Web · NeurIPS 2023
arxiv.org/abs/2306.06070
Monitor-model reference
ModernBERT — Warner et al., 2024
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder · 2024 · design reference for the trajectory monitor (CUDA fine-tuning is a roadmap item)
arxiv.org/abs/2412.09535