ColdStart Public Review: Zero-Shot Generalization Benchmark Harness for Computer-Use Agents (CUA) on Solari MicroVMs.
โšก Solari MicroVM Snapshot Forks ๐Ÿงช 110 Vitest Unit Tests PASS ๐Ÿ”’ Fail-Closed DB Verification (C1โ€“C7) โš–๏ธ MIT Open Source

Benchmarking Zero-Shot Generalization on Computer-Use Agents.

Standard CUA evaluations test repeatability on static forms the agent was already tuned on. ColdStart turns Solari Firecracker microVM snapshots into an Unseen Environment Factoryโ€”generating procedurally mutated web applications to measure out-of-distribution robustness cold.

LIVE SOLARI CLOUD MICROVM RUN
๐Ÿ‘๏ธ Pure Vision-First CUA โ€ข Pixels In, Coordinates Out
๐ŸŽฏ Dynamic Coordinates: Clicks (x, y) on Real Solari Browser
โœ“ Fail-Closed SQLite Check: Row #1 POSTED ($38,880.00)
What you're seeing: Agent completing invoice in an unseen app cold via Solari microVM (16 steps).
Seeded App Variants
14 Variants
5 Mutation Axes โ€ข Click to view
Surface Invariance
100% (2/2)
P5 Dark Serif Skin โ€ข Click to view
Procedural Breakdown
0.00% (0/2)
P2 Two-Step Wizard โ€ข Click to view
Resource Hygiene
0 Leaks
0 Zombie Sandboxes โ€ข Click to view

Visual Replay Engine

16-Step Execution Trace Flow

Interactive breakdown of the deterministic CUA execution loop.

Run: r_mtjqchve_s17
Step: 01 / 16
Solari CUA Live Step Screenshot
CLICK (x: 349, y: 236)
CUA Executed Action
๐ŸŽฏ CLICK at (349, 236)
Target Element
Due Date Input Field
Vision Model Rationale Zero-Shot Reasoning

“Focus the Due Date field to enter the requested due date.”

SQLite Ground-Truth Verifier
Step in progress... Verifier awaits final state.
1.5s per step โ€ข Click any number below to scrub
Steps:
Steps 01 โ€“ 04
Form Discovery
Agent scans viewport, locates client input field, types "ACMECORP", and fills invoice reference ID.
Steps 05 โ€“ 09
Line Item Entry
Enters description, units (480), and unit price ($75.00). Interacts with dynamic row controls.
Steps 10 โ€“ 13
Calculations & Tax
Verifies computed subtotal ($36,000.00), applies 8% tax ($2,880.00), and validates gross balance ($38,880.00).
Steps 14 โ€“ 16
Commit & DB Verify
Clicks Submit button, confirms modal dialogue, and commits SQLite database record inside microVM.

Empirical Scorecard

Causal Axis-Isolated Evaluation

Holding the underlying task constant (ACMECORP) while systematically mutating isolated axes.

Causal Generalization Success Rate

Ground-Truth SQLite Verified
Evaluated on Solari microVM instances n=2 isolated trials per point
๐Ÿ”ฌ Key Research Discovery

Surface Invariance vs Procedural Breakdown

1. Surface Mutations Don't Trick Vision Models: Complete visual restyling (Dark Serif CSS Skin: 100% 2/2) and heavy semantic synonym drift (100% in 16 steps) pose zero difficulty. 2. Procedural Shifts Cause Total Failure: Converting a flat form into a two-step wizard flow (P2: 0.00% 0/2) causes the agent to step-cap at 40 steps without submitting.

Measured Observable Compute:
471.6s Sandbox + 338.6s Browser (0.225 billable hrs)
๐ŸŽฎ Interactive Perturbation Simulator

Test How UI Changes Affect Agent Reliability

Toggle mutation dials to simulate agent generalization performance in real time.

P1 Semantic Synonym Relabeling

e.g. "Customer" โ†’ "Client", "Tax" โ†’ "VAT", "Total" โ†’ "Gross"

P5 Visual CSS Theme Reskin

Dark emerald theme, serif typography, altered visual contrast

P3 Field Order & Input Permutation

Shuffles grid layout, swaps input positions, increases field density

P2 (The Breaker) Two-Step Wizard Flow Shift

Splits form into Step 1 (Customer) + Step 2 (Line Items) requiring "Next"

Simulated Generalization Reliability
100% PASSING (2/2)
Empirical Behavioral Diagnosis:

Baseline: Pure vision-first agent completes the standard invoice in 16 steps. All 7 programmatic database checks pass with zero errors.

Ground-Truth Engine: Direct SQLite (C1โ€“C7)

Audited Live Execution Trace Runs

Fail-Closed Verification Pass
Run ID Perturbation Axis Result Steps Wall Time DB Verifier (C1โ€“C7) Replay Session
r_mtjsm13e_t0 P5_theme:3 (Dark Serif) PASS 16 steps 145.4s 7/7 Green ($38,880.00) Captured S3
r_mtjsnjob_t1 P5_theme:3 (Dark Serif) PASS 16 steps 151.7s 7/7 Green ($38,880.00) Captured S3
r_mtjqchve_s17 P1_relabel:4 (Synonym Drift) PASS 16 steps 139.8s 7/7 Green ($38,880.00) Captured S3
r_mtjsp83o_o0 P3_field_order:4 PASS 17 steps 158.2s 7/7 Green ($38,880.00) Captured S3
r_mtjsqyft_w0 P2_structure:3 (Two-Step Wizard) FAIL 40 steps (cap) 362.4s 0/7 (Never submitted) Timing null

Scientific Transparency & Disclosures

1. Sample Size: Isolated runs were evaluated at $n=2$ per point to establish causal bounds within compute quotas. Additional trials can be scheduled via CLI.
2. Infrastructure Resilience: 2 transient infra-aborts (screenshot channel socket drop & sandbox control reset) were recorded transparently rather than discarded.
3. Grounding Resolution: Free general chat models exhibited ~30โ€“60px coordinate drift; our harness-side hybrid grounding snapped clicks to interactive bounding boxes to achieve deterministic execution.

๐Ÿ›ก๏ธ Plain English & Human Safety Perspective

"Is AI coming for our jobs?"
Why ColdStart is actually the safety net humans need.

Hearing about an AI that can move mice, click buttons, and fill out forms can sound intimidating. But here is the plain truth: today's AI is like an over-eager intern who panics the moment a button moves three inches. ColdStart is the safety crash test that keeps everyone safe.

๐Ÿš—

The Crash-Test Dummy for Software

Before a car with autopilot is allowed on public roads, it undergoes grueling crash tests in rain, fog, and unexpected detours. ColdStart does the exact same thing for AI: it deliberately randomizes buttons, labels, and forms in safe virtual sandboxes to prove where the AI fails before anyone lets it touch real company data.

Safety First โ€ข Zero Production Risk
๐Ÿค

Augmenting Humans, Not Replacing

Our empirical tests proved that AI drops to 0% success the moment a single-page form is split into a two-step wizard. AI cannot replace human judgment, common sense, or real responsibility. What it can do is take away soul-crushing, repetitive data entryโ€”freeing human professionals to focus on high-value strategy and relationships.

Human-in-the-Loop โ€ข Eliminate Drudgery
๐Ÿ”’

Never Trust an AI's Word

When an AI says "I finished the task!", normal software takes its word for it. ColdStart refuses to believe the AI's story. It independently checks the company database behind the scenes with strict mathematical proof. If a single number or invoice total doesn't match 100%, it automatically rejects the action.

Fail-Closed โ€ข Independent Verification
๐Ÿ’ก

Summary for Leadership & Teams: ColdStart doesn't unleash reckless automation. It is the rigorous inspector ensuring your digital tools are safe, accurate, and completely under human control before deployment.

See Where AI Breaks โ†“
๐Ÿ’ผ Real-World Value & The Build Story โ€ข 4h Build Session โ€ข 48-Hour Lifecycle

Solving Real Problems & The Build Story

“We donโ€™t care how you ship, we care that you can ship something great, and if you can ship it faster with AI, even better.”

โšก Why You Need Solari & ColdStart

โœ• The Problem: Fragile Automation Scripts

Old-school scrapers and RPA bots crash the moment a developer tweaks a button class, changes a layout, or moves a field into a popup.

โœ“ How Solari Fixes It

Solari gives the AI visual eyes. It looks at screenshots and clicks real coordinates on the screen just like a human โ€” immune to website code changes.

โœ• The Problem: The "Memorized Demo" Trap

AI agents often look 100% reliable on the exact demo form they were built on, but get confused and fail in the real world when buttons or fields move around.

โœ“ How ColdStart Fixes It

ColdStart automatically scrambles labels, themes, and layouts so you can test your AI cold before putting it in front of real users.

โœ• The Problem: Slow, Heavy Virtual Machines

Setting up full virtual machines takes minutes to boot, costs a lot to run 24/7, and leaves leftover credentials and security risks behind.

โœ“ How Solari Fixes It

Solari boots lightweight, throwaway browser sandboxes in ~10 seconds (measured). Snapshot fast-forks were best-effort in our runs and mostly 409'd, so direct provisioning is the working path. You run your task, close the session, and everything is wiped clean with 0 leftover leaks.

๐Ÿ› ๏ธ 5 Common Real-World Use Cases

Finance & Operations ๐Ÿงพ
Automated Invoice & Data Entry

The Problem: People spend hours copying numbers from PDFs and emails into accounting software (QuickBooks, Xero, ERPs). It costs $2โ€“$4 per invoice in manual time and causes frequent typos.

The Fix: The AI opens the invoice, navigates the web portal, types in the fields, and submits. ColdStart ensures it doesn't crash when the portal updates its layout.
Personal Utility ๐Ÿ“ธ
Smart Photo & Storage Cleanup

The Problem: Google Photos and iCloud get clogged with redundant screenshots, receipts, and duplicate bursts, forcing users to pay for extra storage tiers.

The Fix: An isolated Solari browser agent scans your photo albums, spots blurry bursts and redundant screenshots, and deletes the clutter automatically.
Lead Gen & Research ๐Ÿ•ท๏ธ
Reliable Web Scraping & Lead Enrichment

The Problem: Modern websites block traditional web scrapers or change their HTML tags every few weeks, breaking data feeds and sales lead pipelines.

The Fix: The AI browses like a human with mouse movements and visual reading, collecting verified lead data directly into your spreadsheets.
Software Teams ๐Ÿงช
Autonomous Web App Testing (QA)

The Problem: Developers spend days writing and fixing brittle test scripts every time they redesign a page or add customer-specific themes.

The Fix: ColdStart generates mutated versions of your app (new themes, rearranged inputs, synonym labels) to stress-test your user flows automatically.
Software Teams & Design Ops ๐ŸŽจ
Automated PR Gatekeeping & "Slop" Filtering

The Problem: Developers are drowning in AI-generated pull requests. The bottleneck is no longer writing code; it's reviewing it for "AI Slop" (bad contrast, generic layouts, poor spacing). Running heavy CUAs to check UI aesthetics is cost-prohibitive.

The Fix: The ColdStart Multi-Model Router spins up a Solari microVM in 10s. The lightweight VLM layer checks the PR for accessibility and brand compliance for pennies. If it passes, the heavy CUA verifies the structural flow. Bad PRs are auto-blocked before a human ever reviews them.
Measured ROI & Illustrative Economics

Dual ROI: Time Savings + Economic Advantage

โฑ๏ธ ~6.5-9x vs. assumed baseline ๐Ÿ’ฐ $ cost: not measurable (illustrative)
Metric 1: Time Savings (measured) ~49-69s per task
Manual entry (assumed baseline): 450s (7.5 min) / doc
Solari + Vision AI (measured): ~49-69s / doc

The Impact: The manual baseline is a stated assumption, not a sourced benchmark. Measured agent wall time (isolated set: 471.6s sandbox across 6 runs โ‰ˆ 78.6s/run โ†’ ~21.8h per 1,000; mixed Step 06 set: 616.7s across 5 runs โ‰ˆ 123.3s/run โ†’ ~34.3h; the single showcase run r_mtjqchve_s17 ran 69.4s sandbox / 49.2s browser but the scored sets include step_cap and retried runs), against an assumed 125 hours of human entry.

๐Ÿ“Š Source: artifacts/runs/r_mtjqchve_s17/run.json and artifacts/scorecard.json (wall seconds). Manual baseline is an assumption, not a citation.
Metric 2: Economics (illustrative modeling) not measurable
Human labor (assumed $25/hr × 450s): $3.13 / task (assumed)
Solari MicroVM + Model Compute: $ not measurable*

*Why not measurable: the Solari SDK exposes no credit balance or $/hour rate (credits: null in artifacts/scorecard.json), so no $ cost can be computed. The observable envelope is ~69s sandbox + ~49s browser + 16 LLM calls + ~28.6k token-in / ~0.7k token-out per run. Any revenue/margin figure would be arithmetic on unmeasured inputs.

๐Ÿ“Š Source: src/scorecard/cost.ts, which states there is "no defensible $ conversion without a published rate." The earlier BLS citation was not traceable and has been removed.
โš”๏ธ Competitive Analysis: Why Solari + ColdStart Wins
Evaluation Dimension Legacy RPA (UiPath) Scripted (Playwright) DOM-AI (Browserbase) Solari + ColdStart
Action Space OS Selectors DOM CSS Selectors DOM Tree Parsing Visual Pixels & Coordinates
UI Change Resilience 0% (Breaks on refactor) 0% (Fails on ID/CSS change) Partial (Fails on shadow DOM) Measured directional results (P5 2/2 isolated; P1 1/1 mixed)
Sandbox Boot Time 60โ€“180s (Full Windows VM) ~5โ€“10s (Local Node) ~15โ€“30s (Container) ~10s MicroVM Boot (measured)
Out-of-Distribution Testing None (Warm app only) None (Hardcoded) None (Live site only) ColdStart 5-Axis Engine
Cost Per Execution $1.20โ€“$2.00 + $15k license $0.10โ€“$0.30 (Dev time) $0.10โ€“$0.25 (Token bloat) not measured (no $ rate exposed)
Verification Integrity Self-reported status DOM assertions LLM self-reported 'done' Fail-Closed SQLite Direct Channel
The AI-Accelerated Build Story

From Architecture to Verified Production: The Build Story

Reconciling the Two Timeframes: The initial harness, verifier, and CI landed in one ~4-hour afternoon session on Sep 2 (git log: first commit 14:24, showcase deployed the same afternoon). Docs, showcase media, and the cost router followed that evening through ~03:00 Sep 3. Research, auditing, and architecture took the rest of the 48-hour project lifecycle.

The coding and deployment took one ~4-hour session on September 2nd; docs, media, and polish continued that evening through ~03:00 September 3rd. Research and architecture took the rest of the 48-hour weekend.

๐Ÿ’ก AI Tooling Acceleration: By leveraging an agentic multi-tool stack (Google Antigravity, Claude Code, and Pi Coding Agent routed through OpenCode Go) for rapid TypeScript scaffolding, Vitest generation, and boilerplate orchestration, I compressed weeks of typical benchmarking harness development into one focused afternoon build session.

14:00 โ€“ 14:30 (Hour 0)
Harness Scaffolding & Mutation Engine
Translating Spec into Code: With the architecture locked during the weekend research phase, we leveraged Google Antigravity and Claude Code (two AI coding agents in parallel) to scaffold the 5-axis procedural mutation engine (P1โ€“P5) into TypeScript, enforcing the deterministic PRNG invariant (same seed -> same variant) so every benchmark test is fresh and un-memorized.
14:30 โ€“ 15:00 (Hour 1)
Solari Fast Sandboxes
Connecting Solari MicroVMs: Wired the test harness to Solari's microVM snapshot forks with SDK bindings and strict resource cleanup orchestration generated via Pi Coding Agent through OpenCode Go. Measured boot was ~10 seconds (create queue plus serve); the snapshot fast-fork path returned 409 on 7 of 8 attempts in our runs, so direct provisioning is the working path, with automated cleanup ensuring zero leftover sandboxes after every run.
15:00 โ€“ 15:45 (Hour 2)
The Struggle & Breakthrough
Fixing the "Click-Lock" Bug: In our first live runs, general chat models got stuck โ€” clicking the exact same textbox 24 times without ever typing! Pair-programming with Antigravity, we quickly diagnosed the visual grounding breakdown and engineered smart coordinate snapping to lock the AI's aim onto interactive inputs, achieving a smooth 16-step completion.
15:45 โ€“ 17:00 (Hour 3+)
Verification & Shipped Live
Direct DB Verification & Deployment: Instead of taking the AI's word for it, our verifier checked the real database behind the scenes to confirm every line item was saved. Running benchmark evals against GPT 5.6 Luna via OpenCode Go and generating 110 passing Vitest unit tests via Claude Code, we achieved green CI and deployed the live showcase.
Curious about the technical specs or raw logs?

System Architecture

How ColdStart Evaluates Agents Cold

From seeded ephemeral microVMs to fail-closed SQLite verification.

01

Snapshot-as-Factory

Pre-bakes dependencies into a base microVM snapshot. At evaluation time, it boots new seeded variant environments in ~10 seconds (measured; snapshot fast-forks were best-effort in our runs and mostly 409'd, so direct provisioning is the working path), ensuring clean sandbox isolation without container pollution.

VariantFactory.ts โ€ข Seeded PRNG
02

Vision-First CUA Loop

The agent operates purely on viewport screenshots (1280x800) with zero DOM introspection or selector access. Harness-side hybrid grounding snaps coordinate predictions to active element boundaries.

AgentLoop.ts โ€ข Hybrid Grounding
03

Fail-Closed Verification

Completely ignores agent self-reporting. Recomputes expected mathematical totals from ground-truth task seeds and executes direct SQLite queries inside the microVM to verify SHA256 integrity checks (C1โ€“C7).

Verifier.ts โ€ข Direct DB Channel
๐Ÿง 

๐Ÿง  Phase 2 Prototype: Slop-Catcher & Model Configuration (MOCK ONLY)

Prototype perception/action configuration via src/config/model-router.ts
โœ“ Prototype / evidence-backed core โšก Compute savings: unmeasured

The core benchmark separates Action from Perception at a configuration boundary. The VLM path is an offline prototype; live reliability and compute economics remain unmeasured. Configuration lives in src/config/model-router.ts.

Layer 1: Perception Only
MOCK
The "Slop-Catcher"

The Problem: Developers push code that works functionally but fails experientially (bad spacing, poor contrast, generic AI slop).

The Solution: We don't need a CUA to check alignment. We use a high-fidelity Vision-Language Model (VLM) alongside deterministic CSS audits.

Stack: Gemini 1.5 Flash / GPT-4o via getModelConfig("PERCEPTION")

Workflow: Takes a screenshot of the sandboxed app. The VLM compares it against Design System references, calculates contrast & spacing variance, and flags aesthetic deviations.

๐Ÿ’ฐ Cost impact: unmeasured (MOCK prototype)
Layer 2: Action & Reasoning
FUTURE
Adversarial Red-Team

The Problem: Frontier models are brittle against dark patterns and structural shifts (like our P2 Two-Step Wizard).

The Solution: An asymmetrical Agent-vs-Agent architecture targeting behavioral regressions.

Stack: Claude 3.5 Sonnet (Attacker) vs. UI-TARS / GPT-5.6 Luna (Defender) via getModelConfig("ACTION")

Workflow: The Attacker model dynamically generates adversarial UI traps (e.g., honeypot modals, deceptive flows). The Defender CUA must navigate them.

๐ŸŽฏ Targeted Structural Evaluation
Layer 3: Production Gate
FUTURE
The CI/CD Triage Gate

The Workflow: When a PR is pushed, a lightweight text model triages the diff. If UI components are touched, Solari boots the microVM in ~10s.

Layer 1 (VLM) checks for "slop". If it passes, Layer 2 (CUA) runs the structural generalization tests.

The Result: Proposed workflow; reliability and compute savings require a live implementation and measurement.

โšก Proposed workflow โ€” unmeasured

Slop-Catcher Demo: Clean vs. AI Slop

Below is the live diagnostic output of the demo runner (npm run demo:all). It evaluates the exact same landing page codebase in two states: a human-grade clean design, and an injected 'AI Slop' state.

artifacts/slop-catcher-replay.gif
Action Replay (7 Steps) โ†“ MP4
ColdStart Slop-Catcher Action Replay
artifacts/combined-demo-report.html
Slop-Catcher Diagnostic Report
๐Ÿ›ก๏ธ

๐Ÿ›ก๏ธ Phase 3: QA dogfooding evidence

Live functional QA via src/qa-framework/ โ€” npx coldstart tunnel
โœ“ Prototype / evidence-backed core โšก 30 Live Functional Checks

The Slop-Catcher asks โ€œdoes it look right?โ€ โ€” the QA Framework asks โ€œdoes it work?โ€ Real cloud Chromium sessions drive auth, chat, history, profiles, and workers end-to-end (batches B1โ€“B5), with every claim backed by direct SQLite reads. If it isnโ€™t in the database, it didnโ€™t happen.

Guards: Anti-Flake
MOCK
Zero-Pixel Trap & Copy Drift

The Problem: Tests pass on invisible 0ร—0 elements, and break every time a designer renames a button.

The Solution: expectInteractive() demands visibility + positive dimensions; fuzzyRoleLocator() absorbs copy drift via normalized matching.

๐ŸŽฏ False passes blocked at the assertion layer
Lifecycle: Zero Zombies
FUTURE
Session Guard + Tunnel + SmartReset

The Problem: Cloud browser sessions leak, localhost isnโ€™t reachable from microVMs, and dirty state makes reruns flaky.

The Solution: withSessionGuard() tears down in finally; TunnelDaemon exposes localhost via Cloudflare; SmartReset reseeds fixtures idempotently.

๐Ÿ›ก๏ธ 0 leaked microVMs โ€” same seed, same state
Verdict: Trust Nothing
TEMPLATE
DB Diff + Archiver + Heuristics

The Workflow: DatabaseDiffEngine verifies every UI claim against SQLite (fail-closed D1โ€“D3 style); ArtifactArchiver preserves evidence; HeuristicEngine formats supplied findings into enhancement notes (template only).

The Result: Heuristic output is a report template, not an autonomous page inspector.

โšก Run it: npx coldstart tunnel 4310

QA Framework Demo: Live Functional Batches

Below is the action replay of the QA pipeline (npm run build:qa-gif). Seven stages: tunnel bootstrap โ†’ fixture seeding โ†’ resilient selectors โ†’ zero-pixel trap โ†’ live E2E run โ†’ DB-diff verification โ†’ heuristic verdict.

artifacts/qa-framework-replay.gif
Action Replay (7 Steps โ€ข 1280ร—720) โ†“ MP4
ColdStart QA Framework Action Replay
Try the QA Framework locally
npx coldstart tunnel 4310                      # expose localhost to Solari cloud
TARGET_URL=<url> npx tsx scripts/test-nakama-b1.ts   # run Batch B1 (auth & guards)
npx coldstart heuristics --report qa-evidence/solari-b1/report.json

Procedural Environment Factory

The 5 Out-of-Distribution Perturbation Axes

Seeded PRNG mutations generated on the fly inside fresh Solari microVMs.

Baseline Reference Seed: s0
Baseline Variant UI

Standard Invoice Form

Clean single-page form with canonical field naming (Customer, Description, Quantity, Price, Tax, Subtotal, Gross).

P1: Heavy Synonym Drift Seed: s17
P1 Relabel Variant UI

Semantic Relabeling

Mutates field names: Customer โ†’ Client, Qty โ†’ Units, Tax โ†’ VAT, Total โ†’ Gross, Memo โ†’ Notes.

P2: Procedural Flow Shift Seed: s7
P2 Structure Variant UI

Two-Step Wizard

Transforms flat form into a multi-step wizard requiring explicit "Next", state validation, and "Confirm".

P3: Density & Spatial Order Seed: s3
P3 Field Order Variant UI

Input Permutation

Reorders line items and changes grid arrangement, breaking static coordinate assumptions.

P4: Nested Navigation Seed: s9
P4 Navigation Variant UI

Sidebar Menu Routing

Replaces direct path routing with a dynamic sidebar navigation menu requiring multi-step clicking.

P5: CSS Theme Skin Seed: s21
P5 Theme Variant UI

Dark Serif Aesthetic

Complete visual restyling with dark emerald palette, serif typography, and altered visual contrast.

Build Story & Origin

Why ColdStart, How It Was Built

The competitive discovery, the 48-hour execution log, and the strategic reasoning behind the build.

The Validated Gap

I scanned the other challenge submissions before writing a single line of code:

โŒ Crowded Clusters
Reliability / Verification / Audit 6+
Stealth / Captcha / Proxies 7
Voice Agents 2
Vertical Workflows 2
โœ… Open Gap โ€” Pinetree's Crown Jewel
0
Applicants testing
Zero-Shot Generalization
โ€” The one thing Pinetree actually claims
Pinetree's reported Hallucinate Westworld result (unverified here; add the official link before citing numbers):
Zero-shot on fully-unseen environment, zero prior exposure

Proposal Iteration

REJECTED
v1: Witness

Verification / Reliability Focus

โŸณ

Audited & Pivoted

Found the uncrowded gap

SHIPPED
v2: ColdStart

Zero-Shot Generalization Harness

โฑ๏ธ Fast-Shipping with AI

The 48-Hour Build Timeline

The initial harness, verifier, and CI landed in one ~4-hour afternoon session on Sep 2 (git log: first commit 14:24, showcase deployed the same afternoon). Docs, showcase media, and the cost router followed that evening through ~03:00 Sep 3. Research, auditing, and architecture took the rest of the 48-hour project lifecycle.

From Harry Chow's challenge tweet on Aug 31 to an audited, zero-leak zero-shot benchmark harness on Sep 2: research and architecture took the weekend; the first commit landed Sep 2 at 14:24 and the harness, verifier, and CI shipped in one ~4-hour afternoon session, with docs and media following that evening

โšก AI Tooling Acceleration: By leveraging an agentic multi-tool stack (Google Antigravity, Claude Code, and Pi Coding Agent via OpenCode Go) for rapid TypeScript scaffolding, Vitest generation, and boilerplate orchestration, I compressed weeks of typical benchmarking harness development into one focused afternoon build session.
00
Aug 31 โ€ข Evening
The Spark & Strategic Audit

Harry Chow tweets the $300K challenge: "No resume, no grades. Fork Solari, build a real use case, ship fast with AI."

The Competitive Audit: Scanned the public challenge forks and pull requests in a manual audit (not archived in this repo). Several applicants had already submitted generic "verification CI" entries. Self-rejected v1 ("Witness") to avoid entering a crowded, copycat cluster.

The Discovery: Pinetree's whole identity is zero-shot generalization on unseen environments. Exactly 0 of the other applicants touched it. That was the uncrowded white space.

01
Sep 01 โ€ข Day 1
Research & Architecture (Steps 00 โ€“ 03)

Env Validation & Design Lock (Steps 00-01): Verified Solari API in live environment, proven 0 resource leaks. Locked design: Create-Invoice app with 5 perturbation axes.

Procedural Variant Factory (Step 02): Seeded PRNG matrix generating 14 deterministic web app variants. 15/15 unit tests pass.

Solari Orchestration (Step 03): Wired microVM fork orchestration. Measured boot: ~10.4s per fork (create queue + serve); the snapshot path 409'd on 7 of 8 attempts, so direct provisioning is the working path. Zero container residue. Zero microVMs leaked.

02
Sep 02 โ€ข Build Session
Core Loop & Verifier (Steps 04 โ€“ 05)

Vision-First Agent Loop (Step 04): Built pure vision-first execution (1280x800 screenshots, coordinate actions). Prompted Antigravity and Claude Code to engineer harness-side hybrid grounding and smart coordinate snapping, rapidly diagnosing and resolving the visual "click-lock" bug. Demonstrated 3/3 repeatability on baseline.

Fail-Closed Verifier (Step 05): Accelerated by Claude Code and Pi Coding Agent (via OpenCode Go) to scaffold an independent auditor checking the sandbox SQLite database directly (C1โ€“C7 integrity checks), completely ignoring what the agent claims.

03
Sep 02 โ€ข Build Session
Causal Evaluation (Steps 06 & 06b)

The Breakthrough Finding: Ran causal axis-isolated trials. Discovered the sharp contrast between surface invariance (100% on P1 label synonyms and P5 theme reskins) and procedural collapse (0% on P2 two-step wizard).

Replay Session Wiring: Captured real presigned S3 session replay URLs from Solari driver for every live run.

04
Sep 02 โ€ข Build Session
Packaging & Live Deployment (Steps 07 โ€“ 08)

Packaging & Verification (Step 07): Generated showcase gif (16 keyframes), cleaned secrets, and compiled full test harness.

Final Orchestrator Sign-off (Step 08): PASS. Automated 110 passing Vitest unit tests with zero TypeScript errors and zero leaked microVMs using Claude Code and OpenCode Go (evaluating against GPT 5.6 Luna), deploying the live showcase the same afternoon.

The Ask

"You don't want my resume โ€” you want to know if I can close the gap you care about.

I built ColdStart to measure zero-shot generalization. The Slop-Catcher is an offline mock prototype and future research direction, not a shipped product.

The ColdStart harness uses reproducible variants, harness-side grounding, and fail-closed verification; cost is reported in observable seconds, calls, and tokens because no dollar rate is exposed.

The code runs. The tests pass. The prototype layers remain explicitly unverified. I'd ship it, learn your stack fast, and contribute to Pinetree Agent's reliability story from day one."

About the Author

Ihsan Wanda

Analytics Engineer building at the intersection of data, automation, and real-world workflows.

Ihsan Wanda

Ihsan Wanda

Analytics Engineer

@ SawitPRO

๐Ÿ“ Indonesia (Open to remote/global/SEA)

Future Roadmap

What I'd Build Next

If hired, here's the expansion path from proof-of-concept to production benchmark.

๐Ÿ“

Expand Task Templates

Currently: Create-Invoice only. Add:

  • File Ticket โ€” multi-field form with dropdowns
  • Update Address โ€” navigate, find, update, verify
  • Data Reconciliation โ€” compare, identify, correct
๐Ÿ–ฅ๏ธ

Desktop Variants

Test native GUI applications:

  • Native apps (calc.exe, notepad, file explorer)
  • Cross-platform (Windows, macOS, Linux)
  • Desktop perturbations (DPI, theme, scaling)

Compute Economics & Cost Projections

โšก Router prototype: savings unmeasured
Pipeline Tier / Expansion Evaluator Stack Est. Cost / Run Economic Impact
Layer 1: "Slop-Catcher" (Perception) Gemini 1.5 Flash / GPT-4o unmeasured unmeasured vs. CUA loop
Layer 2: Adversarial Red-Team (Action) Claude 3.5 vs. UI-TARS / Luna ~ $0.15 Targeted CUA execution
Layer 3: Full Continuous Benchmark 4 models ร— 5 axes (Nightly) ~ $4.50 Release qualification gate
Enterprise Real-World Sampling Salesforce / SAP MicroVMs ~10h/mo Production robustness audit

Developer Quickstart

Reproduce Locally in 3 Steps

Full test suite passes with zero external service dependencies.

1. Clone & Install Node >= 18
git clone https://github.com/itw-code/solari-cookbook.git
cd solari-cookbook
npm install
2. Run Full Unit Test Suite (110 Tests PASS) Offline Mocks
npm test
# Output: 110 passed across 11 test files (agent-loop, axes, demo-site, design-qa-orchestrator, model-router, prng, render-demo-report, scan-url, slop-catcher, slop-scoring, verifier)
3. Execute Live Evaluation on Solari API Key Required
export SOLARI_API_KEY="your-solari-api-key"
export OPENAI_API_KEY="your-cua-model-api-key"

# Run isolated causal scorecard evaluation
npm run eval:isolated