Agent S3
by Simular
Open-source computer-use agent that crossed the human baseline on OSWorld
Agent S3 is Simular's Apache-2.0 computer-use agent framework: it drives a real desktop through the GUI, reading the screen with a grounding model and planning with a frontier LLM. Its Behavior Best-of-N scaling runs several attempts in parallel and picks the best, reaching 72.6% on OSWorld against a human baseline of about 72%.
Agent S3 is the third generation of Simular's open-source computer-use agent framework, released 15 December 2025 under Apache-2.0 and installed as the `gui-agents` Python package. It operates a real desktop the way a person does — reading the screen through a visual grounding model, then planning and issuing mouse, keyboard and code actions — and runs on Linux, macOS and Windows. The headline architectural change from Agent S2 is subtraction: S3 removes the hierarchical manager-worker structure of its predecessor and adds a native coding agent that can generate and execute code, letting a single agent mix scripted solutions with GUI manipulation instead of clicking through everything. The second change is Behavior Best-of-N (bBoN), the first wide-scaling method for computer-use agents: rather than trusting one rollout, it runs several in parallel, compresses each into generated facts and a behaviour narrative, and has a judge compare the narratives to select the best outcome. The results are the reason the project matters. On OSWorld, a 369-task benchmark of real computer work, a single run scores 62.6% and bBoN across ten runs reaches 72.6%, edging past the human baseline of roughly 72.36% — a benchmark where the best score was near 20% a year earlier. bBoN also lifts WindowsAgentArena from 50.2% to 56.6% and AndroidWorld from 68.1% to 71.6%. The framework is model-agnostic, with OpenAI, Anthropic, Gemini, OpenRouter and vLLM backends supported; the published numbers use GPT-5 for planning and UI-TARS-1.5-7B for grounding. Simular, backed by a $21.5M Series A led by Felicis in December 2025 with NVIDIA's NVentures participating, sells the managed product built on it.
The applied-AI or automation engineering team that has to drive legacy desktop or web software with no usable API, and wants a self-hosted, Apache-2.0 agent they can run against their own model keys rather than a closed vendor agent.
A computer-use agent at the current state of the art — 72.6% on OSWorld, past the human baseline — that you can install with pip, run locally, and point at whichever model backend you already have contracts for.
At a Glance
- Category
- AI Agents & Orchestration
- Pricing
- Free, Subscription, Usage-based, Contact for pricing
- Target Market
- CTOs, ML/AI Engineers, Automation Engineers, Enterprise Developers, AI Researchers
- Deployment
- Open-source, Self-hosted, Cloud-first
- Headquarters
- San Francisco, California, United States
Key Features
- ✓Behavior Best-of-N (bBoN) wide scaling
Runs N rollouts in parallel, narrates each as a behaviour summary, and has a judge pick the best — turning variance into accuracy.
- ✓Native coding agent
Generates and executes code alongside GUI actions, so a task solvable by a script does not have to be clicked through.
- ✓Flat non-hierarchical architecture
Drops Agent S2's manager-worker hierarchy, which simplified the loop and contributed to a roughly 13% accuracy gain.
- ✓Pluggable model backends
Supports OpenAI, Anthropic, Gemini, OpenRouter and vLLM for planning, with a separate grounding model for screen coordinates.
- ✓Cross-platform desktop control
Runs on Linux, macOS and Windows against real applications, not just a browser sandbox or a virtualised web page.
- ✓Apache-2.0 licence and pip install
Permissively licensed and installable as gui-agents, so it can be embedded in commercial products without licence negotiation.
- ✓Benchmark reproducibility
Published OSWorld, WindowsAgentArena and AndroidWorld results with named model backends, so claimed scores can be independently re-run.
Capabilities
Use Cases
- •Driving legacy desktop software
Automate an internal or vendor application that has no API by operating its interface directly, exactly as a human operator would.
- •Cross-application back-office workflows
Chain steps across a browser, a spreadsheet and a desktop client in one run, which Simular reports across insurance, recruiting and travel.
- •QA and regression testing of GUIs
Exercise a real application through its interface to find breakages that scripted selectors miss when the UI shifts underneath them.
- •Computer-use agent research and benchmarking
Use a permissively licensed state-of-the-art baseline to evaluate new grounding models or planners on OSWorld and its siblings.
- •Reliability engineering for agent pipelines
Apply bBoN's rollout-and-judge pattern to raise task success where a single agent run is too inconsistent to ship.
Ideal For
Best For
- ✓Automating legacy desktop or internal applications that expose no API and cannot be scripted any other way
- ✓Research teams benchmarking computer-use agents who need a reproducible, permissively licensed baseline
- ✓Cross-application workflows that span a browser, a spreadsheet and a desktop client in one task
- ✓Teams that want to self-host the agent loop and keep screen data on their own machines rather than sending it to a vendor's cloud
- ✓Back-office operations work in insurance, recruiting and travel, where Simular reports agents have executed millions of steps
Not Ideal For
- ✗Anyone who can solve the problem with an API or RPA connector — a GUI agent is slower, costlier and far less deterministic than an HTTP call, and should be the fallback rather than the first choice
- ✗Latency- or cost-sensitive production paths: the 72.6% headline requires ten parallel rollouts through a frontier model plus a hosted grounding model, so each task multiplies inference spend by roughly an order of magnitude
- ✗Untrusted or multi-tenant environments — the native coding agent executes generated code locally, and the documentation explicitly warns to run it only in trusted environments
- ✗Multi-monitor workstations and non-standard resolutions; the framework requires a single monitor and the recommended grounding model expects 1920x1080 coordinates
- ✗Buyers who need a vendor-supported, compliance-certified product out of the box rather than a framework — that lives behind Simular's custom-priced Enterprise tier
Integrations
Deployment
Market Analysis
Pros
- ✓Best published OSWorld result at time of writing, and the first past the human baseline on a benchmark that sat near 20% a year earlier
- ✓Apache-2.0 with 12.1k GitHub stars — genuinely usable in a commercial product without licence negotiation
- ✓Model-agnostic across OpenAI, Anthropic, Gemini, OpenRouter and vLLM, so it is not a bet on one provider
- ✓Runs on Linux, macOS and Windows against real applications, with published results on WindowsAgentArena and AndroidWorld as well as OSWorld
- ✓Backed by a December 2025 $21.5M Series A led by Felicis with NVIDIA's NVentures, so the open framework has a funded team behind it
Cons
- ✗The headline number is a best-of-ten result. Solo performance is 62.6%, and Simular itself concedes 'high variance' on long-horizon tasks — single runs stay inconsistent even with stronger models
- ✗Benchmark accuracy is not production reliability: even at 72.6%, better than one task in four still fails, which rules out unattended use on anything consequential
- ✗The local coding agent executes arbitrary generated code and the documentation warns to use it only in trusted environments — a real constraint for enterprise deployment
- ✗Rigid environment requirements: single monitor only, and the recommended grounding model expects a specific 1920x1080 coordinate resolution
- ✗Effectively no independent practitioner discussion — Hacker News threads about Agent S and Simular top out at 4 points with one or two comments, and there is no G2, Capterra or TrustRadius presence to corroborate the vendor's claims
- ✗Compliance claims (SOC 2, HIPAA, SSO, RBAC) appear only as Enterprise-tier bullets on the pricing page with no public attestation
Pricing
Agent S3 (open source)
$0
- ✓Apache-2.0 licence
- ✓pip install gui-agents
- ✓Linux, macOS and Windows
- ✓Bring your own model API keys
- ✓Full bBoN implementation
Simular Plus
From $20/mo per computer
- ✓10,000 reloadable credits
- ✓1 developer agent
- ✓macOS and Windows
- ✓Workflow editor in developer mode
- ✓Community workflow sharing and support
Simular Pro
From $500/mo per computer
- ✓Unlimited credits
- ✓1 free deployment agent
- ✓Virtual machine support
- ✓Private and team workflow sharing
- ✓Zero data retention from LLM providers
- ✓Priority engineering support
Enterprise
Contact for pricing
- ✓Unlimited agents across platforms and private cloud
- ✓Managed scaling
- ✓SOC 2, HIPAA, SSO and RBAC
- ✓Team and audit log sharing
- ✓Dedicated account and engineering support
The framework itself costs nothing under Apache-2.0, but the model bill is the real price: it needs a frontier planner — the published numbers use GPT-5 — plus a hosted grounding model such as UI-TARS-1.5-7B on an inference endpoint, and bBoN multiplies that by the number of rollouts, so the 72.6% configuration costs roughly ten times a single run. Simular's managed product meters per computer rather than per seat: $20/month Plus with 10,000 reloadable credits, $500/month Pro with unlimited credits and zero LLM-provider data retention, and custom-priced Enterprise. SOC 2, HIPAA, SSO and RBAC are Enterprise-only vendor claims; no public attestation was found.
Security & Compliance
Connect
Sources
This page was written from 6 sources, 5 on domains other than github.com.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Sema4.ai
Enterprise AI agents for knowledge work, running natively in your Snowflake or cloud account
Sapiom
Agent infrastructure that routes, runs and meters AI agents in production
Itential FlowAI
Governed AI agents for network and infrastructure operations, with deterministic execution and full audit trails
OpenAI Presence
Deploy production-grade AI voice and chat agents with enterprise policies, guardrails and evals