Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusEmerging
SignalTracked
EvidenceGrade B

OpenHands

59
C -1 vs last quarter

Open-source autonomous coding agent (MIT, 75.2K stars) with enterprise ambitions via All Hands AI. OpenHands Enterprise / Agent Control Plane (announced Mar 30, GA press May 6, 2026) provides centralized workflow orchestration, workflow-scoped least-privilege access, spend tracki

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
75% conf63
Capability
75% conf62
Reliability
79% conf51
Value
50% conf60
Community
79% conf37
Enterprise / Compliance
50% conf55
Autonomy
50% conf75
Integration
50% conf65
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Open-source autonomous coding agent (MIT, 75.2K stars) with enterprise ambitions via All Hands AI. OpenHands Enterprise / Agent Control Plane (announced Mar 30, GA press May 6, 2026) provides centralized workflow orchestration, workflow-scoped least-privilege access, spend tracking, audit trails, and reproducible sandboxes. SWE-bench Verified 68.4% (CodeAct v3, Claude Opus 4.6) holds as an independently corroborated top-tier open-source result (a 77.6 badge on the repo is vendor-reported and not yet corroborated on public leaderboards).

Now confirmed at scale: parallel agent fleets (1000s via RemoteRuntime), sub-agent delegation, headless/SDK (Python + REST). Multiple named enterprise references — US Mobile (80% dev-time reduction), Flextract (87% same-day bug resolution), C3 (data-science scale) — plus AMD partnership.

Model-agnostic via LiteLLM (100+ providers); self-hosted VPC + air-gapped deployment; SAML/SSO + RBAC. GAPS: still no SOC 2/ISO 27001 (path only); historically weak vuln-triage process (2025 'Lethal Trifecta' disclosure went weeks untriaged, mitigations later published; CVE-2026-33718 HIGH command-injection patched in 1.5.0); $23.8M total funding with a ~6-person core team and no Series B/revenue disclosed; Agent Control Plane GA but production-scale deployments still unvalidated by us; OpenCode (~165K stars, ~2.2x OpenHands) dominates open-source coding mindshare, though it occupies a different terminal-native niche.

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

No active caps — no risk flags apply to this tool right now.

Assessment

Status rationale

Tracked is appropriate: strong open-source momentum (75.2K stars, ~490+ contributors, MIT license, 68.4% SWE-bench Verified, parallel-fleet + headless SDK), clear enterprise product direction (Agent Control Plane GA), and now multiple named enterprise references (US Mobile, Flextract, C3). Blockers to Assessed: SOC 2/ISO not yet achieved, Agent Control Plane lacks production deployments validated by us at scale, thin Series A balance sheet / ~6-person team limits runway certainty vs. larger incumbents, and competitive pressure from OpenCode is intensifying.

Evidence is research-based without hands-on testing — depth is thorough (full-research re-baseline, 50+ searches, multi-hop + temporal), lifting the evidence grade to B while the signal level stays Tracked.

Watch for

Movement triggers

Upgrade to Assessed if: (1) SOC 2 Type II certification achieved AND (2) 2+ additional named enterprise customer references documented with production-scale Agent Control Plane deployments. Upgrade to Validated if additionally: hands-on internal testing completed showing reliable complex task performance with no critical open blockers.

Downgrade to Detected if: development stalls (90+ days no releases), key contributors depart, Agent Control Plane revenue traction fails to materialize, new unpatched critical security vulnerability emerges, or OpenCode/Cline/competitor capture causes sustained contributor exodus.

Caution

Risks & limitations

Not yet assessed — this section fills in as ACES research covers the tool.

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.