Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusEmerging
SignalValidated
EvidenceGrade B

Claude Code

56
C -10 vs last quarter

Terminal-native agentic coding at scale with dominant enterprise adoption, now re-baselined on Claude Opus 4.8 (May 28, 2026): SWE-Bench Pro 69.2% category-leading, 4x fewer unflagged code flaws, Dynamic Workflows running hundreds of parallel subagents per session.

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
79% conf59
Capability
79% conf65
Reliability
79% conf40
Value
50% conf85
Community
79% conf0
Enterprise / Compliance
50% conf65
Autonomy
50% conf60
Integration
50% conf75
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Terminal-native agentic coding at scale with dominant enterprise adoption, now re-baselined on Claude Opus 4.8 (May 28, 2026): SWE-Bench Pro 69.2% category-leading, 4x fewer unflagged code flaws, Dynamic Workflows running hundreds of parallel subagents per session. Anthropic financial position exceptional (~$45B annualized revenue, raising at $900B+, Claude Code ARR $2.5B).

Validated maintained: 70% Fortune 100, 8 of Fortune 10, NYSE/Stripe/Ramp/Netflix deployments, overtook ChatGPT in US business AI payments.

Caution

reliability-complaints cap held (autonomy 12) — Opus 4.8 is day-zero with no production-validation window, and the March–May throttling backlash, though actively remediated (5-hour limits doubled May 6, rate limits reset May 15), is not yet a closed 90-day reliability cycle.

Three command-injection CLI CVEs (CVE-2026-35020 CVSS 8.4; CVE-2026-35022 9.9 in CI/CD) remain vendor-dismissed as Informative; CI/CD remote-exploit path is a watch item.

New

Cowork desktop agent (Jan 2026) — runs alongside the CLI for async background tasks. Beta for individual users; GA expected mid-2026.

Recommended

Use cases

  • Terminal-native developers running multi-step refactors from the shell
  • Backend engineers who want a CLI-first companion that respects git semantics
  • Government contractors needing FedRAMP-validated AI tooling
  • Non-coding knowledge work (research summaries, document drafts) via the same CLI
  • Mid-size teams piloting agentic workflows before broader rollout
Score caps

Risk flags

  • Reliability complaints

    trustTemporary

    Widespread reliability complaints (breaks often, unreliable output)

    Caps Autonomy at 60

    Removed when90+ days of improved reliability with community acknowledgment

Assessment

Status rationale

Validated because enterprise adoption criteria fully met: 70% Fortune 100, 8 of Fortune 10, named production deployments (NYSE, Stripe, Ramp, Netflix), comprehensive compliance (SOC 2 Type II, ISO 27001/42001, HIPAA BAA, FedRAMP High + DoD IL4/5 via Bedrock GovCloud, NIST CUI attestation). reliability-complaints cap is temporary and does not trigger demotion. Three command-injection CLI CVEs are vendor-dismissed as Informative with local/CI-precondition exploit paths and remain unpatched; monitored but not status-blocking.

Watch for

Movement triggers

Upgrade autonomy if: Opus 4.8 reliability gains hold through a 90-day production window with no new post-rollout degradation cycle AND community acknowledges improvement AND throttling/usage-limit backlash resolves (real-time weekly-cap visibility shipped). Upgrade compliance if: command-injection CLI CVEs (#35020/#35022) patched or formally mitigated AND FedRAMP on Claude Code CLI ships AND enterprise SLA added.

Downgrade to Assessed if: (a) Opus 4.8 triggers an April-scale community reliability crisis (2,000+ report equivalent, major press), OR (b) named enterprise customers publicly churn citing reliability, OR (c) a new critical unpatched CVE is actively exploited in the wild, OR (d) usage share falls below 50% primary-tool. Downgrade compliance/status if: CLI CVEs exploited in wild; Pentagon designation expands to commercial procurement.

Caution

Risks & limitations

  • Reliability Complaints

    Moderate

    Radar cap: reliability-complaints

Capabilities

Integration surface

Context window

200K tokens standard · 500K Enterprise · 1M beta (gated).

9 access points
  • Terminal CLIGA
  • Web interfaceGA
  • Desktop appGA
  • VS Code extensionGA
  • JetBrains pluginBeta
  • Slack integrationGA
  • GitHub ActionsGA
  • iOS mobileBeta
  • Cowork desktop agentBeta
Proof points

Adoption & benchmarks

80.9%
SWE-bench Verified
Highest in catalog
77.2%
Sonnet 4.5 (fallback)
Same benchmark
59.3%
Terminal-bench
Tooling agility
66.3%
OSWorld
Cross-OS task completion
Enterprise deployments
  • Accenture· 30,000 users
  • Cognizant· 350,000 deployment
  • TELUS· Enterprise rollout
  • Netflix· Engineering teams
  • Spotify· Engineering teams
  • Snowflake· Engineering teams
Government credentials

FedRAMP High, IL4/5, DoD selection, GSA OneGov pricing.

Productivity research

2–6 hours/week saved (Anthropic internal study, n=80). METR RCT (July 2025) showed 19% slowdown for experienced open-source developers in unfamiliar codebases — context-dependent.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.