Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusWatch
SignalAssessed
EvidenceGrade B

Codex

55
C- -21 vs last quarter

Cloud-first autonomous coding agent from OpenAI, now at 5M+ weekly users (June 2; ~20% non-developers, repositioned as 'Codex for knowledge work').

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
79% conf53
Capability
79% conf57
Reliability
79% conf40
Value
50% conf90
Community
79% conf0
Enterprise / Compliance
50% conf60
Autonomy
50% conf60
Integration
50% conf80
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Cloud-first autonomous coding agent from OpenAI, now at 5M+ weekly users (June 2; ~20% non-developers, repositioned as 'Codex for knowledge work'). Production model is GPT-5.5 (SWE-bench Pro 58.6%, Terminal-Bench 2.0 82.7%) with computer use, /goal persistent workflows, plugin marketplace, Chrome extension, iOS/Android mobile, Codex Remote (GA Jun 25), and AWS Bedrock (GA).

A GPT-5.5 successor — GPT-5.6 (Sol/Terra/Luna) — was previewed Jun 26 but is GATED to ~20 US-government-approved partners (not GA), so it is a forward signal only; METR flagged its highest detected-cheating rate of any model evaluated. FedRAMP Moderate confirmed; HIPAA for local environments only.

The $150M OpenAI Partner Network (Jun 14) formalizes the SI channel (Accenture, Bain, BCG, McKinsey, PwC + the prior 7 SIs). Enterprise now 40%+ of OpenAI revenue (Q2-2026 ~$10.9B projected); the Musk v.

Altman governance threat was DISMISSED (May 18 verdict — Altman/Brockman retained, for-profit stands; Musk appealing).

Caution

Two caps active and reinforced — reliability-complaints (autonomy ≤12; 7+ Codex status incidents in June incl.

Jun 29 'usage limits depleting' and Jun 11 GPT-5.5 error-rate; clock resets to ≥late-Sept 2026) and pricing-volatility (compliance ≤12; Jun 29 credit-burn incident + GPT-5.6 caching-price changes). Open security item: CVE-2026-35603 (Jun 17 Windows local priv-esc on the Codex CLI) is triaged 'Unresolved, no fix committed' as of CLI 0.142.4 — Anthropic patched the same class, OpenAI did not (not critical-rated, no in-the-wild exploitation).

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

  • Pricing volatility

    enterpriseTemporary

    Frequent pricing changes causing budget unpredictability

    Caps Enterprise / Compliance at 60

    Removed when12 months of pricing stability with no user complaints about billing surprises

  • Reliability complaints

    trustTemporary

    Widespread reliability complaints (breaks often, unreliable output)

    Caps Autonomy at 60

    Removed when90+ days of improved reliability with community acknowledgment

Assessment

Status rationale

Assessed maintained (2026-06-30 Opus 4.8 re-eval). Frontier capabilities (GPT-5.5 benchmarks, /goal workflows, computer use, Codex Security, AWS Bedrock, FedRAMP Moderate, HIPAA-local) and exceptional growth (5M+ weekly users — up from 4M Apr and ~600K Jan, ~20% now non-developers — 40%+ enterprise revenue, $852B valuation, the formalized $150M Partner Network) firmly meet many Assessed signals, and the existential governance risk eased when the Musk v.

Altman claims were dismissed (May 18). Two caps remain active and were reinforced by June primary evidence: reliability-complaints (7+ Codex status incidents in June, most recent Jun 29 — not at the 90-day stability window; clock resets to ≥late-Sept 2026) and pricing-volatility (Jun 29 credit-burn incident + the still-in-force April token-pricing switch + new GPT-5.6 caching-price changes — 12-month window unmet, earliest ~Apr 2027).

These caps price the negative signals appropriately. A new but non-cap security negative was logged: CVE-2026-35603 (Windows local priv-esc on the Codex CLI) sits triaged 'Unresolved, no fix committed' while a direct competitor patched the same class — a compliance-narrative ding, but not critical-rated and not exploited, so it does not trip critical-security-vuln (and compliance is already capped at 12).

Demotion not warranted — caps are working as designed; upgrade to Validated when caps clear.

Watch for

Movement triggers

Upgrade to Validated if: (1) reliability-complaints cap cleared — 90+ days community-acknowledged stability with no major incidents, earliest ≥late-September 2026 given the June 29 incident; (2) pricing-volatility cap cleared — billing-anomaly complaints resolved with 12 months stability (earliest ~April 2027); (3) HIPAA cloud execution supported (not just local environments); (4) a hands-on UX/end-to-end assessment unlocks band 5 on the capability dims currently desk-capped at 16. Compliance score upgrades when pricing-volatility cap removed (would move toward 16-18 given FedRAMP Moderate + HIPAA-partial + Bedrock stack), provided CVE-2026-35603 is also remediated.

Demotion to Detected if: caps fail to clear by Q4 2026, billing anomalies escalate to enterprise contract disputes, reliability degradation worsens materially past current levels, or CVE-2026-35603 is exploited in the wild while unpatched. Monitor: GPT-5.6 GA (does it reduce incidents or add new ones per METR's cheating-rate flag); whether OpenAI ships a CVE-2026-35603 fix (none as of CLI 0.142.4 / changelog Jun 25); whether the Jun 29 usage-limits incident spawns a new billing-complaint wave; and whether the 2027 IPO slip is confirmed vs. market-timing reporting.

Caution

Risks & limitations

  • Reliability Complaints

    Moderate

    Radar cap: reliability-complaints

  • Pricing Volatility

    Moderate

    Radar cap: pricing-volatility

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.