Codex
Cloud-first autonomous coding agent from OpenAI, now at 5M+ weekly users (June 2; ~20% non-developers, repositioned as 'Codex for knowledge work').
Dimension breakdown
Score · confidenceCloud-first autonomous coding agent from OpenAI, now at 5M+ weekly users (June 2; ~20% non-developers, repositioned as 'Codex for knowledge work'). Production model is GPT-5.5 (SWE-bench Pro 58.6%, Terminal-Bench 2.0 82.7%) with computer use, /goal persistent workflows, plugin marketplace, Chrome extension, iOS/Android mobile, Codex Remote (GA Jun 25), and AWS Bedrock (GA).
A GPT-5.5 successor — GPT-5.6 (Sol/Terra/Luna) — was previewed Jun 26 but is GATED to ~20 US-government-approved partners (not GA), so it is a forward signal only; METR flagged its highest detected-cheating rate of any model evaluated. FedRAMP Moderate confirmed; HIPAA for local environments only.
The $150M OpenAI Partner Network (Jun 14) formalizes the SI channel (Accenture, Bain, BCG, McKinsey, PwC + the prior 7 SIs). Enterprise now 40%+ of OpenAI revenue (Q2-2026 ~$10.9B projected); the Musk v.
Altman governance threat was DISMISSED (May 18 verdict — Altman/Brockman retained, for-profit stands; Musk appealing).
Two caps active and reinforced — reliability-complaints (autonomy ≤12; 7+ Codex status incidents in June incl.
Jun 29 'usage limits depleting' and Jun 11 GPT-5.5 error-rate; clock resets to ≥late-Sept 2026) and pricing-volatility (compliance ≤12; Jun 29 credit-burn incident + GPT-5.6 caching-price changes). Open security item: CVE-2026-35603 (Jun 17 Windows local priv-esc on the Codex CLI) is triaged 'Unresolved, no fix committed' as of CLI 0.142.4 — Anthropic patched the same class, OpenAI did not (not critical-rated, no in-the-wild exploitation).
Use cases
Not yet assessed — this section fills in as ACES research covers the tool.
Risk flags
Pricing volatility
enterpriseTemporaryFrequent pricing changes causing budget unpredictability
Caps Enterprise / Compliance at 60
Removed when — 12 months of pricing stability with no user complaints about billing surprises
Reliability complaints
trustTemporaryWidespread reliability complaints (breaks often, unreliable output)
Caps Autonomy at 60
Removed when — 90+ days of improved reliability with community acknowledgment
Status rationale
Assessed maintained (2026-06-30 Opus 4.8 re-eval). Frontier capabilities (GPT-5.5 benchmarks, /goal workflows, computer use, Codex Security, AWS Bedrock, FedRAMP Moderate, HIPAA-local) and exceptional growth (5M+ weekly users — up from 4M Apr and ~600K Jan, ~20% now non-developers — 40%+ enterprise revenue, $852B valuation, the formalized $150M Partner Network) firmly meet many Assessed signals, and the existential governance risk eased when the Musk v.
Altman claims were dismissed (May 18). Two caps remain active and were reinforced by June primary evidence: reliability-complaints (7+ Codex status incidents in June, most recent Jun 29 — not at the 90-day stability window; clock resets to ≥late-Sept 2026) and pricing-volatility (Jun 29 credit-burn incident + the still-in-force April token-pricing switch + new GPT-5.6 caching-price changes — 12-month window unmet, earliest ~Apr 2027).
These caps price the negative signals appropriately. A new but non-cap security negative was logged: CVE-2026-35603 (Windows local priv-esc on the Codex CLI) sits triaged 'Unresolved, no fix committed' while a direct competitor patched the same class — a compliance-narrative ding, but not critical-rated and not exploited, so it does not trip critical-security-vuln (and compliance is already capped at 12).
Demotion not warranted — caps are working as designed; upgrade to Validated when caps clear.
Movement triggers
Upgrade to Validated if: (1) reliability-complaints cap cleared — 90+ days community-acknowledged stability with no major incidents, earliest ≥late-September 2026 given the June 29 incident; (2) pricing-volatility cap cleared — billing-anomaly complaints resolved with 12 months stability (earliest ~April 2027); (3) HIPAA cloud execution supported (not just local environments); (4) a hands-on UX/end-to-end assessment unlocks band 5 on the capability dims currently desk-capped at 16. Compliance score upgrades when pricing-volatility cap removed (would move toward 16-18 given FedRAMP Moderate + HIPAA-partial + Bedrock stack), provided CVE-2026-35603 is also remediated.
Demotion to Detected if: caps fail to clear by Q4 2026, billing anomalies escalate to enterprise contract disputes, reliability degradation worsens materially past current levels, or CVE-2026-35603 is exploited in the wild while unpatched. Monitor: GPT-5.6 GA (does it reduce incidents or add new ones per METR's cheating-rate flag); whether OpenAI ships a CVE-2026-35603 fix (none as of CLI 0.142.4 / changelog Jun 25); whether the Jun 29 usage-limits incident spawns a new billing-complaint wave; and whether the 2027 IPO slip is confirmed vs. market-timing reporting.
Risks & limitations
Reliability Complaints
ModerateRadar cap: reliability-complaints
Pricing Volatility
ModerateRadar cap: pricing-volatility
Integration surface
Not yet assessed — this section fills in as ACES research covers the tool.
Adoption & benchmarks
Not yet assessed — this section fills in as ACES research covers the tool.
Spotted something wrong or missing here? Suggest a change →
Per-source contributions
Click any dimension to see the underlying sources and citations.
More in this category