Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusWatch
SignalTracked
EvidenceGrade B

Grok Build

47
D 0 vs last quarter

xAI's CLI coding agent — 8 parallel subagents in isolated Git worktrees, Plan Mode (on by default), native MCP + Connectors (GitHub/Linear from the CLI), ACP custom orchestration, and headless mode (-p) for CI/CD. Local-first design (no source code transmitted; air-gap capable).

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
50% conf50
Capability
50% conf50
ReliabilityIncomplete data at this time
Value
50% conf30
CommunityIncomplete data at this time
Enterprise / Compliance
50% conf45
Autonomy
50% conf55
Integration
50% conf50
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

xAI's CLI coding agent — 8 parallel subagents in isolated Git worktrees, Plan Mode (on by default), native MCP + Connectors (GitHub/Linear from the CLI), ACP custom orchestration, and headless mode (-p) for CI/CD. Local-first design (no source code transmitted; air-gap capable). Runs a 2-model pipeline: Composer 2.5 (planning) + grok-build-0.1 (execution, 256K context, released May 20) — not a single model.

Ships /goal long-running autonomous execution (June 22 2026): plan/execute/verify loop with Agent Dashboard for real-time monitoring. Went GA May 25 to all SuperGrok ($30/mo) and X Premium+ ($40/mo) — public pricing, much broader access than the May 14 Heavy-only ($299) beta.

High Risk

all 11 co-founders departed, 50+ staff exits, pre-training lead gone and team gutted, xAI dissolved into SpaceXAI ($4.94B merger loss). grok-build-0.1 has NO independent benchmark validation — the only public figure (70.8% SWE-bench) is vendor-internal and belongs to the now-deprecated grok-code-fast-1 (rivals sit ~17pts higher).

Architecture is genuinely interesting and integration improved; organizational stability and capability validation are not.

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

  • Unvalidated benchmarks

    trustConditional

    Unvalidated benchmark claims

    Caps Autonomy at 70

    Removed whenIndependent benchmark validation (SWE-bench, Aider leaderboard, etc.) published

  • Acquisition uncertainty

    stabilityTemporary

    Acquisition with unclear product roadmap

    Caps Enterprise / Compliance at 60

    Removed whenClear post-acquisition roadmap published with commitment to existing customers

Assessment

Status rationale

HELD at Tracked (Opus 4.8 re-baseline). Research-based monitoring threshold is firmly met: dedicated vendor docs (docs.x.ai/build), a purpose-built backing model (grok-build-0.1, 256K), multiple independent third-party reviews and head-to-head comparisons, and sustained investigative coverage of company stability.

Cannot advance to Assessed without a hands-on team trial, organizational stabilization, and independent benchmark validation of grok-build-0.1 (which currently has none). acquisition-uncertainty caps the maximum signal at Tracked regardless. Integration was upgraded 10→12 this cycle on confirmed native MCP + GitHub/Linear Connectors usable from the CLI, but capability validation (benchmarks) and viability (6, Critical company health: pre-training team gutted, $4.94B merger loss) remain the binding constraints.

Both caps (acquisition-uncertainty, unvalidated-benchmarks) stay active and prevent Validated.

Watch for

Movement triggers

Upgrade to Assessed if: (a) Independent benchmarks improve materially (closing gap to vendor's 70.8% claim), (b) Hands-on team trial completed, (c) SpaceXAI confirms Grok Build as a continued strategic priority post-IPO, (d) Enterprise admin features (SSO/RBAC/audit) ship for Grok Build CLI specifically. Downgrade back to Detected if: (a) Product development visibly stalls post-IPO (no releases in 60+ days), (b) SpaceXAI publicly deprioritizes coding agent, (c) Independent benchmarks degrade further or critical security findings surface.

Caution

Risks & limitations

  • Acquisition Uncertainty

    Moderate

    Radar cap: acquisition-uncertainty

  • Unvalidated Benchmarks

    Moderate

    Radar cap: unvalidated-benchmarks

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.