Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusWatch
SignalAssessed
EvidenceGrade B

Factory

50
D -18 vs last quarter

[2026-06 refresh: Factory Router (model routing, ~20-25% token-cost cut, private preview), Factory 2.0 'Software Factory' repositioning + new customers Wipro/Blackstone/RBC/Comarch, Automated Security Review (GA, STRIDE/OWASP) + AutoWiki (GA), Deferred Context Engine (GA); CRO Ma

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
79% conf40
Capability
79% conf51
Reliability
79% conf14
Value
50% conf75
Community
79% conf0
Enterprise / Compliance
50% conf75
Autonomy
50% conf70
Integration
50% conf75
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

[2026-06 refresh: Factory Router (model routing, ~20-25% token-cost cut, private preview), Factory 2.0 'Software Factory' repositioning + new customers Wipro/Blackstone/RBC/Comarch, Automated Security Review (GA, STRIDE/OWASP) + AutoWiki (GA), Deferred Context Engine (GA); CRO Marcello Gallo + board member Lila Tretikov added. Context 15->16, viability 15->16, rating 75->77.] Enterprise autonomous agent platform; $150M Series C at $1.5B unicorn valuation (Khosla/Sequoia/Blackstone, April 2026, 5x in 7 months) and revenue doubling MoM for 6 months.

Missions runs true multi-day orchestration (orchestrator/workers/validators; median ~2hr, 14% >24hr, longest 16 days) with 10+ parallel Droids. Customers include Nvidia, Morgan Stanley, Adobe, EY, Palo Alto Networks, Adyen.

Surfaces: CLI/TUI, VS Code (+Cursor/Windsurf), JetBrains, Vim, Web, Desktop App (macOS/Win), Slack, Linear, plus Droid Exec headless mode for CI/CD. Compliance posture: ISO 42001, SOC 2 Type I (verified at factory.ai/security), GDPR/CCPA, SAML/OIDC SSO + SCIM, RBAC, BYOK, zero-data-retention, OTEL audit, EU data residency (shipped v0.126, May 2026), hybrid + fully airgapped deployment.

Monitoring

Terminal-Bench 2.0 leadership LOST — Droid+GPT-5.3-Codex 77.3% now ~#6; Codex CLI+GPT-5.5 leads at ~82% (still beats OpenAI's own Codex agent by 2.2pts on same model).

SOC 2 Type II claimed in some secondary outlets but NOT confirmed at primary source as of 2026-05-28 — Type I remains verified; FedRAMP and HIPAA BAA unconfirmed this cycle. Token-cost unpredictability, stuck-session, and response-latency complaints persist across 5+ independent sources; code quality remains repo-discipline dependent; always-on agents on roadmap but not shipped.

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

No active caps — no risk flags apply to this tool right now.

Assessment

Status rationale

Assessed because the platform has strong commercial momentum ($1.5B unicorn, doubling MoM for 6 months, blue-chip enterprise roster incl. Nvidia/Morgan Stanley/EY), genuine multi-day agentic orchestration (Missions), a broad interface surface (CLI/IDE/Web/Slack/Linear/Desktop/Headless Droid Exec), and a credible Enterprise-Ready governance composite (ISO 42001, SAML/OIDC SSO+SCIM, RBAC, BYOK, ZDR, OTEL audit, EU data residency, hybrid/airgapped).

Held at Assessed rather than Validated because (1) Terminal-Bench 2.0 leadership was lost this cycle (77.3% now ~#6 vs ~82% leaders), (2) primary-source SOC 2 Type II verification gap persists (Type I only at factory.ai/security), (3) reliability/token-cost complaint pattern spans 5+ independent reviews, (4) no hands-on testing has been performed, and (5) always-on agents remain announced not shipped.

Watch for

Movement triggers

Upgrade to Validated if: (1) SOC 2 Type II verified at primary source, (2) hands-on testing validates code quality and reliability, (3) token-cost predictability improves or pricing model adjusts, (4) always-on agents ship and perform reliably, (5) FedRAMP achieved, (6) regains independent benchmark leadership. Downgrade to Tracked if: SOC 2 Type I lapses, named enterprise departure, further benchmark regression, security incident with customer impact, or reliability complaints accelerate into a cap trigger (reliability-complaints if autonomy drops to <=12 band).

Caution

Risks & limitations

Not yet assessed — this section fills in as ACES research covers the tool.

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.