Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Workflow Tool
StatusEmerging
SignalTracked
EvidenceGrade B

CrewAI

61
C+ 0 vs last quarter

Leading open-source multi-agent orchestration framework with role-based crew architecture, native MCP/A2A protocol support, and AMP enterprise platform. 52.4K GitHub stars (up from 47.8K in April), 27M+ PyPI downloads, ~450M agentic executions/month (2B+ cumulative over 12 months

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
50% conf60
Capability
50% conf60
ReliabilityIncomplete data at this time
Value
50% conf55
CommunityIncomplete data at this time
Enterprise / Compliance
50% conf65
Autonomy
50% conf60
Integration
50% conf65
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Leading open-source multi-agent orchestration framework with role-based crew architecture, native MCP/A2A protocol support, and AMP enterprise platform. 52.4K GitHub stars (up from 47.8K in April), 27M+ PyPI downloads, ~450M agentic executions/month (2B+ cumulative over 12 months), used by 40-63% of the Fortune 500 (IBM, Microsoft, PwC, NVIDIA, Deloitte, Oracle, KPMG, Accenture, Capgemini, P&G, Walmart, SAP, Adobe, PayPal). SOC2/HIPAA compliance confirmed on Enterprise tier.

The four March 2026 CVEs (VU#221883) are now RESOLVED — CERT/CC advisory Rev 3 (May 20) confirms vendor remediation: Code Interpreter Tool removed (PR #4791/#5309, allow_code_execution deprecated) plus centralized path/URL validators (PR #5310/#5315). LangGraph still leads independent production benchmarks (76% vs CrewAI 71% medium-complexity) and a recurring 'prototype in CrewAI, ship in LangGraph' migration pattern persists (CrewAI carries ~18-56% token overhead).

Series A only ($18M Oct 2024) with no confirmed Series B — runway watch warranted for a 29-person team scaling enterprise commitments. Tracked for production multi-agent orchestration; advancement to Assessed needs internal hands-on pilot and Series B / SOC2 Type II clarity.

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

No active caps — no risk flags apply to this tool right now.

Assessment

Status rationale

Tracked maintained. Positive: SOC2/HIPAA enterprise compliance, March 2026 CVEs (VU#221883) now fully remediated per CERT/CC Rev 3, stable v1.14.6 shipped May 28 with further hardening, ~450M executions/month at scale, GitHub momentum (52.4K stars), 40-63% Fortune 500 penetration with an expanding named roster (IBM, Microsoft, PwC, NVIDIA, Deloitte, Oracle, KPMG, Accenture, Capgemini, P&G, Walmart, SAP, Adobe, PayPal).

Not yet Assessed because: (1) no internal hands-on evaluation performed (handsOn=not_tested) — now the primary gate, (2) only $18M Series A with no confirmed Series B raises runway questions for a 29-person team, (3) LangGraph still leads independent production benchmarks (76% vs 71%) and the 'prototype in CrewAI, ship in LangGraph' migration pattern persists on token-cost/conditional-logic grounds, (4) SOC2 Type II not formally confirmed (Type I implied).

Watch for

Movement triggers

Upgrade to Assessed if: (1) internal hands-on evaluation completed with documented pilot results (the now-binding gate — handsOn=not_tested) AND (2) SOC2 Type II audit (vs current Type I implied) confirmed AND (3) Series B funding closed or alternative runway clarity provided.

Note

the prior CVE-closure condition is now SATISFIED — CERT/CC VU#221883 Rev 3 confirms full vendor remediation.

Downgrade to Detected if: new unpatched critical CVEs emerge OR Series A runway exhausted without follow-on funding OR sustained community exodus to LangGraph beyond the current benchmark/token-overhead gap.

Caution

Risks & limitations

Not yet assessed — this section fills in as ACES research covers the tool.

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.