cto.new
Free multi-model agentic coding agent (Alloys: 7+ frontier models including Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, Kimi K2.6, GLM 5.1, MiniMax, Grok 4.20 — Claude Opus 4.8 added June 2026 alongside DeepSeek V4 Flash + Gemini 3.5 Flash, now the top in-product model) with SOC 2
Dimension breakdown
Score · confidenceFree multi-model agentic coding agent (Alloys: 7+ frontier models including Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, Kimi K2.6, GLM 5.1, MiniMax, Grok 4.20 — Claude Opus 4.8 added June 2026 alongside DeepSeek V4 Flash + Gemini 3.5 Flash, now the top in-product model) with SOC 2 Type II via Delve trust portal. Apr 2026 expansion: CTOClaw no-code agent builder, AI business multi-agent teams, agent marketplace, mobile dev via Expo, email-triggered agents, browser-during-run — moving from 'free code agent' toward broader agent platform.
Surfaces: web (primary), GitHub App (26K+ installs), a newly advertised free CLI, and MCP; no REST/headless automation API. Premium tier at $20/mo (rolling 24h/7d usage limits); individual adoption at Adobe/Docker/Twilio/PwC. Tracked for cost-conscious async delegation and multi-model experimentation.
No SSO/SAML (enterprise deployment blocker), no enterprise tier, no SLAs, $5.7M seed with no Series A, free-tier compute funded by data-collection/ads — sustainability and execution risk on the expanded product surface remain unproven; condition for Assessed is hands-on validation plus enterprise governance.
Use cases
Not yet assessed — this section fills in as ACES research covers the tool.
Risk flags
No active caps — no risk flags apply to this tool right now.
Status rationale
Tracked on Opus 4.8 re-baseline. SOC 2 Type II (Delve trust portal), Alloys multi-model orchestration across 7+ frontier models, the Apr 2026 product expansion (CTOClaw, agent marketplace, AI business teams), a new free CLI, 26K+ GitHub App installs, and individual adoption at Adobe/Docker/Twilio/PwC demonstrate genuine capability and traction; Premium ($20/mo) adds concrete monetization.
Evidence remains documentation- and third-party-based (not_tested, no enterprise-contract deployment), and the SSO/SAML gap, absent enterprise tier/SLAs, $5.7M seed with no Series A, and data-monetized free-tier compute hold it below Assessed. Promotion to Assessed requires hands-on validation plus enterprise governance maturity.
Movement triggers
Upgrade to Assessed if: (1) hands-on internal/third-party evaluation produces documented trial findings, OR (2) SSO/SAML launched AND enterprise tier with SLAs announced, OR (3) Series A closed with disclosed revenue, OR (4) independent third-party benchmark validates multi-agent team success rate at production scale. Downgrade to Detected if: funding distress signals emerge, compute-cost overrun forces free-tier restriction (triggering funding-concerns cap), security incident on agent marketplace/BYOK data flows (critical-security-vuln cap), or community-exodus signals appear (mass switching to alternatives).
Risks & limitations
Not yet assessed — this section fills in as ACES research covers the tool.
Integration surface
Not yet assessed — this section fills in as ACES research covers the tool.
Adoption & benchmarks
Not yet assessed — this section fills in as ACES research covers the tool.
Spotted something wrong or missing here? Suggest a change →
Per-source contributions
Click any dimension to see the underlying sources and citations.
More in this category