Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusEmerging
SignalTracked
EvidenceGrade B

Kimi Code

55
C 0 vs last quarter

Terminal-first agentic coding CLI from Moonshot AI, now powered by Kimi K2.7 Code (released June 12 2026; supersedes K2.6).

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
50% conf60
Capability
50% conf65
ReliabilityIncomplete data at this time
Value
50% conf60
CommunityIncomplete data at this time
Enterprise / Compliance
50% conf25
Autonomy
50% conf75
Integration
50% conf45
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Terminal-first agentic coding CLI from Moonshot AI, now powered by Kimi K2.7 Code (released June 12 2026; supersedes K2.6). K2.7 vendor SWE-bench Pro score is 58.6 — a different metric from K2.6's SWE-bench Verified 80.2 (do not conflate).

K2.6 baseline capabilities: 300-sub-agent swarm, 80.2% SWE-bench Verified, 12-hour autonomous runs. Open source (Apache-2.0 CLI / open weights) with aggressive API pricing ($0.60/$2.50 per M on the official API; ~$0.73/$3.49 via OpenRouter — still multiples cheaper than the frontier Claude Opus line).

Commercial momentum is strong: Moonshot raised $2B at $20B valuation (May 7), ARR topped $200M (April, +100% MoM), and Kimi K2.6 is the #1 model by token volume on OpenRouter (1.88T tokens/week).

Caution

zero enterprise governance, a confirmed user-data-isolation breach (OECD.AI, Apr 2026), CVE-2026-25046 (dev-script scope), an active U.S. congressional investigation into PRC AI models naming Moonshot, and Italy GDPR block plus Australia/Canada/Netherlands/South Korea restrictions make this unsuitable for enterprise use without significant risk mitigation.

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

No active caps — no risk flags apply to this tool right now.

Assessment

Status rationale

Tracked retained. K2.6 capabilities are genuinely frontier-adjacent (80.2% SWE-bench Verified vs Opus 4.6 at 80.8%; the current Claude Opus line still leads on the hardest single-shot benchmarks, 300-agent swarm, #1 on OpenRouter by raw token volume) and commercial momentum is materially strong after the $2B/$20B round and $200M+ ARR. However, the risk profile remains the binding constraint: a confirmed user-data-isolation breach (OECD.AI, Apr 2026), CVE-2026-25046, an active congressional investigation now formally naming Moonshot among PRC AI models, multi-country regulatory restrictions, and zero enterprise governance.

Harmonic Security data showing Kimi's enterprise shadow usage far ahead of other Chinese AI tools indicates adoption is materially ahead of governance — elevating organizational risk for any enterprise that has not explicitly evaluated. Tracked reflects research-based monitoring without hands-on testing; advancement to Assessed would require named enterprise references and a remediated security posture.

Watch for

Movement triggers

Upgrade if: Enterprise features launched (SSO, SAML, SCIM, RBAC, audit logs, SOC 2 or ISO 27001 certification), congressional investigation resolved favorably, named Western enterprise customer references published, data-isolation breach remediated with post-mortem, Kimi Claw moves from research preview to GA with an audited security model. Downgrade if: additional or unremediated security incidents, formal U.S. sanctions or enterprise bans, congressional investigation escalates to enforcement action, first-party Kimi Code CLI tool-calling reliability deteriorates (not just OpenClaw integration layer), CLI development stalls (>90 days no release), or sentiment deteriorates below 3.

Caution

Risks & limitations

Not yet assessed — this section fills in as ACES research covers the tool.

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.