Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusWatch
SignalAssessed
EvidenceGrade B

Kilo Code

51
C- 0 vs last quarter

Open-source agentic engineering platform (Apache 2.0) on a shared CLI core across VS Code, JetBrains, CLI, and KiloClaw cloud agents. ~1.5M users, #1 on OpenRouter by token volume (6T+/month).

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
77% conf66
Capability
77% conf59
Reliability
79% conf39
Value
50% conf50
Community
79% conf10
Enterprise / Compliance
50% conf50
Autonomy
50% conf60
Integration
50% conf75
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Open-source agentic engineering platform (Apache 2.0) on a shared CLI core across VS Code, JetBrains, CLI, and KiloClaw cloud agents. ~1.5M users, #1 on OpenRouter by token volume (6T+/month). Agent Manager runs 2-8+ parallel agents in git worktree isolation; 500+ models via BYOK. INDEPENDENT VALIDATION: Martian Code Review Bench (DeepMind/Anthropic/Meta researchers) ranks Kilo Code Reviewer #1 among open-source code review tools across precision- and recall-optimized profiles (CodeRabbit leads overall F1).

Named enterprise customers: Plug&Pay, Ida Infront (Swedish government, 70 seats), Dona Dowlan, Andrey Guenov. Enterprise tier adds SSO/SCIM/RBAC/audit.

Caution

No SOC 2 / ISO 27001 / FedRAMP. GitLab right-of-first-refusal active through August 24, 2026.

Still $8M seed only. Quality strain with ~1,175+ open issues; recurring stuck-loop / token-burn reports (Trustpilot 2.7, low volume).

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

  • Acquisition uncertainty

    stabilityTemporary

    Acquisition with unclear product roadmap

    Caps Enterprise / Compliance at 60

    Removed whenClear post-acquisition roadmap published with commitment to existing customers

Assessment

Status rationale

Assessed because (a) strong agentic capabilities now have partial independent benchmark validation (Martian Code Review Bench #1 open-source), (b) named customer base diversifying (4 cases including Swedish government 70-seat deployment), (c) accelerating adoption (~1.5M users, 6T+ tokens/month, #1 OpenRouter, stars/forks growing), BUT (d) governance gaps persist (no SOC 2/ISO 27001), (e) acquisition uncertainty material (GitLab ROFR through Aug 24, 2026), (f) seed-stage funding modesty ($8M, no Series A), (g) enterprise pricing opacity, and (h) reliability/billing complaints isolated but persistent. Promotion to Validated requires SOC 2 or ISO 27001 cert AND F500-class customer reference AND ROFR resolution.

Watch for

Movement triggers

Upgrade if: SOC 2 Type II or ISO 27001 achieved (compliance 10→12+, status candidate); GitLab ROFR expires Aug 24 2026 without exercise (remove acquisition-uncertainty cap, viability could move to 11); Series A funding announced (viability 10→11-12); F500-class named customer confirmed; independent end-to-end benchmark (SWE-bench Verified or Aider Leaderboard) entry published (autonomy could move toward 16). Downgrade if: GitLab exercises ROFR (acquisition disruption, viability/compliance downward pressure); reliability/token-burn complaints intensify to a widespread pattern (reliability-complaints cap on autonomy, autonomy≤12); code indexing regresses again; Martian open-source result disputed or not reproduced.

Caution

Risks & limitations

  • Acquisition Uncertainty

    Moderate

    Radar cap: acquisition-uncertainty

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.