Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusEmerging
SignalTracked
EvidenceGrade B

Cosine Genie

59
C 0 vs last quarter

Fully autonomous AI software engineer -- first-cohort partner in UK's £500M Sovereign AI programme (April 2026), granted 500K GPU hours on Isambard-AI. Deployed across UK defence primes and critical national infrastructure including nuclear deterrent programmes.

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
50% conf55
Capability
50% conf55
ReliabilityIncomplete data at this time
Value
50% conf50
CommunityIncomplete data at this time
Enterprise / Compliance
50% conf60
Autonomy
50% conf70
Integration
50% conf65
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Fully autonomous AI software engineer -- first-cohort partner in UK's £500M Sovereign AI programme (April 2026), granted 500K GPU hours on Isambard-AI. Deployed across UK defence primes and critical national infrastructure including nuclear deterrent programmes. Platform runs entirely within customer infrastructure with zero external dependencies (on-prem, VPC, air-gapped, zero-egress).

Lumen own-model family (production-first models post-trained for COBOL, Fortran, ABAP, Verilog, Rust, SQL) layers model IP on top of the Genie agent. Pricing now fully public incl. a Free tier (80 tasks, no card), Hobby $20/seat, Professional $200/seat, Enterprise custom.

Genie 2.1: ~$107K of $250K SWE-Lancer (independent). Caution: Lumen Niche-Bench claims remain vendor-only; SOC2/ISO self-described not certified; ~$6-8M raised, 32 staff against Cognition/Devin's now-closed $1B/$26B round that won US Army/Navy, Goldman, Citi -- the exact regulated/defence niche Cosine targets.

Strong UK-sovereign differentiator; competitive asymmetry widening.

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

No active caps — no risk flags apply to this tool right now.

Assessment

Status rationale

Tracked maintained (Opus 4.8 re-baseline, net-flat). Cosine now owns both the agent layer and a model layer (Lumen), with niche/legacy-language specialization, and pricing is now fully public including a Free tier -- both reduce TCO/transparency concerns.

UK Sovereign cohort membership, 500K GPU hours, and air-gapped defence/CNI deployments remain the strongest validation signal and a genuine sovereignty moat. However, the moat is under intensifying pressure: Cognition/Devin's $1B/$26B raise (May 27, 2026) and named US Army/Navy, Goldman, Citi customers directly target the regulated/defence segment, and the independent SWE-Bench gap to frontier agents (Claude Code 87.6% vs ~43.8%) is wide.

Lumen Niche-Bench claims stay vendor-only, SOC2/ISO are self-described not certified Type II, Enterprise pricing is contact-sales, the company is ~$6-8M raised / 32 staff with no priced Series A, and no hands-on internal evaluation has been done. Sovereign positioning differentiates but narrows the addressable market for general enterprise adoption.

Watch for

Movement triggers

Upgrade to Assessed if: independent third-party validation of Lumen Outpost/Niche-Bench claims published, first named Fortune 500 / major commercial (non-government) customer disclosed, a priced Series A closes (or UK Sovereign Fund exercises its round option), SOC2 Type II fully certified, or hands-on internal evaluation completed. Downgrade if: UK Sovereign programme encounters political headwinds or replacement, funding runway exhausted without a new round, Lumen benchmark claims debunked by independent testing, development stalls (>90d), or a well-capitalized competitor (e.g.

Cognition/Devin) displaces Cosine in the UK sovereign/defence programme.

Caution

Risks & limitations

Not yet assessed — this section fills in as ACES research covers the tool.

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.