Devin Review
AI code reviewer for GitHub (full parity) / GitHub Enterprise Server (limited) / GitLab (GA June 19 2026: MR diffs, inline comments, AI chat) PRs with semantic diff organization, severity-ranked bug detection, codebase-aware chat that proposes and commit-applies edits, an auto-fi
Dimension breakdown
Score · confidenceAI code reviewer for GitHub (full parity) / GitHub Enterprise Server (limited) / GitLab (GA June 19 2026: MR diffs, inline comments, AI chat) PRs with semantic diff organization, severity-ranked bug detection, codebase-aware chat that proposes and commit-applies edits, an auto-fix agent loop (autofixes its own and other bots' review comments until CI passes), and a REST API v3 with an enterprise review-status polling endpoint. Auto-review triggers, auto-merge, and CI status checks.
Inherits Cognition's enterprise compliance (SOC 2 Type II, SSO/SAML/OIDC, VPC/air-gapped, US/EU data residency). Real-world validation: flagged the March 31 2026 malicious-axios npm supply-chain attack for customers under an hour before public disclosure.
still no independent accuracy/false-positive benchmarks for Review specifically, GHES gaps (no comment posting / review submission / merge), and enterprise references are for the broader Devin agent, not Review.
Use cases
Not yet assessed — this section fills in as ACES research covers the tool.
Risk flags
Unvalidated benchmarks
trustConditionalUnvalidated benchmark claims
Caps Autonomy at 70
Removed when — Independent benchmark validation (SWE-bench, Aider leaderboard, etc.) published
Status rationale
Assessed because Devin Review continues strong iteration (REST API v3 + enterprise polling endpoint, GHES support, closed-loop auto-fix, codebase-aware chat with commit-apply, GitLab interactive read-write preview) and Cognition's company health is now exceptional (closed $1B round at $26B post-money, $492M ARR, blue-chip + government customers), but the product still lacks independent accuracy benchmarks for Review specifically, GitLab is not GA, GHES has documented capability gaps (no comment posting / review submission / merge), and it has not been hands-on tested. Enterprise adoption signals (Goldman Sachs, Nubank, Mercedes, NASA, Itaú, US Army/Navy) remain for the broader Devin agent — not Devin Review specifically.
CodeRabbit still dominates the dedicated AI code-review category (2M+ repos, 13M+ PRs, 8,000+ paying companies, GitHub/GitLab/Bitbucket/ADO). Holding at Assessed pending independent validation of review accuracy, GitLab GA, GHES full parity, and Review-specific named enterprise references.
Movement triggers
(1) Independent benchmarks published for Devin Review specifically (accuracy, false-positive rate) — removes unvalidated-benchmarks cap and is the primary gate to Validated. (2) GitLab reaches full GA for Review. (3) GHES reaches full feature parity (comment posting, review submission, merge). (4) Review-specific named enterprise customers with production metrics (current references are for the Devin agent). (5) Hands-on internal trial with measured catch/false-positive rate.
(1) Accuracy/false-positive complaints emerge at scale on Review specifically. (2) Pricing restructure adds opacity or billing-surprise complaints spike. (3) GitLab preview dropped or stalls. (4) CodeRabbit/competitor signs major joint deals locking out Review. (5) Material adverse change in Cognition health (now low risk post-$1B raise).
Risks & limitations
Unvalidated Benchmarks
ModerateRadar cap: unvalidated-benchmarks
Integration surface
Not yet assessed — this section fills in as ACES research covers the tool.
Adoption & benchmarks
Not yet assessed — this section fills in as ACES research covers the tool.
Spotted something wrong or missing here? Suggest a change →
Per-source contributions
Click any dimension to see the underlying sources and citations.
More in this category