Greptile
Full-codebase-aware AI code review agent with semantic code graph indexing and multi-hop agentic investigation (Claude Agent SDK) — genuine technical differentiation catching cross-file bugs competitors miss. v4 (March 2026) is the current release; vendor A/B tests show 74% more
Dimension breakdown
Score · confidenceFull-codebase-aware AI code review agent with semantic code graph indexing and multi-hop agentic investigation (Claude Agent SDK) — genuine technical differentiation catching cross-file bugs competitors miss. v4 (March 2026) is the current release; vendor A/B tests show 74% more addressed comments and 43% acceptance. Production-deployed at NVIDIA, Brex, and Coinbase (confirmed via Anthropic's own customer story) plus 2,000+ orgs / 9,000+ teams.
A fresh independent 3-week, 146-PR benchmark ranked Greptile #1 on precision with a 0% false-positive rate — a positive counter-signal to the older community narrative.
Per-review pricing ($1/review after 50/month) controversy unresolved through end of May 2026 — HN front page, dedicated greptile.fail critique site, at least one documented customer cancellation, predatory math for high-velocity/agentic teams.
OSS free-review promise still violated (maintainers billed, refunds only after public pressure). SOC 2 claim remains unverified post-Delve (Greptile named among 58 firms; 493/494 reports fabricated) with no AICPA-registered re-certification announced.
No in-app cancel button, no spending caps, possible California ARL exposure. Vendor headline benchmarks (82% catch rate) still lack first-party independent validation.
Use cases
Not yet assessed — this section fills in as ACES research covers the tool.
Risk flags
Pricing volatility
enterpriseTemporaryFrequent pricing changes causing budget unpredictability
Caps Enterprise / Compliance at 60
Removed when — 12 months of pricing stability with no user complaints about billing surprises
Severe negative sentiment
trustTemporaryWidespread negative sentiment (sentiment score 1-2)
Removed when — Sustained sentiment improvement over 90+ days with community acknowledgment of fixes
Unvalidated benchmarks
trustConditionalUnvalidated benchmark claims
Caps Autonomy at 70
Removed when — Independent benchmark validation (SWE-bench, Aider leaderboard, etc.) published
Status rationale
Tracked — genuine technical differentiation (semantic code graph, multi-hop agentic investigation) confirmed in production at named enterprise customers (NVIDIA, Brex, Coinbase) and now corroborated by an independent precision benchmark. $25M Series A and ~28-person team intact. But unresolved pricing controversy with at least one documented cancellation, unverified SOC 2 post-Delve, OSS promise violations, missing cancel/spending safeguards, and a high vendor-reported false-positive rate compound trust and retention risks.
No path to Assessed without hands-on validation and resolution of compliance/pricing gaps.
Movement triggers
Upgrade if: Legitimate SOC 2 re-certification from AICPA-registered firm published with auditor name and report date; per-review pricing revised or grandfathered; OSS free-review promise enforced in practice; in-app cancel button shipped; independent benchmark validates the vendor's headline bug catch rate. Downgrade to Detected if: customer churn confirmed at scale in financial/employee data; SOC 2 fraud confirmed; significant data breach; sentiment continues to deteriorate without vendor response.
Risks & limitations
Pricing Volatility
ModerateRadar cap: pricing-volatility
Unvalidated Benchmarks
ModerateRadar cap: unvalidated-benchmarks
Severe Negative Sentiment
ModerateRadar cap: severe-negative-sentiment
Integration surface
Not yet assessed — this section fills in as ACES research covers the tool.
Adoption & benchmarks
Not yet assessed — this section fills in as ACES research covers the tool.
Spotted something wrong or missing here? Suggest a change →
Per-source contributions
Click any dimension to see the underlying sources and citations.
More in this category