Skip to main content
World Wide TechnologyBenchAI tool benchmarks
Autonomous Agent
StatusEmerging
SignalTracked
EvidenceGrade B

Goose

56
C 0 vs last quarter

Open-source (Apache 2.0) AI agent under the Agentic AI Foundation (AAIF) at the Linux Foundation.

UX / DXCapabilityReliabilityValueCommunityEnterpriseAutonomyIntegration

Dimension breakdown

Score · confidence
UX / DX
79% conf55
Capability
79% conf62
Reliability
79% conf93
Value
50% conf60
Community
79% conf31
Enterprise / Compliance
50% conf35
Autonomy
50% conf50
Integration
50% conf60
Sources blend review platforms, community sentiment, the Signal Radar and practitioner ratings. Scores re-blend each quarter.

Open-source (Apache 2.0) AI agent under the Agentic AI Foundation (AAIF) at the Linux Foundation. Repo moved to github.com/aaif-goose/goose (AAIF/Linux Foundation donation, April 7 2026). ~49.8K GitHub stars, 474+ contributors, 136 releases — sustained velocity (v1.34 -> v1.36 in 2 weeks).

Runs locally with 25+ LLM providers. v1.36.0 (May 27, 2026) added a '/goal' self-evaluation step before task completion, a local 'goose review' code-analysis command, Hooks with pre/post tool-execution denial, deep linking (goose://resume), and a TUI command interface + diff viewer; v1.35.0 (May 22) shipped the Hooks system, ACP slash commands, unified thinking-effort control, and proactive OAuth token refresh. YAML Recipes for composable workflow automation, headless mode for CI/CD, and 'goose serve' background service distinguish it from competitors.

Strongest open-source data sovereignty story (fully air-gapped capable via Ollama). Block-proven at 12,000 engineers, 40% more code shipped; Block Q1 2026 record results (gross profit $2.91B, +27% YoY) reinforce commercial backing.

AAIF governance materially active: governing-board chair (David Nalley, AWS), 146+ member orgs. No enterprise compliance for the agent itself (no SOC 2, SSO, RBAC, audit logs) — a third-party-review claim of a SOC 2 Type II hosted instance is uncited and unconfirmed by official sources.

No independent SWE-bench validation. No native codebase semantic indexing.

Recommended

Use cases

Not yet assessed — this section fills in as ACES research covers the tool.

Score caps

Risk flags

  • No codebase indexing

    capabilityConditional

    No codebase indexing or semantic search capability

    Caps Capability at 55

    Removed whenCodebase indexing or RAG capability launched

  • Unvalidated benchmarks

    trustConditional

    Unvalidated benchmark claims

    Caps Autonomy at 70

    Removed whenIndependent benchmark validation (SWE-bench, Aider leaderboard, etc.) published

Assessment

Status rationale

Tracked because active development under AAIF governance (~49.8K stars, 474+ contributors, 136 releases — v1.34 to v1.36 in 2 weeks; repo moved to github.com/aaif-goose/goose April 7 2026), strong Block commercial backing (Q1 2026 record results, 'intelligence company' strategy), and strongest open-source data sovereignty story. AAIF governance materially active (governing-board chair, 146+ members) — not just a press-release shell.

Assessment is research-based, not hands-on, which caps confidence at the Tracked band. Blocked from advancement by: (1) zero enterprise compliance infrastructure for the agent (no SOC 2, SSO, RBAC, audit logs); (2) no independent benchmark validation; (3) no native codebase semantic indexing.

Watch for

Movement triggers

Upgrade if: (a) Enterprise features ship for the agent — SSO/SAML, audit logs, SOC 2 (would lift compliance out of Band 1); (b) Independent SWE-bench results published (removes unvalidated-benchmarks cap, autonomy could move to 15-16); (c) Native codebase indexing/embeddings shipped (removes no-codebase-indexing cap, context could move to 13-14); (d) Sustained release velocity + viability evidence justifies viability 13->14; (e) Hands-on validation (handsOn -> demo/tested) would raise evidence grade and could support an Assessed signal level. Downgrade if: AAIF/Block reduces investment, major unpatched security vulnerability, or reliability issues (latency, stuck agents, loops) escalate to warrant a reliability-complaints cap.

Caution

Risks & limitations

  • Unvalidated Benchmarks

    Moderate

    Radar cap: unvalidated-benchmarks

  • No Codebase Indexing

    Moderate

    Radar cap: no-codebase-indexing

Capabilities

Integration surface

Not yet assessed — this section fills in as ACES research covers the tool.

Proof points

Adoption & benchmarks

Not yet assessed — this section fills in as ACES research covers the tool.

Spotted something wrong or missing here? Suggest a change →

Per-source contributions

Click any dimension to see the underlying sources and citations.