Skip to main content
AIDiveForge AIDiveForge

Jaybase vs Shepherd

Jaybase and Shepherd are both agent frameworks tracked by AIDiveForge. Below is a side-by-side comparison of pricing, capabilities, platforms, and ownership — sourced from each tool's live website and verified before publishing.

Jaybase

Jaybase

Jaybase stores every agent-generated fact as an immutable, time-stamped record, which means the full sequence of what an agent wrote, when, and why is always recoverable. The vendor describes it as designed for accounting, compliance, and approval workflows where you cannot afford to lose the paper trail. Because it is append-only, there is no overwrite risk — replaying a sequence from any point is a native operation. The library is self-hostable and open-source under AGPL, so it runs inside your own infrastructure without a call home. The project has a small contributor footprint, which means production teams should expect to own gaps in documentation rather than wait for the maintainer to fill them.

Shepherd

Shepherd

SHEPHERD is a Python substrate from Stanford and Northeastern that turns an agent's execution into a Git-like, reversible trace — so a supervising meta-agent can observe, intercept, fork, and revert any step without rebuilding that capability from scratch each time. The vendor-published benchmark numbers are specific: a supervisor meta-agent lifted pair-coding pass rate from 28.8% to 54.7% on CooperBench; a counterfactual repair meta-agent beat MetaHarness on Terminal-Bench 2.0 by 12.8% while cutting wall-clock time by 58%. The framework is research-grade and open-source, installed via pip. Teams outside the specific use cases the paper targets — runtime intervention, counterfactual optimization, and agentic RL training — will find precious little guidance on how far the substrate stretches.

AttributeJaybaseShepherd
PricingFreeFree
Free trialNoNo
Open sourceYesYes
Has APIYesNo
Self-hosted optionYesYes
PlatformsLinux, DockerPython
Released2026-072026
Pros
  • Append-only storage model, so no agent write operation can silently overwrite a prior fact — which means compliance audits have a complete, tamper-evident record rather than the last-write-wins snapshot most databases produce.
  • Native replay support, so when a non-deterministic agent produces a different result on a rerun, you can trace the exact prior sequence and compare it against the new one rather than guessing what changed.
  • Self-hosted with no vendor dependency, so financial or regulated data never leaves your own infrastructure and you are not subject to a hosted service's availability or policy changes.
  • AGPL open-source license, so the full codebase is auditable — which matters in regulated environments where black-box dependencies fail security review.
  • API available, so agents written in any language can append facts without being locked to a specific SDK or framework.
  • Git-like reversible execution traces built into the substrate, so a meta-agent can revert a worker to any prior state without custom snapshot logic that teams would otherwise rebuild from scratch on every project.
  • Fork-and-replay from any past checkpoint, which means a counterfactual optimizer can test a corrected decision path without re-running the entire prior sequence — the vendor reports 58% lower wall-clock versus MetaGarness on Terminal-Bench 2.0.
  • Meta-agents and worker agents share the same @task code interface, so the control layer does not require a separate DSL or framework to learn — it is plain Python decorated functions.
  • Open-source with pip install and self-hosting support, so teams running sensitive codebases can keep execution fully on-premise with no data leaving their environment.
  • Intercept hooks let a meta-agent catch a destructive action before it lands, rather than reading about it in a post-mortem transcript — the supervisor use case lifted CooperBench pass rate from 28.8% to 54.7%.
Cons
  • The append-only model has no native support for queries beyond sequential log reads — teams that need to filter, aggregate, or join across fact types have to build that layer themselves, and at any meaningful fact volume that becomes a non-trivial engineering project.
  • With a single maintainer and a small contributor base, documentation gaps are yours to resolve: the README covers the core path, but edge cases in approval workflow design or multi-agent sequencing have no community forum depth to fall back on.
  • Teams that eventually need multi-tenant isolation, role-based fact access, or a managed hosted option will find none of those features in the current architecture — that is the condition under which teams move to a purpose-built audit log service with a commercial support tier.
  • The framework's documented capabilities cover exactly three use cases from the paper; teams that need meta-agent patterns outside runtime intervention, counterfactual optimization, or agentic RL training will find no templates, examples, or community patterns to lean on — they are extending a research prototype.
  • There is no API, which means SHEPHERD cannot be called from a non-Python orchestration layer or integrated into an existing service mesh without a custom wrapper — teams with polyglot architectures hit this wall immediately and typically reach for a framework with a REST interface instead.
  • The Claude CLI dependency in the interactive demo signals the substrate's current depth of LLM provider integration; teams that cannot or will not use Anthropic models during onboarding face an underdocumented offline path before they have validated the tool for their use case.
  • Research-grade codebase with no paid support tier means production incidents land entirely on the team's own debugging of the substrate — organizations that need an SLA or vendor escalation path will abandon SHEPHERD before the first outage.
Bottom line

Only Jaybase exposes a public API. Choose based on which difference matters most for your workflow.

Frequently asked questions

What is the difference between Jaybase and Shepherd?

Jaybase is Free and open source, while Shepherd is Free and open source. Compare pricing, free trial, API, platforms, and pros/cons in the table above on AIDiveForge.

Is Jaybase better than Shepherd?

It depends on your workflow. Use the side-by-side attributes (pricing, open source, API, self-hosted, platforms) to decide. AIDiveForge does not rank a universal winner — we publish verified facts so you can choose.

Jaybase vs Shepherd: which should I pick?

Pick Jaybase if its pricing model, openness, or platform fit matches your constraints; pick Shepherd otherwise. Check free-trial availability on each listing if you want to test before committing.

Comparison data is sourced and verified by the AIDiveForge data pipeline. AIDiveForge is editorially independent.