Skip to main content
AIDiveForge AIDiveForge

Memharness vs Shepherd

Memharness and Shepherd are both agent frameworks tracked by AIDiveForge. Below is a side-by-side comparison of pricing, capabilities, platforms, and ownership — sourced from each tool's live website and verified before publishing.

Memharness

Memharness

The core premise is storing facts, not strings, with two independent time axes: when something became true in the world and when the agent learned it — so querying past agent states is a real query, not archaeology through logs. Everything lives in a single SQLite file, which means the storage layer makes zero LLM or network calls and stays auditable. Recall combines hybrid vector search and full-text search with a source-staleness signal, so older or superseded sources rank down automatically. Where it breaks: the SQLite backend is a hard ceiling for teams expecting distributed writes or high-concurrency production deployments. Teams hitting that ceiling will need to treat memharness as a pattern to port, not a service to scale horizontally.

Shepherd

Shepherd

SHEPHERD is a Python substrate from Stanford and Northeastern that turns an agent's execution into a Git-like, reversible trace — so a supervising meta-agent can observe, intercept, fork, and revert any step without rebuilding that capability from scratch each time. The vendor-published benchmark numbers are specific: a supervisor meta-agent lifted pair-coding pass rate from 28.8% to 54.7% on CooperBench; a counterfactual repair meta-agent beat MetaHarness on Terminal-Bench 2.0 by 12.8% while cutting wall-clock time by 58%. The framework is research-grade and open-source, installed via pip. Teams outside the specific use cases the paper targets — runtime intervention, counterfactual optimization, and agentic RL training — will find precious little guidance on how far the substrate stretches.

AttributeMemharnessShepherd
PricingFreeFree
Free trialNoNo
Open sourceYesYes
Has APIYesNo
Self-hosted optionYesYes
PlatformsSQLite, MCPPython
Released2026
Pros
  • Bi-temporal storage tracks both world-time and agent-learn-time independently, so you can reconstruct exactly what the agent believed at any past moment — which means post-incident reviews and compliance audits have an actual record to query instead of inferring from logs.
  • Provenance-scoped deletion lets you remove all facts derived from a specific source in one operation, so GDPR takedown requests or source revocations do not require a full memory wipe that destroys unrelated facts.
  • The storage layer makes zero LLM or network calls, so memory reads and writes have no latency dependency on external APIs and no token cost — which means memory operations do not blow your inference budget.
  • Hybrid vector-plus-full-text recall with a built-in staleness signal means older or superseded sources rank lower automatically, so the agent surfaces the most current relevant facts without you writing custom re-ranking logic.
  • MCP exposure and a self-hosted SQLite backend mean the tool drops into any agent stack that speaks MCP without requiring a separate managed service, so you retain full data ownership and avoid a vendor dependency in the memory layer.
  • Git-like reversible execution traces built into the substrate, so a meta-agent can revert a worker to any prior state without custom snapshot logic that teams would otherwise rebuild from scratch on every project.
  • Fork-and-replay from any past checkpoint, which means a counterfactual optimizer can test a corrected decision path without re-running the entire prior sequence — the vendor reports 58% lower wall-clock versus MetaGarness on Terminal-Bench 2.0.
  • Meta-agents and worker agents share the same @task code interface, so the control layer does not require a separate DSL or framework to learn — it is plain Python decorated functions.
  • Open-source with pip install and self-hosting support, so teams running sensitive codebases can keep execution fully on-premise with no data leaving their environment.
  • Intercept hooks let a meta-agent catch a destructive action before it lands, rather than reading about it in a post-mortem transcript — the supervisor use case lifted CooperBench pass rate from 28.8% to 54.7%.
Cons
  • SQLite is a single-writer database: the moment two agent processes attempt concurrent writes — a parallelized pipeline, a multi-worker deployment, any architecture where more than one process holds the file — writes will collide or block. Teams with concurrent-write requirements either serialize all memory operations through a single process (adding a bottleneck) or abandon memharness for a Postgres- or Redis-backed alternative.
  • The project has 2 stars and 1 fork on GitHub at time of curation, with 19 commits and no open issues, which means community-sourced debugging, third-party integrations, and production war stories are essentially nonexistent. Teams that hit an edge case are reading the source, not a Stack Overflow thread.
  • There is no built-in access control or multi-tenant isolation: if multiple agents or users share the same SQLite file, provenance-scoped deletion could become a liability rather than a feature — one delete call wipes facts for every tenant who learned from that source. Teams building multi-user applications will need to implement per-user database files or a sharding layer before going to production.
  • The framework's documented capabilities cover exactly three use cases from the paper; teams that need meta-agent patterns outside runtime intervention, counterfactual optimization, or agentic RL training will find no templates, examples, or community patterns to lean on — they are extending a research prototype.
  • There is no API, which means SHEPHERD cannot be called from a non-Python orchestration layer or integrated into an existing service mesh without a custom wrapper — teams with polyglot architectures hit this wall immediately and typically reach for a framework with a REST interface instead.
  • The Claude CLI dependency in the interactive demo signals the substrate's current depth of LLM provider integration; teams that cannot or will not use Anthropic models during onboarding face an underdocumented offline path before they have validated the tool for their use case.
  • Research-grade codebase with no paid support tier means production incidents land entirely on the team's own debugging of the substrate — organizations that need an SLA or vendor escalation path will abandon SHEPHERD before the first outage.
Bottom line

Only Memharness exposes a public API. Choose based on which difference matters most for your workflow.

Frequently asked questions

What is the difference between Memharness and Shepherd?

Memharness is Free and open source, while Shepherd is Free and open source. Compare pricing, free trial, API, platforms, and pros/cons in the table above on AIDiveForge.

Is Memharness better than Shepherd?

It depends on your workflow. Use the side-by-side attributes (pricing, open source, API, self-hosted, platforms) to decide. AIDiveForge does not rank a universal winner — we publish verified facts so you can choose.

Memharness vs Shepherd: which should I pick?

Pick Memharness if its pricing model, openness, or platform fit matches your constraints; pick Shepherd otherwise. Check free-trial availability on each listing if you want to test before committing.

Comparison data is sourced and verified by the AIDiveForge data pipeline. AIDiveForge is editorially independent.