Skip to main content
AIDiveForge AIDiveForge

LoopTroop vs Shepherd

LoopTroop and Shepherd are both agent frameworks tracked by AIDiveForge. Below is a side-by-side comparison of pricing, capabilities, platforms, and ownership — sourced from each tool's live website and verified before publishing.

LoopTroop

LoopTroop

The tool orchestrates a local pipeline — LLM council planning, an iterative execution loop called Ralph, and OpenCode worktree isolation — designed for multi-file feature work where correctness matters more than turnaround time. Every ticket goes through an interview phase before a line is code is written, resolving ambiguities via adaptive question batches that the vendor describes as intentionally taking over an hour. You review diffs and sign off before anything reaches your main branch. The tradeoff is explicit: LoopTroop is slow by design. Teams treating it as a fast pair-programmer will be frustrated inside the first session.

Shepherd

Shepherd

SHEPHERD is a Python substrate from Stanford and Northeastern that turns an agent's execution into a Git-like, reversible trace — so a supervising meta-agent can observe, intercept, fork, and revert any step without rebuilding that capability from scratch each time. The vendor-published benchmark numbers are specific: a supervisor meta-agent lifted pair-coding pass rate from 28.8% to 54.7% on CooperBench; a counterfactual repair meta-agent beat MetaHarness on Terminal-Bench 2.0 by 12.8% while cutting wall-clock time by 58%. The framework is research-grade and open-source, installed via pip. Teams outside the specific use cases the paper targets — runtime intervention, counterfactual optimization, and agentic RL training — will find precious little guidance on how far the substrate stretches.

AttributeLoopTroopShepherd
PricingFreeFree
Free trialNoNo
Open sourceYesYes
Has APINoNo
Self-hosted optionYesYes
PlatformsLocal desktop (JavaScript GUI)Python
Released2026
Pros
  • 100% local execution with no cloud routing, so proprietary codebases never leave the host and there is no per-request cost accumulating against an API quota.
  • Git worktree isolation for every in-progress change, which means reviewing or discarding a bad AI-generated diff is a clean branch delete rather than a manual undo across modified files.
  • Multi-model council planning before any code is written, so spec ambiguities surface as explicit questions you answer rather than silent assumptions that break three files later.
  • Manual approval gate on every bead of changes before commit, so no AI-generated code reaches your main branch without your explicit sign-off — eliminating the 'it shipped before I reviewed it' failure mode.
  • Free and MIT-licensed, so there is no vendor lock-in and the orchestration logic is auditable and forkable by the team maintaining it.
  • Git-like reversible execution traces built into the substrate, so a meta-agent can revert a worker to any prior state without custom snapshot logic that teams would otherwise rebuild from scratch on every project.
  • Fork-and-replay from any past checkpoint, which means a counterfactual optimizer can test a corrected decision path without re-running the entire prior sequence — the vendor reports 58% lower wall-clock versus MetaGarness on Terminal-Bench 2.0.
  • Meta-agents and worker agents share the same @task code interface, so the control layer does not require a separate DSL or framework to learn — it is plain Python decorated functions.
  • Open-source with pip install and self-hosting support, so teams running sensitive codebases can keep execution fully on-premise with no data leaving their environment.
  • Intercept hooks let a meta-agent catch a destructive action before it lands, rather than reading about it in a post-mortem transcript — the supervisor use case lifted CooperBench pass rate from 28.8% to 54.7%.
Cons
  • Speed is architecturally sacrificed: the interview phase alone is described as taking over an hour by design, which means LoopTroop is the wrong tool for any task where you need a working diff in minutes rather than hours — teams with fast-iteration workflows will abandon it for a standard AI coding assistant after the first blocked sprint.
  • No external API surface is available, so the pipeline cannot be triggered from CI, scripts, or external tooling — every run starts from the local GUI, which blocks any team wanting to embed AI coding steps into an automated workflow.
  • The pipeline stages are fixed — interview, plan, execute, review — and the docs describe no mechanism for custom branching or conditional routing between stages; teams whose tasks require dynamic mid-run replanning must intervene manually or restart the ticket.
  • The framework's documented capabilities cover exactly three use cases from the paper; teams that need meta-agent patterns outside runtime intervention, counterfactual optimization, or agentic RL training will find no templates, examples, or community patterns to lean on — they are extending a research prototype.
  • There is no API, which means SHEPHERD cannot be called from a non-Python orchestration layer or integrated into an existing service mesh without a custom wrapper — teams with polyglot architectures hit this wall immediately and typically reach for a framework with a REST interface instead.
  • The Claude CLI dependency in the interactive demo signals the substrate's current depth of LLM provider integration; teams that cannot or will not use Anthropic models during onboarding face an underdocumented offline path before they have validated the tool for their use case.
  • Research-grade codebase with no paid support tier means production incidents land entirely on the team's own debugging of the substrate — organizations that need an SLA or vendor escalation path will abandon SHEPHERD before the first outage.
Bottom line

LoopTroop and Shepherd are closely matched on pricing model, openness, and API availability — pick by feature set and platform support in the table above.

Frequently asked questions

What is the difference between LoopTroop and Shepherd?

LoopTroop is Free and open source, while Shepherd is Free and open source. Compare pricing, free trial, API, platforms, and pros/cons in the table above on AIDiveForge.

Is LoopTroop better than Shepherd?

It depends on your workflow. Use the side-by-side attributes (pricing, open source, API, self-hosted, platforms) to decide. AIDiveForge does not rank a universal winner — we publish verified facts so you can choose.

LoopTroop vs Shepherd: which should I pick?

Pick LoopTroop if its pricing model, openness, or platform fit matches your constraints; pick Shepherd otherwise. Check free-trial availability on each listing if you want to test before committing.

Comparison data is sourced and verified by the AIDiveForge data pipeline. AIDiveForge is editorially independent.