Skip to main content
AIDiveForge AIDiveForge

DiffForge vs improv.sh

DiffForge and improv.sh are both cli coding agents tracked by AIDiveForge. Below is a side-by-side comparison of pricing, capabilities, platforms, and ownership — sourced from each tool's live website and verified before publishing.

DiffForge

DiffForge

The tool runs Codex, Claude Code, and OpenCode side by side in local terminals, with a kernel that leases files so concurrent agents cannot touch the same path at once. Loop Spaces add scheduled blueprint graphs — think cron jobs, but the steps are agent handoffs and verification scripts rather than shell commands. Voice dictation runs locally via Whisper or through the cloud, and screen snips can be dragged directly into a prompt, so you can point at a bug rather than describe it. Token usage and credit events stay visible per provider in real time, which matters the moment you are running three agents against three different API accounts simultaneously. The self-hosted option keeps code on your machine — only commands travel over the wire.

improv.sh

improv.sh

improv operates as a task harness: the @im call pulls repo context, detects your test commands, writes acceptance criteria, and packages shell validation steps into one spec the agent can implement on turn one. The loop infrastructure is the distinguishing piece — judges run your actual exit-code commands (npm test, typecheck, build), so done means your tests pass, not that the agent says it's done. The tool installs locally via curl with no external API keys required, and the Chrome extension brings the same engine into web-based chat interfaces. The 920-skill library and daily auto-research loop suggest the routing layer will keep growing — but the page offers no independent benchmarks to validate the token-savings figures cited.

AttributeDiffForgeimprov.sh
PricingPaidFree
Free trialNoNo
Open sourceNoYes
Has APIYesNo
Self-hosted optionYesYes
PlatformsDesktop app with web dashboard and device syncVS Code, Cursor, Claude Code, terminal, Chrome
Pros
  • File lease coordination at the kernel level, so three agents editing the same repo never produce a simultaneous write conflict — without this, you are manually partitioning work or running agents sequentially.
  • Loop Spaces schedule agent handoffs as blueprint graphs, so repetitive development cycles — run agent, verify output, trigger next step — run unattended instead of requiring you to babysit each transition.
  • Local-first execution with remote command queuing, which means code stays on your machine while you steer the session from a phone or second device — avoiding the data-exposure tradeoff of fully cloud-hosted alternatives.
  • Per-provider token metering with live pace forecasts, so you catch a runaway agent burning through API credits before the bill arrives rather than after.
  • Local Whisper dictation plus screen snip injection, which means you can describe a visual problem by showing it to the agent instead of translating it into text — cutting prompt-writing time on UI and design-adjacent tasks.
  • Repo-aware spec compilation pulls your actual test commands and package scripts into the task before the agent starts, which means the agent implements against your real constraints instead of inventing them mid-run.
  • Exit-code judges close the loop on real shell commands — npm test, typecheck, build — so you are not relying on the agent's self-assessment of whether it finished.
  • Task memory persisted under .improv/tasks/ survives session boundaries, so an agent restarted mid-task picks up status and spec instead of starting the discovery cycle again.
  • Local-first install with no external API keys required, which means the harness runs in air-gapped or locked-down environments where cloud tooling is blocked.
  • Chrome extension and VS Code/Cursor Marketplace extension share the same local engine, so the spec compilation and judge loop work whether you are in the IDE or a browser-based chat interface — without switching tabs.
Cons
  • The coordination architecture is built around a single local desktop runtime. Teams expecting multiple developers to share one forge session — running agents collaboratively from separate machines — will find this model does not fit; at that scale, teams move to server-side orchestration platforms designed for multi-user access.
  • Loop Spaces blueprint graphs are a visual scheduling layer. When your agent pipeline requires branching logic that responds to dynamic output — agents that fork differently based on what the previous step returned — the blueprint canvas is the constraint. Community-reported workarounds involve scripting the branching externally and invoking Loop Spaces as leaf nodes, which means maintaining coordination logic in two places.
  • The tool is closed-source, so the coordination kernel, file lease logic, and Loop Spaces scheduler cannot be audited, patched, or extended at the source level. Teams in regulated environments that require full-stack auditability of execution infrastructure treat this as a disqualifying constraint and evaluate open-source alternatives instead.
  • Task state is written to .improv/tasks/ on the local machine. Teams with more than one developer working the same codebase have no shared task state — there is no sync layer described on the page — so parallel agent runs on different machines produce divergent task records with no reconciliation path.
  • The tool exposes no API surface, so teams that want to trigger improv from a CI pipeline or wrap it in a custom orchestration layer cannot. Teams hitting this wall move to harness frameworks that expose programmatic interfaces — at which point they are maintaining the prompt compilation logic themselves.
  • The token-savings figures on the page (~613 tokens median) are vendor-reported with no independent reproduction methodology described. Teams making adoption decisions based on cost reduction should treat these numbers as illustrative until they run their own baseline comparison.
  • Chrome extension installation requires either the Chrome Web Store or a manual sideload script — neither path is available in Firefox or Safari. Teams on non-Chromium browsers are limited to the terminal install, losing the browser chat integration entirely.
Bottom line

DiffForge is paid while improv.sh is free; improv.sh is open source; only DiffForge exposes a public API. Choose based on which difference matters most for your workflow.

Frequently asked questions

What is the difference between DiffForge and improv.sh?

DiffForge is Paid, while improv.sh is Free and open source. Compare pricing, free trial, API, platforms, and pros/cons in the table above on AIDiveForge.

Is DiffForge better than improv.sh?

It depends on your workflow. Use the side-by-side attributes (pricing, open source, API, self-hosted, platforms) to decide. AIDiveForge does not rank a universal winner — we publish verified facts so you can choose.

DiffForge vs improv.sh: which should I pick?

Pick DiffForge if its pricing model, openness, or platform fit matches your constraints; pick improv.sh otherwise. Check free-trial availability on each listing if you want to test before committing.

Comparison data is sourced and verified by the AIDiveForge data pipeline. AIDiveForge is editorially independent.