LLM Observability With a Free Trial
As of September 2026, AIDiveForge tracks 5 llm observability with a free trial. The top three by verified-data score are Selfship.ai, Descles agent control plane, and Argosvix. Curated llm observability with a free trial tracked by AIDiveForge. Each tool listed is currently paid. Each tool below offers a time-limited free trial. Listings are verified against each tool's live website and re-checked regularly.
Last updated September 16, 2026 · 5 tools
Ranked by AIDiveForge's verified-data score: data completeness, verification recency, community rating, and real visitor engagement. How we rank · No tool can pay for placement.

1. Selfship.ai
The observe-diagnose-fix-verify cycle runs on live traffic without you assigning tasks or writing prompts. When a failure pattern crosses a threshold, Selfship identifies the exact conversation step that breaks, generates a proposed code change with eval evidence attached, and opens a PR in your GitHub. You review, you merge — nothing touches your code without your sign-off. After the merge, it watches real traffic to confirm the fix held; if it didn't, it opens a new PR with what it observed, up to a retry limit, then flags the item for human attention.
PaidFree Trial · 14 days$20/mo baseVerified Sep 8, 2026
2. Descles agent control plane
Descles sits between your existing SDK and your provider, requiring only an endpoint swap and a scoped key. Every request is bound to an identity — organization, group, agent — before it reaches the model, so attribution is structural, not reconstructed from logs after the fact. Budget enforcement and tool-call holds happen in the request path, not as post-hoc alerts. The live surface covers Model I/O for OpenAI and Anthropic traffic; Tool I/O and Resource I/O enforcement require runtime integration with the Action API or MCP adapter, which is a real integration step, not a config toggle. Teams that need enforceable tool-call governance without that runtime integration will find the approval loop is illustrative rather than operational.
PaidFree Trial · 30 daysAPIVerified Sep 16, 2026
3. Argosvix
Argosvix sits between your application and your LLM provider calls, scoring each response for quality, safety, PII exposure, and cost efficiency without requiring you to rewrite your call logic — the vendor describes setup as a single line pasted into Claude Code or Cursor. The dashboard surfaces call traces, latency trends, and error rates across OpenAI, Anthropic, Gemini, Mistral, Grok, Kimi, DeepSeek, and Qwen in a unified view. Anomalies surface in minutes, the vendor states. The free tier caps at 50,000 calls per month with 30-day retention — enough for eval and indie projects, but production volumes at any meaningful scale push you to a paid tier fast. There is no self-hosted option, which matters the moment your security team asks where the call data lives.
PaidFree Trial · 7 daysAPIVerified Aug 16, 2026
4. Latitude LLM
Latitude is an open-source AI agent monitoring platform that captures full conversation traces, clusters similar failures into triage-ready issue groups, and turns confirmed failure modes into automated evaluations that run against every new trace. The vendor states it ingests via OpenTelemetry, so teams already using OTEL pipelines point their existing setup at Latitude without reformatting data. Semantic search runs across 100% of traces — no sampling — which means finding 'frustrated users on a specific model version after a specific release' takes filters, not queries. The ceiling appears when your team needs the monitoring layer to also drive prompts or chain agents: that is not what this tool does.
PaidOpen SourceFree Trial · 30 days$99/monthAPISelf-hostedVerified Jun 24, 2026
5. Voker
Voker is a passive observability platform for conversational AI agents: it ingests chat session data, surfaces frustration patterns and knowledge gaps, and ties agent behavior to downstream metrics like conversion and retention. The self-hosted deployment path means your conversation data stays on your infrastructure — a hard requirement for many enterprise teams that competing SaaS observability tools cannot meet. The platform targets teams running at least 1,000 monthly sessions; below that threshold the pattern-detection signal is thin and the tooling is underutilized. Non-engineering teams can query agent insights without filing a ticket, which removes the bottleneck between product decisions and session data. Note: the scraped page content did not match Voker's product — factual claims here are drawn from the structured tool data provided.
PaidFree Trial · 30 days$80/moAPISelf-hostedVerified Jun 1, 2026
Listings on this page are sourced and verified by the AIDiveForge data pipeline. AIDiveForge is editorially independent — inclusion and rank are not for sale. Labeled ads are separate.