Skip to main content
AIDiveForge AIDiveForge

Live Captions by Subanana vs VoxRT Wake-Word

Live Captions by Subanana and VoxRT Wake-Word are both audio & voice tracked by AIDiveForge. Below is a side-by-side comparison of pricing, capabilities, platforms, and ownership — sourced from each tool's live website and verified before publishing.

Live Captions by Subanana

Live Captions by Subanana

The tool covers four distinct workflows under one interface: video subtitling with glossary enforcement, verbatim transcription with word-level speaker separation, meeting capture without requiring a bot to join the call, and live captioning for in-room or public-display audiences across 95+ languages. Dual ASR engines run per language pair with millisecond timecodes, and custom glossaries correct terminology before translation — each substitution logged. Export options include SRT, VTT, FCPXML, XLSX, Markdown, and burned-in video up to 4K. The free tier caps projects at 15 minutes, which surfaces the wall fast for anyone processing long-form content. No API is available, so teams that need to wire this into an existing pipeline hit a dead end and look elsewhere.

VoxRT Wake-Word

VoxRT Wake-Word

The SDK ships a Rust runtime under 1 MB with wake-word models around 100 KB, so it fits on mobile and IoT targets without gutting your memory budget. Audio stays on the device — the vendor states models are encrypted at rest and the system works offline by default, which means GDPR and HIPAA conversations get simpler, not harder. The published models are free for commercial use; custom models trained to your phrase, accent profile, or domain vocabulary are a paid engagement. iOS and Android are available in v1; Windows, WebAssembly, microcontrollers, automotive, and wearables are listed as v2, meaning shipping on those targets today is not an option. Teams that need a language other than English are also waiting — multilingual support is post-v1 on the roadmap.

AttributeLive Captions by SubananaVoxRT Wake-Word
PricingPaidPaid
Free trialNoNo
Open sourceNoNo
Has APINoNo
Self-hosted optionNoYes
PlatformsWeb, Chrome extensioniOS 16+, Android 8.0+, Linux, macOS, Windows, microcontrollers (ARM Cortex-M), Raspberry Pi, Jetson
Released2026
Pros
  • Dual ASR engines run per language pair with millisecond timecodes and silence recovery, so the exported file stays frame-accurate even when audio quality dips between speakers.
  • Glossary enforcement corrects domain-specific terms before translation and logs every substitution, which means brand names, product terms, and specialized vocabulary survive the language switch without a manual review pass.
  • Word-level speaker diarization splits overlapping voices at word boundaries and carries named roster labels through every export format, so a two-hour interview with four speakers arrives as a quotable, attributed transcript rather than an undifferentiated wall of text.
  • Export covers SRT, VTT, FCPXML, XLSX, Markdown, and burned-in video up to 4K, which means the same processed file hands off to a video editor, a data analyst, and a publishing workflow without conversion steps.
  • A no-bot browser extension captures meetings without joining as a participant, so teams whose platforms block third-party bots can still get a transcript without requesting IT exceptions.
  • Runtime under 1 MB with wake-word models around 100 KB, so the SDK fits on memory-constrained mobile and embedded targets where competing runtimes cannot be installed.
  • No cloud round-trip and no per-detection fees, which means always-on listening stays within battery and cost budgets that would make a cloud-dependent architecture unshippable.
  • Audio never leaves the device and models are encrypted at rest, so voice features pass privacy and compliance reviews that would block any SDK sending audio to a third-party server.
  • Voice activity detection gates the heavier models, so the battery drain of continuous microphone monitoring is cut to the minimum — critical for wearables and IoT where always-on is the use case.
  • Published models are free for commercial use with no account required, so a team can validate accuracy on real hardware before committing to a paid custom-model engagement.
Cons
  • The free tier caps each project at 15 minutes, so a 90-minute interview or a two-hour event recording hits the wall on the first upload — teams processing long-form content regularly are immediately into paid territory and need to budget accordingly before starting.
  • No API is available, which means every file requires a manual upload through the web interface. Teams that generate transcription jobs programmatically — automated ingest pipelines, post-production workflows triggered by a CI step — cannot integrate this tool and move to a competitor that exposes an endpoint.
  • Live captioning and meeting transcription depend on a stable connection to the Subanana service with no self-hosted option, so organizations under strict data-residency requirements or operating in environments where outbound connections to third-party SaaS are restricted cannot deploy this tool.
  • Microcontroller targets — ARM Cortex-M4, M7, M33, M55, M85 — are listed as v2 and not available. Teams building firmware for these chips today cannot use VoxRT and will need a competitor like Picovoice Porcupine or Arm's ML Embedded Evaluation Kit, which already ship no_std-compatible binaries.
  • English is the only supported language in v1. A product shipping to Spanish or French-speaking markets has no path forward with VoxRT until post-v1 multilingual support lands — no timeline is stated on the vendor page.
  • Custom model training — tuning the wake phrase to your brand name, accent distribution, or noise profile — is a paid vendor engagement, not a self-service pipeline. Teams that expected to iterate on model accuracy independently will find themselves dependent on VoxRT's turnaround cycle for each training run.
  • Windows and WebAssembly support is v2, meaning browser-based demos and Windows desktop apps cannot ship with VoxRT in v1. Teams prototyping on the web before committing to a mobile build lose the ability to test the actual SDK in that environment.
Bottom line

Live Captions by Subanana and VoxRT Wake-Word are closely matched on pricing model, openness, and API availability — pick by feature set and platform support in the table above.

Frequently asked questions

What is the difference between Live Captions by Subanana and VoxRT Wake-Word?

Live Captions by Subanana is Paid, while VoxRT Wake-Word is Paid. Compare pricing, free trial, API, platforms, and pros/cons in the table above on AIDiveForge.

Is Live Captions by Subanana better than VoxRT Wake-Word?

It depends on your workflow. Use the side-by-side attributes (pricing, open source, API, self-hosted, platforms) to decide. AIDiveForge does not rank a universal winner — we publish verified facts so you can choose.

Live Captions by Subanana vs VoxRT Wake-Word: which should I pick?

Pick Live Captions by Subanana if its pricing model, openness, or platform fit matches your constraints; pick VoxRT Wake-Word otherwise. Check free-trial availability on each listing if you want to test before committing.

Comparison data is sourced and verified by the AIDiveForge data pipeline. AIDiveForge is editorially independent.