Video Tools With an API
As of August 2026, AIDiveForge tracks 42 video tools with an api. The top three by verified-data score are Palmier Pro, PixVerse, and Vidmoat. Curated video tools with an api tracked by AIDiveForge. Listings are verified against each tool's live website and re-checked regularly.
Last updated July 29, 2026 · 42 tools
Ranked by AIDiveForge's verified-data score: data completeness, verification recency, community rating, and real visitor engagement. How we rank · No tool can pay for placement.

1. Palmier Pro
The editor runs on macOS and requires macOS 26 (Tahoe) — that version gate will stop teams on older hardware cold. Inside that constraint, the core promise holds: generation from Kling, Seedance, Veo, and other models happens inline, and the resulting clips land on a multi-track timeline alongside your own footage. The MCP server is the differentiating detail — Claude Desktop, Cursor, and Codex can read your project structure and write edits back to the timeline directly, which means an agent can do the grunt work while you review the cut. When you need to hand off to an established post-production suite, NLE XML export covers Premiere Pro and DaVinci Resolve.
PaidOpen SourcePro $29/mo, Max $69/moAPIVerified Jul 24, 2026
2. PixVerse
PixVerse covers the full content-creation surface: text-to-video, image-to-video, multi-shot scene structuring, lip sync with emotion-driven character performance, and style-level video editing. Character Reference lets you anchor a face or subject across shots from one image, which is the feature that collapses when you try to approximate it with generic generation models. The API makes it scriptable for teams running batch or production workflows. Where it breaks: fine-grained directorial control — precise camera paths, physics fidelity, frame-by-frame timing — stays shallow compared to dedicated compositing pipelines. Teams that outgrow the canvas-level controls end up wrapping the API in a custom layer.
Paid$4.80/minAPIVerified Jul 14, 2026
3. Vidmoat
Vidmoat's Auto-Cut feature ingests long raw files, removes silences, and assembles an editable cut without manual trimming. The Moat AI agent accepts plain-language prompts — 'make this a punchy TikTok' — and executes multi-step edits: captions, color grade, dead-air removal, in sequence, narrating each step. The MCP server layer is the actual differentiator: external agents like Claude Code or Cursor connect via a single API key and drive the full timeline — 65+ commands, frame previews returned as images, rendered MP4 out the other side. Where it breaks: teams needing granular manual control over complex narrative structures will hit the ceiling of what a prompt-driven agent can reliably interpret. No self-hosted option exists, so regulated industries with strict data residency requirements cannot use this.
PaidAPIVerified Jul 21, 2026
4. FableCut
FableCut is a browser-based, Premiere-style non-linear video editor with zero npm dependencies, designed from the ground up so that AI agents — Claude Code, Claude Desktop, or anything that speaks MCP or REST — can drive the timeline directly. The JSON document *is* the project: agents write to it, the UI reflects the change live. That's the promise. The wall appears when you need effects, color grading, audio mixing, or any of the post-production work that professional editors expect — the docs describe a lean, agent-first tool, not a full-featured studio. Teams that hit that ceiling move to a traditional NLE and use FableCut only for the automated rough-cut stage.
FreeOpen SourceAPISelf-hostedVerified Jul 9, 2026
5. BatchEdits
BatchEdits runs that pipeline on up to 50 clips at once — silence removal, auto captions across 50+ languages, auto zoom, and platform-specific cropping for 9:16 or 16:9 output — in what the vendor describes as under five minutes per batch. The workflow is upload, pick your style settings once, export. One Product Hunt user reported silence trimming across 12 talking-head clips landed clean without cutting mid-syllable, with captions close to accurate on the first pass. The ceiling appears when your editing needs move beyond that fixed pipeline: custom cuts, color grading, or anything requiring per-clip creative decisions stay manual. Teams with those requirements use BatchEdits for the mechanical layer and a separate editor for everything else.
PaidAPIVerified Jul 24, 2026
6. Clipin AI
The vendor describes a pipeline where a text idea or reference photo moves through script generation, automatic storyboarding, and clip rendering without leaving the platform. Switching between text-to-video and image-to-video in the same generation flow means you can anchor a scene to a specific frame when the prompt alone won't get you there. The free tier runs on a starter-credit policy — enough to evaluate output quality before committing. Where the ceiling shows up: output is fixed-length clips between 3 and 15 seconds, so anything requiring longer continuous footage or editorial sequencing lands outside what this tool handles. Teams needing that reach for a proper NLE after export.
PaidAPIVerified Jul 20, 2026
7. Fliki
The core workflow is paste-and-generate: drop in a URL, script, or prompt, and Fliki writes the script, selects stock visuals, attaches an AI voice, adds music and subtitles, and returns a video in minutes. The vendor reports 2,000+ AI voices across 80+ languages, and the digital twin feature lets you record once and generate presenter videos in any supported language from that single recording. Where the ceiling appears is creative control — the visual selection is AI-driven, so when brand-specific imagery or precise scene composition matters, you are fighting the defaults. Teams producing high-volume, format-consistent social content hit their stride here; teams whose brand guidelines require custom motion graphics or granular edit control hit the wall fast.
Paid$28/mo or $14/mo annuallyAPIVerified Jun 29, 2026
8. Riffkit
The core workflow is three steps: paste a winning TikTok link, let Riffkit extract the hook-turn-payoff structure, and receive a post-ready video with caption, hashtags, and cover frame. The vendor states outputs ship with your product woven into the story, a locked persona for consistent characters across a series, and native English and Spanish generation — not subtitle overlays. Billing runs per second of rendered video, so a nine-second clip costs nine seconds at the rated price. The ceiling appears when you need a language beyond the two supported, or when your creative brief requires a structure the source formula doesn't contain. Teams with higher-volume pipelines pipe it through an AI agent using the described command interface, but that integration path has no published SDK docs on the vendor page.
Paid$8–9.9 per video or from $99/mo on plan; billed by the secondAPIVerified Jul 3, 2026
9. SynthCut
SynthCut exposes a full multi-track, frame-based video editor as an MCP server, so any MCP-compatible AI client — Claude Desktop being the documented example — drives real FFmpeg operations locally, fully offline. The architecture sidesteps the cloud-dependency and privacy concerns that come with hosted video AI tools. Where it breaks: the AI client is doing the driving, which means your workflow ceiling is whatever your MCP client can reason about and whatever tools SynthCut exposes. The project has 21 commits and 3 stars at time of scrape — early-stage by any measure. Teams that need a mature plugin ecosystem or a GUI-first fallback will hit that wall fast.
FreeOpen SourceAPISelf-hostedVerified Jul 23, 2026
10. Topview
TopView assembles that missing crew into a single canvas: an AI Video Agent that takes a prompt, a script, or a reference URL and routes the work through scene planning, character handling, lip sync, and style matching without you stitching tools together. The Smart Canvas workspace keeps projects, assets, and credits in one place, which matters when an agency is running parallel ad sets for five clients. The wall appears at the editorial layer — fine-grained timeline control, frame-accurate cuts, and complex audio mixing are not here. Teams that need broadcast-grade post-production finish in a dedicated editor; TopView is where the generation happens, not where the finishing does.
PaidAPIVerified Jul 24, 2026
11. Vidu S1
The platform delivers streaming video avatars — human, anime, or mascot — driven by voice, text, and visual input over a WebRTC plus WebSocket control flow. The integration sequence is explicit: create a session server-side, join the RTC room, open the control channel, maintain heartbeats, hang up deliberately, then pull billed usage. Provider routing lets teams slot HeyGen API for presenter-style video and Replicate for async model jobs alongside the core Vidu S1 stream, all behind a single server-side orchestration layer that keeps credentials off the browser. This is a pilot-scoped tool — the docs describe a flow built for evaluation, not a drop-in widget for existing platforms. Teams shipping to production at scale will need to instrument session state, retry logic, and usage review themselves.
PaidAPIVerified Jul 8, 2026
12. VlogMe
VlogMe threads those pieces together through a chat-based director workflow: you describe the goal, the AI prepares a full scene plan with script, voice, music, and captions, and you approve it before anything renders. Each scene stays independently editable after the fact, so fixing one line does not mean starting the whole production over. The Video Studio layer adds eight purpose-built single-shot workflows — text to scene, still image to motion, lip sync, restyle — feeding results back into the larger project. The model roster pulls from Google, ByteDance, Kuaishou, Kling, and xAI, letting you route each shot to the engine that handles it best. The ceiling shows up when your production logic gets complex: the director workflow is a linear approval loop, not a branching system, so anything requiring conditional structure or non-linear scene logic goes beyond what the chat interface was built for.
PaidAPIVerified Jul 21, 2026
13. Skryber
The tool ingests video from YouTube, TikTok, Instagram, and 1,800-plus other sources or direct uploads up to 2 GB, then auto-reframes to 9:16 with speaker tracking, strips silences and filler words, applies karaoke-style captions, and exports in 4K. AI dubbing across 33 languages uses voice cloning so each original speaker retains their own sound. The narration feature watches silent footage and writes synced copy — useful for drone or B-roll channels that have no on-camera voice. The minute-based billing model means you only pay for successfully exported content, not for processing that fails. The ceiling appears when you need editorial judgment the pipeline cannot make: which three clips from a two-hour recording are actually worth publishing.
PaidAPIVerified Jul 22, 2026
14. VideoInPrompt
The tool accepts MP4, MOV, or WEBM uploads, samples keyframes, runs vision-model analysis on scene context, and returns either natural language prompts or structured JSON schemas ready for downstream LLMs and image generators. The JSON output — covering scene, lighting, motion, and a ready-to-paste AI prompt — is the differentiating artifact for developers wiring this into automation pipelines via API. It fits tightly scoped, single-video jobs: repurposing a TikTok, cloning a competitor ad's visual language, pulling SEO metadata from a product demo. The vendor does not describe batch processing, multi-video comparison, or any output editing layer on the page, so teams processing hundreds of videos per day will hit workflow gaps that a single-conversion tool cannot close.
Paid$12/moAPIVerified Jun 28, 2026
15. A2E Canvas
A2E generates avatar-led videos from text scripts, letting marketing teams, L&D professionals, and developers produce localized video at volume without cameras, microphones, or actors on set. The core workflow is text-in, video-out: write a script, pick or clone an avatar, select a language, and export. The vendor states support for 40+ languages with voice cloning that retains original tone across translations. The free tier provides 30 daily credits, which is enough to prototype but falls short of production-scale batch generation — that requires a paid-only tier. Teams hitting the canvas on throughput or needing white-labeled output in their own applications route through the API.
Paid$14.9 one-time or $0 freeAPISelf-hostedVerified Jun 1, 2026
16. Akool
The platform covers avatar video generation, face swap, video translation with lip-sync, image generation, background replacement, and voice cloning — meaning a marketing team can take one asset through localization, persona swap, and audio rebrand without leaving the tool. The vendor states 4K diffusion-based rendering with temporal consistency, which matters when your avatar needs to hold the same face across a 90-second spot. Where the ceiling appears: AKOOL is a one-shot generation and editing suite, not an autonomous agent, so any workflow requiring conditional logic between steps gets built outside — in your own orchestration layer. Self-hosting is not an option, which means your assets and voice clones live on AKOOL's infrastructure. Teams with strict data-residency requirements hit that wall fast.
Paid$21/mo for ProAPIVerified Jun 20, 2026
17. D-ID
D-ID lets you feed a script, image, and voice into its API or web interface and get back a finished video of a digital human delivering your message. The core problem it solves is that video content takes time and money to produce at scale—hiring talent, booking studios, managing post-production. D-ID collapses that into minutes and a API call. Pricing starts free (limited credits monthly) with paid tiers around $10–100/month depending on video minutes and API volume; enterprise pricing available on request. The honest limitation: avatars work best for straightforward messaging and explainers, not narrative performance or high emotional nuance.
PaidFree Trial · 14 days$4.7/moAPIVerified Apr 7, 2026
18. Descript
The core idea: transcribe the recording, edit the transcript, and Descript makes the matching cuts in the timeline automatically. The AI layer — Descript calls it Underlord — goes further, offering to remove filler words in bulk, generate show notes, recut long-form content into social clips, and apply scene design without manual timeline work. That pipeline holds well for solo creators and small teams producing one or two videos a week. The ceiling appears when output volume scales or when a project needs frame-level precision editing — at that point, editors reach for a traditional NLE alongside Descript, not instead of it.
Paid$16/moAPIVerified Jun 1, 2026
19. Frontier AI for Motion Design
The vendor describes a five-stage pipeline — brief, brand research, storyboard, build, iterate — run by one agent without switching tools. Motion reads your site to extract real colors, type, and style references before it storyboards anything, which means the output starts from your actual brand rather than a generic template. The MCP and API surface lets other agents — Claude, Cursor, ChatGPT — call Motion directly and receive a rendered video back. There is no self-hosted option and no free tier, so teams that need on-premise deployment or want to prototype before committing are blocked at the door. The studio service exists for launches where the agent output alone is not enough, but that path involves booking a call — it is not self-serve.
PaidAPIVerified Jun 23, 2026
20. Haiper AI Video
Haiper handles text-to-video, image-to-video, and clip extension in a single browser-based workflow — no installation, no model configuration. You submit a prompt or upload a source image, pick a duration, and receive a short generated clip. The credit-based model means free-tier usage runs out faster than it looks on a project with multiple revisions. Teams doing more than a handful of generations per week hit the credit ceiling and move to paid usage or reconsider the economics against per-seat competitors. API access exists, so embedding generation into a pipeline is possible — the vendor states this, though production rate limits and SLA details require checking documentation directly.
PaidAPIVerified Jun 22, 2026
21. HeyGen
HeyGen addresses a real friction point: creating video content at scale without the logistics of hiring talent or renting studios. You write a script, pick an avatar (or upload your own), select a voice, and the tool generates a finished video in minutes. The core pitch is speed and repeatability for marketing teams, HR onboarding, and e-learning shops. Free tier covers basic exports; paid plans start around $25/month and unlock premium avatars, higher quality, and batch processing. The honest catch is that output still reads as synthetic—useful for internal comms or explainer videos, less so if you need to convince skeptics that a real human endorses your product.
Paid$29/moAPIVerified Dec 1, 2025
22. HeyGen Avatar 5
The core workflow is script-in, video-out: paste a script or upload a PDF, pick an avatar, and the platform generates a 1080p or 4K video with lip-synced narration and auto-subtitles. Translation into 175+ languages runs through the same pipeline, which means a training video recorded once can ship localized without re-recording. The ceiling appears when you need precise editorial control — avatar gestures, pacing, or emotional beats beyond what the text-based editor exposes. Teams doing high-volume, tightly branded content typically find themselves exporting and finishing in a dedicated editor. For output that depends on a human face behaving exactly right on camera, the gap between generated and filmed is still noticeable.
Paid$29/moAPIVerified Jun 1, 2026
23. Kling
Kling AI generates video from text prompts and images, with a documented focus on photorealistic human motion and native 4K output rather than upscaled resolution. Built-in audio synthesis and lip-sync are included, which removes the external toolchain that most comparable generators require. The free tier provides 66 daily credits — enough for experimentation and low-volume testing. The wall appears when you push toward high-volume batch output or need fine-grained control over scene composition across a multi-shot sequence; the one-shot generation model does not chain shots autonomously. Teams running high-volume e-commerce catalogs typically schedule generation in batches and manage sequencing outside the tool.
Paid$6.99–$159.99/monthAPIVerified Jun 1, 2026
24. LTX Studio
The platform covers the full arc from script upload to timeline edit inside a single workspace — storyboard generation, text-to-video, image-to-video, camera control with keyframes, and sound design are all connected rather than siloed. The vendor states that AI Characters, Objects, and Locations persist as named elements across scenes, which is where most competing tools quietly fail. The camera control and keyframe tools give directors shot-level precision without dropping into a code environment. The ceiling appears when you need fine-grained post-production compositing or when brand audio requirements exceed what the built-in sound design layer can handle — teams at that stage are exporting to dedicated editing pipelines.
Paid$12-$100/moAPISelf-hostedVerified Jun 9, 2026
25. motionvid.ai
Motionvid lets you submit a text prompt or reference image and receive a rendered motion graphics output — YouTube intros, branded explainers, animated infographics, TikTok clips — without touching a keyframe. The workflow is one-shot generation with optional text-based refinement, so iteration means re-prompting, not scrubbing a timeline. That speed is real for standard formats. The ceiling appears when output needs frame-precise control, custom character rigs, or motion that diverges from what the model was trained to produce. Teams with those requirements end up exporting and finishing in a traditional editor, which partially defeats the time savings.
Paid$9/monthAPIVerified Jun 1, 2026
26. Opus Clip
OpusClip takes a long-form video URL or upload, runs it through a scoring model that identifies high-engagement moments, and returns ranked short clips ready for TikTok, Reels, or Shorts — without an editor in the loop. The vendor states the model evaluates hooks, speaker energy, and topic coherence to rank clips automatically. That works well for talking-head content: interviews, podcasts, webinars. It starts to slip on footage that depends on visual context the model doesn't read — sports highlights with complex action, heavily edited narrative video, or anything where the audio alone doesn't carry the moment. Teams hitting that ceiling typically add a manual review pass or offload to a dedicated video editor for those asset types.
PaidFree Trial · 7 days$15/moAPIVerified Jun 3, 2026
27. Pictory
Pictory takes a URL, script, or long-form article and converts it into a video by matching your text to stock footage, adding captions, and assembling a timeline — no editing software required. The workflow is fast for standard marketing clips and social cuts. Where it strains is in creative control: the stock footage matching is automated, which means the tool picks the visual, not you, and correction rounds add up quickly. Teams producing one-off brand videos find the output acceptable at speed; teams with strict visual identity standards spend significant time overriding selections. When the asset library and auto-matching stop fitting the brief, teams move to a dedicated editor or a custom motion graphics workflow.
PaidFree Trial · 14 days$25/moAPIVerified Jun 5, 2026
28. Pika
Pika sits in the crowded space of generative video tools, competing with Runway and OpenAI's Sora by offering faster inference and a focus on ease of use over photorealism. You describe what you want in text or upload an image, and it outputs a video clip—useful for social content, product demos, or storyboarding. The free tier lets you generate a handful of videos monthly; paid plans start around $10/month for creators needing batch exports and longer clips. The biggest friction: video quality remains noticeably synthetic, and render times can stretch depending on server load, making it less suitable for deadline-critical work.
Paid$8/monthAPIVerified Oct 1, 2023
29. Reeloop
The tool takes a single line of text — or a URL, or a pasted script — and produces a short-form video with a Claude-written script, Seedance-rendered cinematic scenes, an AI voiceover synced to the script, and word-by-word captions. Six style presets (Viral Story, Educational, Documentary, ASMR, Italian Brainrot, UGC Ad) shape tone and pacing without manual configuration. The vendor states average generation time is around four minutes. Auto-posting to TikTok and YouTube Shorts is available as a scheduled feature, so a channel can publish on cadence without manual uploads. The ceiling appears when you need footage that diverges from what Seedance generates — there is no manual scene override described in the docs.
Paid$23/moAPIVerified Jun 22, 2026
30. Runway
Runway lets you generate, edit, and transform video and images using AI without touching code—think Photoshop meets a generative model API. The core problem it solves: professional-grade AI video editing takes weeks of learning or hiring engineers. You get access to models for background removal, motion synthesis, upscaling, and text-to-video generation. The free tier covers basic monthly credits, but real work requires a paid plan starting around $12–$28/month depending on resolution and model access. The honest friction: the free tier shrinks fast, and output quality still lags human-made footage for broadcast work.
Paid$12/monthAPIVerified Oct 1, 2023
31. Seedance 2.0 Mini Generator
The tool takes a single text prompt or up to nine reference images and produces a multi-shot video up to 15 seconds long, with built-in camera movement controls — zooms, pans, tracking shots — described in plain language rather than keyframe editors. Resolution goes to 1080p; aspect ratios cover 16:9, 9:16, and 1:1, so the same prompt can output for YouTube, TikTok, and Instagram without re-shooting. Where it breaks: duration caps at 15 seconds, which rules out anything longer than a punchy ad cut. Teams needing minute-long narrative pieces hit that wall fast and have no native workaround inside the platform.
Paid$39.9/moAPIVerified Jun 18, 2026
32. seedancee2.ai
The core loop is blunt and fast: write a prompt with camera direction and mood, generate a clip, tune duration and format, export. The vendor states outputs reach 4K at 1920x1080, and community examples on the showcase page support that claim without obvious post-production polish. Character Lock — the ability to hold a subject consistent across shots — is the feature that separates this from one-shot generators when you need to build a scene sequence rather than a single clip. The ceiling appears when a project demands shot-to-shot editorial precision that a prompt cannot fully specify; fine-grained control over timing, cut points, or dialogue sync still requires an editor downstream. For ad variant production and pre-visualization, the speed arithmetic works — for anything requiring locked timing against audio, it doesn't.
Paid$24.9/mo - $83.25/moAPIVerified Jun 9, 2026
33. Spatius
Spotter is a point-and-shoot identification app: you photograph a landmark, street food, animal, or foreign-language sign, and the app returns an AI-generated synopsis plus a chat thread anchored to that specific subject. Each identification saves as a 'Spot,' accumulating into a personal travel journal. The free tier caps snaps sharply, so teams building travel or education products on top of this API hit the credit ceiling fast during any meaningful test cycle. There is no self-hosted option, which means all image data routes through Spatius infrastructure — a deal-breaker for enterprise deployments where data residency matters.
Paid$19/moAPIVerified Jun 1, 2026
34. SwiftThumbnail
SwiftThumbnail takes a YouTube link or style input and generates thumbnail variants you can download or edit manually — no design canvas to learn, no export settings to configure. The core workflow is single-shot: input in, image out. That speed holds for solo creators and agencies running high weekly output. The ceiling appears when a project demands fine-grained layout control or brand consistency across dozens of assets — at that point, the one-shot model leaves you cycling through generations rather than directing them. Teams with strict brand guidelines end up supplementing with a dedicated design tool.
Paid$15/moAPIVerified Jun 7, 2026
35. Synthesia
The core workflow is script-in, video-out: you write or paste text, select an avatar and language, and the platform renders a presenter-led video. This holds up well at volume — L&D teams producing dozens of compliance or onboarding modules report genuine throughput gains over traditional recording. The ceiling appears when you need emotional range, off-script spontaneity, or branded visuals that go beyond slide-style backgrounds. Avatar consistency across a long series is solid; voice consistency across sessions is less so, and for customer-facing content where callers hear the same agent repeatedly, that gap registers. Teams needing custom avatar likeness or advanced brand control hit a paid-only gate.
Paid$14/moAPIVerified Jun 7, 2026
36. Tavus
Tavus lets developers deploy conversational video agents—digital replicas that see, hear, and respond with emotional nuance—without building a video stack from scratch. The core problem it solves is latency: most video AI feels choppy or requires heavy post-production. Tavus delivers near-synchronous interaction through proprietary rendering, critical for sales calls or live support where lag breaks trust. Pricing starts at the API tier but exact costs aren't published upfront, requiring a direct conversation with sales. The main friction: this isn't a no-code tool. You need engineering resources to integrate the API and train custom replicas.
Paid$59/moAPIVerified Apr 7, 2026
37. Veed.io
The platform handles the full production chain in a browser: text-to-video generation, AI avatar talking heads, automatic subtitles, dubbing into other languages, background removal, and noise reduction — none of which require a local install or handoff to a separate tool. Brand kits let you lock colours, fonts, and logos so a junior marketer and a senior designer produce outputs that are visually indistinguishable. The subtitle engine is frequently cited in community feedback as the strongest single feature. The wall appears when you need frame-level precision — VEED is not a replacement for a timeline editor, and high-volume programmatic generation routes through the API, which is a paid-only feature.
Paid$12-$39/moAPIVerified Jun 9, 2026
38. ViMax
The framework orchestrates four autonomous agents — Director, Screenwriter, Producer, and Video Generator — that take a text input and carry it through scripting, scene planning, and clip generation without you manually handing off between steps. The agents call external APIs under the hood: Google Veo for video output, Nanobana for image generation, and your LLM provider of choice for script and direction logic. That architecture means the framework code itself costs nothing, but every scene rendered incurs API charges from those third-party services. Narrative-coherent multi-scene output — the problem the tool exists to solve — is what you get when the pipeline runs cleanly. Where teams hit friction is in the dependency chain: configuration across multiple API keys, rate limits from external providers, and limited community support for edge-case pipeline failures.
FreeOpen SourceAPISelf-hostedVerified Jun 9, 2026
39. ViralMint
ViralMint is an open-source pipeline that chains scout, download, clip, and generate into a single workflow ending in a finished mp4. The outlier detection compares each video against its own channel baseline rather than a global average, so a 3× spike on a small channel surfaces next to a 20× monster on a large one — and you decide which matters. The Clip Studio extracts 30–60 second moments from long-form video; the Smart Video pipeline assembles originals from a text idea using AI script, Pexels stock, voiceover, and captions. The 58 MCP tools let Claude Code run the full pipeline hands-off. The wall appears when you need direct publishing to platforms — ViralMint produces the mp4 and stops there.
PaidOpen SourceFree open-source + pay-as-you-go AI usageAPISelf-hostedVerified Jun 9, 2026
40. Vivijure
Vivijure is a self-hosted module host for AI film production, built on a thin Cloudflare Workers core that runs on the free tier and routes tasks to whatever GPU backend you wire up — your own box, RunPod, or a cloud motion API. The typed hook contract between modules is what keeps the pipeline from becoming a tangle of bash scripts: each step — storyboard rendering, TTS narration, lip-sync, music bed — is a discrete worker you attach or replace without touching the rest. The AGPL-3.0 license means the source is auditable and forkable, but it also means any hosted derivative you ship has to stay open. The repo shows 31 open issues and 8 open pull requests against 278 commits — active, but not mature.
FreeOpen SourceAPISelf-hostedVerified Jun 23, 2026
41. Vmake AI
Vmake is a cloud-only video and image enhancement platform built for sellers, creators, and agencies who need polished output without a post-production pipeline. The core workflow is one-shot: upload a video, select an enhancement task — upscaling, background removal, watermark cleanup, avatar generation — and receive processed output. Batch processing handles volume jobs without manual queuing. The free tier provides a credit pool sufficient for light experimentation, but production-volume workflows hit the credit ceiling fast. Teams running daily content schedules will exhaust free credits within hours and need to account for that in their tooling budget from the start.
PaidFree tier + $10–$30/monthAPIVerified Jun 1, 2026
42. wavreel
The pipeline is linear and intentional: upload an MP3, WAV, or M4A narration, and Whisper transcribes it with timestamps. Every three seconds of audio gets a visual description, which drives a stock image query against Pexels. You review scenes in a browser editor, swap any image that misses, then render. The vendor states the median session runs around 15 minutes. That speed holds for narration-driven Shorts; the ceiling shows when a project needs original footage, licensed music, or anything the Pexels catalog cannot cover.
Paid$19/moAPIVerified Jun 19, 2026
Listings on this page are sourced and verified by the AIDiveForge data pipeline. AIDiveForge is editorially independent — no money changes hands for inclusion.