🎉 Early access is open — every signup gets 5 free minutes of dubbing. No credit card required.
中文 CapabilitiesHow it worksPricingCompareFAQ
AI video translation · dubbing · subtitles

Put your video in
10 languages.

Paste a link or drop a file. Minutes later you have a dubbed, subtitled master in every language you publish — in your own cloned voice. Like hiring a localization studio, a hundred times faster.

★ The part nobody else does: fix one line, re-render only that line see it below ↓

Credits never expire · Your footage auto-deleted in 7 days · Commercial license · No credit card

Live UI demo · switch languages, click any line, edit text, re-record just that line (real dubbed samples for all 10 languages in production — they'll replace this soon)
app.voxtakes.com / projects / zelos-x9-launch · Spanish

Line review · 46 lines

QC via re-transcription
Mixdown unlocks when all lines pass · this job used 12.4 / 60 min
Free tier exports carry a corner watermark · removed on any paid plan
Why we exist

Black-box dubbing makes you pay for their mistakes.

Today's tools treat a finished video as a one-shot generation. But real work always has that one line that's off — the question was never whether you can edit it, but what it costs to do so.

✕

One word fix = full re-render

Regenerate the entire job. On per-minute tools that means paying for 20 minutes of credits because of one mispronounced brand name — plus another render wait.

Voxtakes: only the edited line is re-recorded, and QC re-dos are on us
✕

Subtitles that break mid-sentence

ASR chops one sentence into two fragments, viewers read half-thoughts, and names get misspelled. Nobody fixes it because the timeline is locked in.

Voxtakes: full-sentence segmentation + per-line re-transcription QC + locked glossary
✕

Credits that expire at midnight

Monthly credit pools reset whether you used them or not. Creators publish in bursts — flat monthly plans punish exactly that rhythm.

Voxtakes: credits never expire, top-ups valid for 12 months
Capabilities

Line-level control is just the start — every layer of the master is engineered.

Voice, multiple speakers, subtitle timing — the three places dubbing goes wrong. Each one has real engineering behind it.

Voice cloning

Still your voice —
just speaking another language.

15 seconds of clean audio is enough to clone timbre, pitch range and speaking habits across languages. Long passages render block-by-block instead of line-by-line, so your voice never "drifts" between sentences. Viewers should think "you speak Spanish" — not "someone dubbed this."

Source→Español
Multi-speaker

A two-person interview should not
become one voice.

Voiceprint clustering detects every speaker in interviews, podcasts and dramas, clones each one separately, and keeps identities bound through the whole video. Host is host, guest is guest. This is where most dubbing tools fail hardest — we made it the default.

Host · ASo, what changed since the last prototype?Host voice, cloned separately
Guest · BCost. And battery life doubled.Guest voice, kept throughout
Host · ATwo voices, two clones — zero crosstalk.Speaker identity never swaps mid-video
Subtitle sync

The caption leaves
when the sentence ends.

Fragments are stitched into complete sentences before translation; each caption's on-screen time follows the real dubbed audio length, with auto-wrapping that never overflows or overlaps. Bilingual layout: translation large and high, source small and low — viewers never piece half-sentences together.

✕Typical tools: cut mid-sentence
It can walk over rocky terrain for thi…
raw ASR fragments on screen
✓Voxtakes: full sentence, real audio length
It can walk over rocky terrain for thirty straight minutes.
它能在乱石地形上连续行走三十分钟。
Also in the box
Pull from URL — no upload Resume from checkpoint — never re-render what passed Glossary lock for brand names SRT / VTT / ASS export Title & description drafts with every master Bulk queue · API (beta) Priority render lane
How it works

Three steps. Ten languages.

STEP 01

Paste a link, or drop a file

YouTube links are pulled server-side at full source quality — nothing to download or re-upload. Local files go in over resumable chunked upload; close the tab, it still finishes.

10-min video ≈ 15-min turnaround
STEP 02

Review line by line. Fix what's off.

Translation and cloned voiceover arrive line by line. Play a line, edit the text, roll a new take — each change re-records that line only. Everything else is untouched.

Edit one line, pay for one line
STEP 03

Ship the whole package

Burned-in bilingual subtitles + SRT/VTT/ASS files + localized title & description drafts. Run all target languages in parallel from one approved script.

1 video × N languages
Pricing

Simple math. Credits that stay yours.

Metered in finished-video minutes (rounded up). Edits are charged for changed lines only. Every plan's credits roll over forever.

Monthly
AnnualSave 20%

Starter

Weekly creators
$29 / mo
60 dubbed minutes · ≈ $0.48/min
  • 10 target languages, parallel outputs
  • Line-by-line review & per-line billing
  • URL pull / resumable chunked upload
  • Burned bilingual subs + SRT/VTT export
  • Credits never expire
Get started
Most popular

Creator

Full-time YouTubers & podcasters
$79 / mo
200 dubbed minutes · ≈ $0.40/min
  • Everything in Starter
  • Voice cloning — keep your timbre & delivery
  • Multi-speaker separation
  • Glossary & brand-name lock
  • Priority render queue
  • 1080p masters · no watermark
Start free — 5 min on us

Studio

Agencies & localization teams
$199 / mo
600 dubbed minutes · ≈ $0.33/min
  • Everything in Creator
  • 3 seats · shared workspace & approvals
  • Client review links
  • API & bulk queues (beta)
  • Usage reports · configurable retention
Talk to us

Overage $0.89 / finished minute (list price elsewhere: $2.20–$3.00) · failed-QC re-dos never billed · USD pricing, applicable tax collected by our merchant of record

Compare

Run the numbers against the big names.

Real scenario: a 20-minute master that needs 2 line fixes after review.

VoxtakesRaskElevenLabsSubtitle-only tools
First 20 min $9–16($0.40–0.48/min in-plan) ≈$60–150($2.4–3.0/min) ≈$44(Dubbing v2, $2.20/min) ≈$4(subtitles, no voiceover)
Cost of fixing 2 lines 2 lines only · ≈$0.04per-line billing, no re-render Full re-render in web UI ≈$48segment redubbing is API-only Studio supports per-clip regenerationproduct is now in maintenance mode Back to the NLE to fix timing by hand
Line-by-line review flow ✓ QC badges + approval state + client review links ✕ No QC loop ◐ Editor exists, no QC ✕
Subtitle segmentation ✓ Full sentences, real audio timing ◐ Raw ASR chunks ◐ Raw ASR chunks ◐ Frequent mid-sentence cuts
Unused credits Never expire (top-ups valid 12 months) Reset monthly Rollover capped at 3× —
Pull from URL, no upload ✓ And not metered ✓ ◐ Some entry points ◐

Figures from vendors' published 2026 pricing pages; products evolve — tell us if anything moved and we'll update.

FAQ

Questions people actually ask.

Real masters for all 10 target languages are in production right now and will sit in the demo area as soon as they land. Want one on your footage today — book the 15-minute live demo and we'll run a line while you watch.

Our masters are assembled from per-line audio tracks on a timeline. Editing a line's text or voice re-generates that line's audio only and drops it back in place — everything else is reused, so billing is that line's duration. It also means no 20-minute re-render for a two-word fix.

By finished-video minutes, rounded up. One video dubbed into two languages = two units. Edits bill changed lines only; re-dos caused by failed QC are free. Every signup includes 5 free minutes — no card.

Source files and masters are auto-deleted after 7 days by default (configurable shorter or longer), encrypted at rest, and never used to train models. GDPR-compliant processing; DPA available on request.

We launched with 10 languages tuned and natively evaluated — English, Spanish, German, French, Portuguese, Italian, Japanese, Korean, Mandarin, plus Arabic in beta — instead of advertising "135 languages" with three usable ones. Live per-pair quality ratings are published on our status page.

15+ seconds of clean speech is enough. Timbre, pitch range and delivery carry across languages. For long-form coherence (audiobooks, narration-heavy videos) we render in blocks rather than isolated lines, so the voice doesn't drift between sentences.

Yes — all paid plans include a commercial license and you own the output. You're responsible for rights in the content you submit; we run content moderation and a DMCA process. When localizing work you don't own, get permission and credit the source.

Currently 60 minutes / 2 GB per job; bulk work goes through Studio's queue. A public API (REST + webhook callbacks) is in beta — suited to automation stacks and agency pipelines.

Your next video speaks ten languages.

5 free minutes on signup. No card, nothing to install.