Everything is included, on every plan
Not all localization is equal. Plans differ only by how many minutes you get — every capability below is available on all of them.
Localization
Full AI localization
Upload once and the agent takes it end to end.
Voice & background separation
Splits the audio into clean dialogue and a separate music/effects track, so the original score survives the dub.
Speech to text
Full sentences with punctuation, ready for subtitles and closed captions — in 84 source languages across 170 regional variants.
Transcreation
More than translation — each line is adapted for context, register, and the length of the original delivery, so it fits the source audio utterance by utterance.
Phrase timing
Pauses and delivery pacing matched to the timing of the source performance.
Voice matching
Keeps the speaker's vocal character while sounding native in the target language.
Emotional tone
Carries the emotional delivery of each utterance across into the target language.
Gender detection
Our own model detects each speaker's vocal gender for every segment, so an appropriate voice is selected automatically rather than guessed.
Matched through a performance
Every line is matched to the voice as it sounded in that moment. A speaker who drops to a whisper or slips into falsetto is matched the way they actually sounded — not flattened into one voice for the whole video.
Subtitles
Automatic subtitle generation in both the source and target languages.
Multiple target languages
Localize a single upload into as many languages as you need — 173 target locales across 81 languages, with 2,782 voices.
Editor
The AI works for you, not the other way around
Everything Addavox produces is a starting point you can change.
Hear every layer separately
A three-track mixer puts the background, the original voice, and the dub under independent control. Solo any one of them, or balance all three.
Shape any line
Every segment has its own EQ, gain, and pan. Adjust one line without touching the rest, then copy that treatment to the whole project.
Retime by hand
Drag a waveform to move it. Pull an edge to trim it. Split a line at the playhead. Re-transcribe a passage by dragging across it.
Ask instead of hunting
Tell the assistant what you want. “Retell this joke so it works culturally in French.” “This is too long, make it shorter.” “Rename this term everywhere.” Project-wide changes are previewed for your confirmation before anything is applied.
Check the meaning
Back-translate any line to see what it actually says in your own language. Paraphrase it shorter or longer without retranslating from scratch.
Direct the performance
Record the target-language read yourself, in the browser, then have it re-voiced in the original speaker's matched voice. Blend up to four reference segments for a steadier match.
Subtitles that match
Set characters-per-line for source or target. What you preview is what the exported SRT and VTT contain.
Work the way you like
Every keyboard shortcut and touch gesture is documented beside the editor. It is genuinely usable on a tablet.
A complete localization editor
- Edit transcription and translation side by side
- Merge and split segments in one click
- Adjust timing — move, compress, expand, and trim audio to fit the source
- Add and edit pauses between segments
- Regenerate voice-matched audio for any segment
- Normalize delivery across a speaker to even out emotional tone
- Direct voice talent: record the target-language performance in the editor, then re-voice it in the original speaker's matched voice
- Blend up to four reference segments for a steadier voice match
- Shorten or elaborate a line to fit the timing, without retranslating
- Back-translate any line to check the meaning in your own language
- Full audio mixer — 3-band EQ, high-pass filter, gain, pan, and bypass per utterance
- Built-in reference for every keyboard shortcut and touch gesture, beside the editor
AI agent
An AI localization agent
Ask for changes in plain language instead of hunting through menus.
- Fix a transcription or translation by describing the problem
- Retranslate in a more formal or more casual register
- Adapt a joke or idiom so it lands culturally
- Shorten a line that runs past the source timing
- Correct gender and pronoun choices
- Jump to any segment by what it says
- Make project-wide changes — rename a term everywhere, restyle a batch of segments — previewed for your confirmation before anything is applied
Collaboration
Bring in a native speaker without the setup cost
No accounts. No licenses. No onboarding call.
Invite by email
A reviewer gets a link, not an account. They open it and start working.
Scoped to one language
A reviewer invited for French sees French. That boundary is enforced on our servers, not hidden in the interface — so a reviewer cannot reach a language they were not invited to.
Work at the same time
Several reviewers can work in one project simultaneously. You can see who is online and what they are working on as they do it.
Changes take effect immediately
When a reviewer edits a line, that line's audio is regenerated automatically. No re-run, no re-export.
You stay in control
Links expire in seven days and can be revoked instantly. Every change records who made it, what it was, and what it became.
Voice
2,782 voices. Or the one you already have.
Keep the speaker your audience already knows, in a language they do not.
Their voice, in 173 locales
Carry a speaker's vocal character into 173 target locales across 81 languages, from about five seconds of clean speech.
Matched moment by moment
People do not speak in one register. They drop to a whisper, jump to falsetto, lean into a line. Addavox matches each segment to how the speaker actually sounded at that moment, instead of averaging them into a single voice for the whole video.
Emotion carries
The emotional delivery of each line carries into the target language, not just the words.
Or pick from the catalogue
2,782 preset voices for when you do not need — or do not have consent for — the original speaker's voice.
Steadier when you want it
Blend up to four additional reference segments to build a more consistent match for a difficult line.
Direct it yourself
Record the target-language performance in the editor, then have it re-voiced in the original speaker's matched voice.
Trust
Built for localization. Not for impersonation.
Voice AI earns trust by what it refuses to do.
A voice cannot be regenerated in its own language
Addavox carries a voice into a language the speaker is not speaking. It will not regenerate that voice in its native language. That single restriction is what makes the technology useful for dubbing and unusable for impersonation.
Consent is a gate, not a checkbox
Voice-matched localization cannot begin until someone attests, by name, that they have the speaker's consent and the rights to the content. That attestation is timestamped, has to be current, and is recorded with the project. Standard synthetic-voice localization processes no speaker's voice at all, and requires no consent.
Voice data is never stored
Voice matching happens in memory while the job runs. The voice profile is not kept afterward. There is no voice library building up behind your projects.
If anything goes wrong, you are refunded
Automatically, on every failure path — a job that fails, a file over the limit, a cancellation. You are never charged for work you did not receive, and you do not have to ask.
Every change is attributable
Every edit, yours or a reviewer's, records who made it, what it was before, and what it became.
Reviewers see only their language
Enforced on our servers. Links expire in seven days and can be revoked instantly.