Cutting to a voiceover means scrubbing for words. VoxMark ends that. Point it at your comp or sequence — not at a file you exported — and it renders the real audio mix, transcribes it with the AI provider of your choice — in any of 99 languages, auto-detected — and drops the result onto the timeline as markers or subtitles. Multiple speakers come back named and colour-coded. The transcript then lives in the panel as a clickable navigator: click the sentence, the playhead is there.
Most transcription tools want a wav. VoxMark wants your composition. Scan Comp lists every audio layer it can find — including the ones buried in nested precomps — and Render & Transcribe renders the actual mix you would hear, music bed and all, then sends that off. Nothing to export, nothing to line back up afterwards, and the timings stay honest because they came from the timeline in the first place.
VoxMark transcribes with Whisper, which was trained on 99 languages, and it asks nothing of you to use them. Leave the dropdown on Auto-detect and the language is identified from the audio itself — a Japanese voiceover, an Arabic interview, a Spanish read under an English music bed — and the markers arrive on the timeline in that script.
Groq and OpenAI both run Whisper, so the range is identical on either — Groq is simply faster, and free. AssemblyAI brings its own multilingual model, and auto-detects too. And a straight word on accuracy: Whisper is at its best on the widely spoken languages and softens on the smaller ones. That is the model, not the tool — every product built on Whisper shares the same ceiling. What differs is what happens to the words afterwards: here they become frame-accurate, colour-coded markers you can click.
Turn on speaker identification and a two-hander comes back as two colours — in the panel and on the markers themselves, using After Effects' own label colours. Rename "Speaker A" to the person who actually said it and every marker repaints live, in a single undo.
Once a comp is transcribed, the Navigate tab is how you get around it. Click any caption and the playhead jumps to it. Scrub the timeline and the current line highlights as you pass it — in Premiere it follows live playback too. The transport row steps marker to marker, so you stop dragging along the timeline ruler hunting for the start of a sentence.
One dropdown decides how finely the transcript is cut, and it is worth setting deliberately — the right granularity is the difference between markers you cut against and markers you animate to.
Either way, the language is auto-detected unless you set it, and you can clear the existing markers first if you want a clean slate rather than a second set layered over the first.
A full transcript on the timeline is a lot of text. Two toggles calm it down without losing anything: hide the caption text so the markers stay but the words live only in the panel, or collapse ranged markers to single frames. Both are reversible, and neither touches the transcript.
Export a proper SRT for subtitles and captions anywhere. VoxMark's carries two things a standard SRT does not: the exact frame rate of the project it came from, and the voice each line belongs to. That is what makes a round trip survive.
There is no VoxMark subscription and no middleman server. The panel talks straight to the transcription service you choose, using a key you hold, stored locally. Three are supported, and the panel links you to each one's key page.
The same panel loads in both hosts and relabels itself to suit — Comp in After Effects, Sequence in Premiere. Scan the active sequence, render its audio mix, drop the markers on the timeline; in Premiere the transcript also follows live playback. You are not buying the same tool twice.
Your download is a zip holding the panel — VoxMark_v2.0.4.zxp — and a short INSTALL.txt. Three steps, about two minutes:
One thing After Effects makes us ask for: VoxMark renders through its own VoxMark 16k Mono output template, and After Effects does not let a plugin add one by itself. The panel spots this on first run and walks you through loading it — five clicks, once, and it never comes up again. Premiere needs nothing.
Updates come to you. When a newer release is out an ↑ Update button appears in the panel's header; click it and the new .zxp lands in your Downloads folder, ready to drop onto your ZXP installer. Perpetual licence, free updates — version 2.0.4 today, and every version after it.
Not for you? There is a 30-day money-back guarantee — ask and you get a refund.
After Effects 2022 to 2026 and Premiere Pro 2022 to 2026, on Windows and macOS. One licence covers both applications.
No. VoxMark is a one-off purchase and there is no VoxMark service to subscribe to. Transcription runs on a provider you choose, with an API key you hold — and Groq's free tier is enough for most work, so for many people the running cost is nothing.
Start with Groq — free, and the fastest of the three on the same Whisper model. Use OpenAI if you already have a key and would rather keep everything in one account; it bills per minute of audio. Use AssemblyAI when you want speakers identified automatically, or when the file is long — it gives five hours free. The panel links you straight to each provider's key page.
It is stored locally on your machine and the panel talks directly to the provider. There is no VoxMark server in between.
Your actual comp. Scan Comp finds the audio layers — including ones inside nested precomps — and Render & Transcribe renders the real mix, music bed and all, then sends that. Nothing to export by hand, and nothing to re-sync afterwards, because the timings come from the timeline itself.
Both, from one dropdown. Segments give you phrase markers, which is what you want for cutting and for subtitles. Words give you a marker per word, which is what kinetic typography needs. Long lines can also be split at a character limit you set.
Yes, with AssemblyAI selected and "Identify speakers" ticked — each voice comes back as its own colour, on the markers themselves as well as in the panel. Rename "Speaker A" to a real name and every marker repaints in one undo. On Groq or OpenAI, where there is no diarization, you can select captions in the panel and assign voices by hand; the result is identical.
Yes, two ways, both reversible. Aa hides the caption text so the markers remain but the words live only in the panel, and ◆ collapses ranged markers to single-frame ones. Neither affects the transcript.
Save & Clear parks a comp's markers in memory and empties the timeline; Restore puts them back, with the count shown on the button. Each comp keeps its own memory, and that memory survives renaming the comp — so you cannot restore one comp's markers into another by accident.
It is a standard .srt and works anywhere one does. It also carries the exact frame rate of the project it came from, and the voice each line belongs to — neither of which a plain SRT records. That is what lets a transcript move between applications without the timings drifting.
Yes. Import SRT builds markers from any subtitle file with no transcription, no API key and no cost. The imported markers go into marker memory too, so Restore works on them.
VoxMark for Cinema 4D reads VoxMark's SRT natively, frame rate and voices included — so you can transcribe in After Effects and animate to the same markers, in the same colours, in C4D.
Ninety-nine. VoxMark transcribes with Whisper, which covers 99 languages. Leave the language on Auto-detect and it is identified from the audio on every provider, or pick any of the 99 from the dropdown — the fourteen most-used are pinned at the top. Markers and the exported SRT are Unicode throughout, so Japanese, Arabic, Hindi or Russian land on the timeline in their own script.
Your download is a zip holding VoxMark_v2.0.4.zxp and a short INSTALL.txt. Drag the ZXP onto a ZXP installer — the free ZXP/UXP Installer does it in one drop. The certificate is self-signed, so there is a one-time security prompt to allow. Restart After Effects or Premiere Pro, open Window ▸ Extensions ▸ VoxMark, and paste the licence key from your Gumroad receipt into the screen that appears, then click Activate.
There is a 30-day money-back guarantee. Buy it, try it on real work, and if it does not fit the way you cut, ask for a refund within 30 days and you get one.
They are free, for as long as VoxMark exists — it is a perpetual licence, not a subscription. The panel watches for new releases itself: when one lands, an ↑ Update button appears in its header, and clicking it drops the new .zxp into your Downloads folder for you to install over the old one. Nothing to check for, no upgrade fee.
Yes. The key is not tied to a particular computer, so your desktop and your laptop are both fine. It is a licence for one artist rather than a whole studio — but that one licence covers After Effects and Premiere Pro, so you are never buying the same tool twice.
The Chroma Discord — that is where VoxMark is supported, and where you will get an answer quickest. The panel also keeps a log you can open from its own menu, which usually says exactly what a provider objected to.
One click turns your voiceover into frame-accurate, colour-coded markers — in After Effects and Premiere Pro, on the AI provider of your choice, including a free one.
Perpetual licence · free updates · 30-day money-back guarantee