More VoxMark on YouTube ↗

Voiceover to markers for After Effects & Premiere Pro

Your voiceover becomes frame-accurate markers. One click.

Cutting to a voiceover means scrubbing for words. VoxMark ends that. Point it at your comp or sequence — not at a file you exported — and it renders the real audio mix, transcribes it with the AI provider of your choice — in any of 99 languages, auto-detected — and drops the result onto the timeline as markers or subtitles. Multiple speakers come back named and colour-coded. The transcript then lives in the panel as a clickable navigator: click the sentence, the playhead is there.

After EffectsPremiere ProGroqOpenAIAssemblyAI99 languagesSRT in & outCinema 4D
£39.99
One licence for After Effects and Premiere Pro · Windows and macOS · perpetual, with free updates · pay in three instalments if you would rather · 30-day money-back guarantee
5 1 · version 2.0.4 · After Effects & Premiere Pro 2022–2026
Before / after

The old way, and the VoxMark way.


Without VoxMark
  • Export the VO, upload it somewhere, wait, download a transcript
  • Read the transcript in one window and scrub the timeline in another
  • Type marker comments by hand, one line at a time
  • Nudge every marker because the export didn't line up with the comp
  • Work out who said what, then colour the markers yourself — if you bother
  • Hand off an SRT and hope the frame rate matches on the other end
With VoxMark
  • Scan the comp, hit Render & Transcribe — it renders the real mix itself
  • Markers land on the timeline already timed and already worded
  • Words or phrases, your choice, with a max-length split
  • Speakers identified, named and colour-coded automatically
  • Click a line in the panel and the playhead goes there
  • Export an SRT that carries your exact frame rate
Whole-mix transcription

Point it at the comp, not at a file.


Most transcription tools want a wav. VoxMark wants your composition. Scan Comp lists every audio layer it can find — including the ones buried in nested precomps — and Render & Transcribe renders the actual mix you would hear, music bed and all, then sends that off. Nothing to export, nothing to line back up afterwards, and the timings stay honest because they came from the timeline in the first place.

  • Scans the comp or sequence for audio, nested precomps included
  • Renders the real mix — VO under music, layered tracks, all of it
  • Renders at 16 kHz mono, which is what the speech models listen to anyway — so a ninety-minute sequence uploads in a fraction of the time, with nothing lost
  • Auto-detect the language, or set it yourself
  • Markers land timed to the timeline, not to an exported file
  • Runs in After Effects and Premiere Pro from the same panel
Any language

Ninety-nine languages. It works out which.


VoxMark transcribes with Whisper, which was trained on 99 languages, and it asks nothing of you to use them. Leave the dropdown on Auto-detect and the language is identified from the audio itself — a Japanese voiceover, an Arabic interview, a Spanish read under an English music bed — and the markers arrive on the timeline in that script.

  • Auto-detect reads the language off the audio — there is no setting to get wrong
  • Or pin it: all 99 are in the dropdown, with the fourteen most-used at the top
  • Non-Latin scripts land on the markers as they are — CJK, Cyrillic, Arabic, Devanagari
  • The SRT is written as UTF-8, so the script survives the handoff to Premiere, Cinema 4D or anywhere else
  • Words or phrases, Navigate, marker memory — every feature behaves the same whatever the language
99
languages · auto-detected
EnglishEspañolFrançaisDeutschItalianoPortuguêsNederlandsPolskiРусский日本語한국어中文العربيةहिन्दी+ 85 more, auto-detected

Groq and OpenAI both run Whisper, so the range is identical on either — Groq is simply faster, and free. AssemblyAI brings its own multilingual model, and auto-detects too. And a straight word on accuracy: Whisper is at its best on the widely spoken languages and softens on the smaller ones. That is the model, not the tool — every product built on Whisper shares the same ceiling. What differs is what happens to the words afterwards: here they become frame-accurate, colour-coded markers you can click.

Every voice, its own colour

Multi-speaker audio, sorted out for you.


Turn on speaker identification and a two-hander comes back as two colours — in the panel and on the markers themselves, using After Effects' own label colours. Rename "Speaker A" to the person who actually said it and every marker repaints live, in a single undo.

Andy Judy Jay

Three voices, three colours, on the timeline itself
Automatic
  • AssemblyAI tells the voices apart; VoxMark colours them
  • Colours land on the real markers, not just in the panel
  • Rename a voice and every marker follows, in one undo
  • Click a swatch to cycle the colour
Or by hand — any provider
  • Ctrl-click, shift-click or ctrl-drag to paint a selection across captions — the selected lines go bold
  • Assign a voice; the colours land immediately. The dropdown starts empty every time, so a voice is only ever applied because you picked it
  • Leave the name blank and the voice is colour only — the markers take the colour, the captions carry no name. Handy for sections and takes
  • clears the selection and leaves the controls where they are
  • Works on Groq and OpenAI too, where there is no diarization
  • Worth doing on a single-voice VO as well, just for the colour
Navigate

The transcript is the controller.


Once a comp is transcribed, the Navigate tab is how you get around it. Click any caption and the playhead jumps to it. Scrub the timeline and the current line highlights as you pass it — in Premiere it follows live playback too. The transport row steps marker to marker, so you stop dragging along the timeline ruler hunting for the start of a sentence.

  • Click a caption, the playhead goes there
  • Scrub and the current line highlights as you pass it
  • Transport buttons step marker to marker, or snap to the ends
  • Timecodes shown per line, tabular and readable
  • Captions are banded by voice — one voice's block shares a background, the next voice's takes the other, so who says what reads at a glance
  • Set preview to caption snaps the work area — in and out points in Premiere — to the captions you have selected
  • In Premiere the transcript follows live playback
Words or phrases

Per-word for kinetic type. Per-phrase for subtitles.


One dropdown decides how finely the transcript is cut, and it is worth setting deliberately — the right granularity is the difference between markers you cut against and markers you animate to.

Segments (phrases)
  • A marker per phrase — what you want for editing and for subtitles
  • Split long lines at a character limit you set, so no caption arrives too wide to read
  • The default, and the right answer most of the time
Words
  • A marker on every single word
  • What kinetic typography actually needs — timing you would otherwise key by ear
  • Same transcript, same audio; just a finer cut

Either way, the language is auto-detected unless you set it, and you can clear the existing markers first if you want a clean slate rather than a second set layered over the first.

Housekeeping

Markers when you want them. Out of the way when you don't.


A full transcript on the timeline is a lot of text. Two toggles calm it down without losing anything: hide the caption text so the markers stay but the words live only in the panel, or collapse ranged markers to single frames. Both are reversible, and neither touches the transcript.

Declutter
  • Aa — hide caption text on the timeline; the words stay in the panel
  • — collapse ranged markers to single-frame markers
  • Both toggle straight back, and the transcript is untouched either way
  • Both work in Premiere Pro as well as After Effects
Marker memory
  • Save & Clear parks a comp's markers and empties the timeline
  • Restore puts them back, with the count on the button
  • Every comp keeps its own memory — and it survives a rename
  • It refuses to restore one comp's markers into another
Handoff

SRT in, SRT out — with your frame rate baked in.


Export a proper SRT for subtitles and captions anywhere. VoxMark's carries two things a standard SRT does not: the exact frame rate of the project it came from, and the voice each line belongs to. That is what makes a round trip survive.

Export
  • A standards-compliant .srt, usable anywhere
  • Frame rate written into the header — no guessing on the other end
  • Speaker names carried per line
Import
  • Load any existing .srt and get markers from it
  • No transcription, no API key, no cost
  • Imported markers go into memory too, so Restore works on them
Cinema 4D
  • VoxMark for Cinema 4D reads the frame rate and voices natively
  • Transcribe in After Effects, animate to the same markers in C4D
  • Colours and names intact across the jump

Bring your own AI

Your key, your provider, your costs.


There is no VoxMark subscription and no middleman server. The panel talks straight to the transcription service you choose, using a key you hold, stored locally. Three are supported, and the panel links you to each one's key page.

Groq — recommended
  • Free tier, and genuinely fast
  • The same Whisper model, served quicker
  • Where most people should start
OpenAI Whisper
  • Pay-per-use, billed per minute of audio
  • Sensible if you already have a key
AssemblyAI
  • Five hours free, good on long files
  • The one that identifies speakers for you
Both editors

After Effects and Premiere Pro. One licence.


The same panel loads in both hosts and relabels itself to suit — Comp in After Effects, Sequence in Premiere. Scan the active sequence, render its audio mix, drop the markers on the timeline; in Premiere the transcript also follows live playback. You are not buying the same tool twice.

  • After Effects 2022–2026 and Premiere Pro 2022–2026
  • UI labels swap between Comp and Sequence depending on the host
  • Markers, colours and SRT behave the same in both
  • Convert to Captions turns the markers into a native Premiere caption track — no SRT round trip
  • Voices by hand, hidden caption text and single-frame markers all work in Premiere too, and the marker colours match the panel
  • Windows and macOS
Specification

What it needs, and what it works with.


Host and platform
  • After Effects 2022, 2023, 2024, 2025, 2026
  • Premiere Pro 2022, 2023, 2024, 2025, 2026
  • Windows and macOS
  • Version 2.0.4, and a perpetual licence with free updates

Under the hood
  • A CEP panel; the audio mix is rendered by the host itself
  • Transcription runs on your chosen provider with your own API key, stored locally
  • Supported providers: Groq Whisper, OpenAI Whisper, AssemblyAI
  • 99 languages through Whisper — auto-detected, or set per job from the dropdown; markers and SRT are Unicode
  • Audio is rendered at 16 kHz mono and streamed to the provider, so sequence length does not drive memory use
  • In After Effects the panel needs its own VoxMark 16k Mono output template — a one-time load, and the panel walks you through it

In and out
  • Comp and sequence markers, ranged or single-frame, with or without caption text
  • SRT import and export; exported files carry frame rate and voice data
  • Premiere Pro caption tracks, built from the markers with Convert to Captions
  • Round-trips to VoxMark for Cinema 4D with colours and names intact
Installing VoxMark

Install it, activate it, and never chase an update.


Your download is a zip holding the panel — VoxMark_v2.0.4.zxp — and a short INSTALL.txt. Three steps, about two minutes:

  1. Drag the .zxp onto a ZXP installer — the free ZXP/UXP Installer is the easy one. The certificate is self-signed, so you get a one-time security prompt; allow it and it never asks again.
  2. Restart After Effects or Premiere Pro and open Window ▸ Extensions ▸ VoxMark.
  3. VoxMark asks for a licence on first launch: paste the key from your Gumroad receipt, click Activate, then paste an API key from your chosen provider. Done.

One thing After Effects makes us ask for: VoxMark renders through its own VoxMark 16k Mono output template, and After Effects does not let a plugin add one by itself. The panel spots this on first run and walks you through loading it — five clicks, once, and it never comes up again. Premiere needs nothing.


Updates come to you. When a newer release is out an ↑ Update button appears in the panel's header; click it and the new .zxp lands in your Downloads folder, ready to drop onto your ZXP installer. Perpetual licence, free updates — version 2.0.4 today, and every version after it.

Not for you? There is a 30-day money-back guarantee — ask and you get a refund.

Questions

Questions, answered.


Which versions does VoxMark support?

After Effects 2022 to 2026 and Premiere Pro 2022 to 2026, on Windows and macOS. One licence covers both applications.

Do I need a subscription?

No. VoxMark is a one-off purchase and there is no VoxMark service to subscribe to. Transcription runs on a provider you choose, with an API key you hold — and Groq's free tier is enough for most work, so for many people the running cost is nothing.

Which transcription provider should I use?

Start with Groq — free, and the fastest of the three on the same Whisper model. Use OpenAI if you already have a key and would rather keep everything in one account; it bills per minute of audio. Use AssemblyAI when you want speakers identified automatically, or when the file is long — it gives five hours free. The panel links you straight to each provider's key page.

Where does my API key go?

It is stored locally on your machine and the panel talks directly to the provider. There is no VoxMark server in between.

Does it transcribe a file, or my actual comp?

Your actual comp. Scan Comp finds the audio layers — including ones inside nested precomps — and Render & Transcribe renders the real mix, music bed and all, then sends that. Nothing to export by hand, and nothing to re-sync afterwards, because the timings come from the timeline itself.

Word-level or phrase-level markers?

Both, from one dropdown. Segments give you phrase markers, which is what you want for cutting and for subtitles. Words give you a marker per word, which is what kinetic typography needs. Long lines can also be split at a character limit you set.

Can it tell speakers apart?

Yes, with AssemblyAI selected and "Identify speakers" ticked — each voice comes back as its own colour, on the markers themselves as well as in the panel. Rename "Speaker A" to a real name and every marker repaints in one undo. On Groq or OpenAI, where there is no diarization, you can select captions in the panel and assign voices by hand; the result is identical.

The markers clutter my timeline. Can I calm them down?

Yes, two ways, both reversible. Aa hides the caption text so the markers remain but the words live only in the panel, and collapses ranged markers to single-frame ones. Neither affects the transcript.

Can I clear the markers and get them back later?

Save & Clear parks a comp's markers in memory and empties the timeline; Restore puts them back, with the count shown on the button. Each comp keeps its own memory, and that memory survives renaming the comp — so you cannot restore one comp's markers into another by accident.

What is different about VoxMark's SRT?

It is a standard .srt and works anywhere one does. It also carries the exact frame rate of the project it came from, and the voice each line belongs to — neither of which a plain SRT records. That is what lets a transcript move between applications without the timings drifting.

Can I make markers from an SRT I already have?

Yes. Import SRT builds markers from any subtitle file with no transcription, no API key and no cost. The imported markers go into marker memory too, so Restore works on them.

Does it work with Cinema 4D?

VoxMark for Cinema 4D reads VoxMark's SRT natively, frame rate and voices included — so you can transcribe in After Effects and animate to the same markers, in the same colours, in C4D.

Which languages does it handle?

Ninety-nine. VoxMark transcribes with Whisper, which covers 99 languages. Leave the language on Auto-detect and it is identified from the audio on every provider, or pick any of the 99 from the dropdown — the fourteen most-used are pinned at the top. Markers and the exported SRT are Unicode throughout, so Japanese, Arabic, Hindi or Russian land on the timeline in their own script.

How do I install and activate it?

Your download is a zip holding VoxMark_v2.0.4.zxp and a short INSTALL.txt. Drag the ZXP onto a ZXP installer — the free ZXP/UXP Installer does it in one drop. The certificate is self-signed, so there is a one-time security prompt to allow. Restart After Effects or Premiere Pro, open Window ▸ Extensions ▸ VoxMark, and paste the licence key from your Gumroad receipt into the screen that appears, then click Activate.

What if it is not for me?

There is a 30-day money-back guarantee. Buy it, try it on real work, and if it does not fit the way you cut, ask for a refund within 30 days and you get one.

How do I get updates, and do they cost anything?

They are free, for as long as VoxMark exists — it is a perpetual licence, not a subscription. The panel watches for new releases itself: when one lands, an ↑ Update button appears in its header, and clicking it drops the new .zxp into your Downloads folder for you to install over the old one. Nothing to check for, no upgrade fee.

Can I use it on more than one machine?

Yes. The key is not tied to a particular computer, so your desktop and your laptop are both fine. It is a licence for one artist rather than a whole studio — but that one licence covers After Effects and Premiere Pro, so you are never buying the same tool twice.

Where do I get help if something goes wrong?

The Chroma Discord — that is where VoxMark is supported, and where you will get an answer quickest. The panel also keeps a log you can open from its own menu, which usually says exactly what a provider objected to.

Stop scrubbing for words. Click the sentence.


One click turns your voiceover into frame-accurate, colour-coded markers — in After Effects and Premiere Pro, on the AI provider of your choice, including a free one.

Perpetual licence · free updates · 30-day money-back guarantee