VoxMark - After Effects & Premiere Voiceover to Markers (Win/macOS)

Your voiceover becomes frame-accurate markers. One click.Cutting to a voiceover means scrubbing for words. VoxMark ends that. Point it at your comp or sequence — not at a file you exported — and it renders the real audio mix, transcribes it with the AI provider of your choice, in any of 99 languages, auto-detected, and drops the result onto the timeline as markers or subtitles. Multiple speakers come back named and colour-coded. The transcript then lives in the panel as a clickable navigator: click the sentence, the playhead is there.After Effects · Premiere Pro · Groq · OpenAI · AssemblyAI · 99 languages · SRT in and out · Cinema 4DVersion 2.0.4 — After Effects and Premiere Pro 2022 to 2026, Windows and macOS. One licence covers both applications.The old way, and the VoxMark wayWithout VoxMark Export the VO, upload it somewhere, wait, download a transcript Read the transcript in one window and scrub the timeline in another Type marker comments by hand, one line at a time Nudge every marker because the export did not line up with the comp Work out who said what, then colour the markers yourself — if you bother Hand off an SRT and hope the frame rate matches on the other end With VoxMark Scan the comp, hit Render & Transcribe — it renders the real mix itself Markers land on the timeline already timed and already worded Words or phrases, your choice, with a max-length split Speakers identified, named and colour-coded automatically Click a line in the panel and the playhead goes there Export an SRT that carries your exact frame rate Point it at the comp, not at a fileMost transcription tools want a wav. VoxMark wants your composition. Scan Comp lists every audio layer it can find — including the ones buried in nested precomps — and Render & Transcribe renders the actual mix you would hear, music bed and all, then sends that off. Nothing to export, nothing to line back up afterwards, and the timings stay honest because they came from the timeline in the first place. Scans the comp or sequence for audio, nested precomps included Renders the real mix — VO under music, layered tracks, all of it Renders at 16 kHz mono, which is what the speech models listen to anyway — so a ninety-minute sequence uploads in a fraction of the time, with nothing lost Auto-detect the language, or set it yourself Markers land timed to the timeline, not to an exported file Runs in After Effects and Premiere Pro from the same panel Ninety-nine languages. It works out which.VoxMark transcribes with Whisper, which was trained on 99 languages, and it asks nothing of you to use them. Leave the dropdown on Auto-detect and the language is identified from the audio itself — a Japanese voiceover, an Arabic interview, a Spanish read under an English music bed — and the markers arrive on the timeline in that script. Auto-detect reads the language off the audio — there is no setting to get wrong Or pin it: all 99 are in the dropdown, with the fourteen most-used at the top Non-Latin scripts land on the markers as they are — CJK, Cyrillic, Arabic, Devanagari The SRT is written as UTF-8, so the script survives the handoff to Premiere, Cinema 4D or anywhere else Words or phrases, Navigate, marker memory — every feature behaves the same whatever the language Groq and OpenAI both run Whisper, so the range is identical on either — Groq is simply faster, and free. AssemblyAI brings its own multilingual model, and auto-detects too. And a straight word on accuracy: Whisper is at its best on the widely spoken languages and softens on the smaller ones. That is the model, not the tool — every product built on Whisper shares the same ceiling. What differs is what happens to the words afterwards: here they become frame-accurate, colour-coded markers you can click.Multi-speaker audio, sorted out for youTurn on speaker identification and a two-hander comes back as two colours — in the panel and on the markers themselves, using After Effects' own label colours. Rename "Speaker A" to the person who actually said it and every marker repaints live, in a single undo.Automatic AssemblyAI tells the voices apart; VoxMark colours them Colours land on the real markers, not just in the panel Rename a voice and every marker follows, in one undo Click a swatch to cycle the colour Or by hand — any provider Ctrl-click, shift-click or ctrl-drag to paint a selection across captions — the selected lines go bold Assign a voice; the colours land immediately. The dropdown starts empty every time, so a voice is only ever applied because you picked it Leave the name blank and the voice is colour only — the markers take the colour, the captions carry no name. Handy for sections and takes The clear button drops the selection and leaves the controls where they are Works on Groq and OpenAI too, where there is no diarization Worth doing on a single-voice VO as well, just for the colour The transcript is the controllerOnce a comp is transcribed, the Navigate tab is how you get around it. Click any caption and the playhead jumps to it. Scrub the timeline and the current line highlights as you pass it — in Premiere it follows live playback too. The transport row steps marker to marker, so you stop dragging along the timeline ruler hunting for the start of a sentence. Click a caption, the playhead goes there Scrub and the current line highlights as you pass it Transport buttons step marker to marker, or snap to the ends Timecodes shown per line, tabular and readable Captions are banded by voice — one voice's block shares a background, the next voice's takes the other, so who says what reads at a glance Set preview to caption snaps the work area — in and out points in Premiere — to the captions you have selected In Premiere the transcript follows live playback Per-word for kinetic type. Per-phrase for subtitles.One dropdown decides how finely the transcript is cut, and it is worth setting deliberately — the right granularity is the difference between markers you cut against and markers you animate to.Segments (phrases) A marker per phrase — what you want for editing and for subtitles Split long lines at a character limit you set, so no caption arrives too wide to read The default, and the right answer most of the time Words A marker on every single word What kinetic typography actually needs — timing you would otherwise key by ear Same transcript, same audio; just a finer cut Either way, the language is auto-detected unless you set it, and you can clear the existing markers first if you want a clean slate rather than a second set layered over the first.Markers when you want them. Out of the way when you don't.A full transcript on the timeline is a lot of text. Two toggles calm it down without losing anything: hide the caption text so the markers stay but the words live only in the panel, or collapse ranged markers to single frames. Both are reversible, and neither touches the transcript.Declutter Hide caption text on the timeline; the words stay in the panel Collapse ranged markers to single-frame markers Both toggle straight back, and the transcript is untouched either way Both work in Premiere Pro as well as After Effects Marker memory Save & Clear parks a comp's markers and empties the timeline Restore puts them back, with the count on the button Every comp keeps its own memory — and it survives a rename It refuses to restore one comp's markers into another SRT in, SRT out — with your frame rate baked inExport a proper SRT for subtitles and captions anywhere. VoxMark's carries two things a standard SRT does not: the exact frame rate of the project it came from, and the voice each line belongs to. That is what makes a round trip survive.Export A standards-compliant .srt, usable anywhere Frame rate written into the header — no guessing on the other end Speaker names carried per line Import Load any existing .srt and get markers from it No transcription, no API key, no cost Imported markers go into memory too, so Restore works on them Cinema 4D VoxMark for Cinema 4D reads the frame rate and voices natively Transcribe in After Effects, animate to the same markers in C4D Colours and names intact across the jump Your key, your provider, your costsThere is no VoxMark subscription and no middleman server. The panel talks straight to the transcription service you choose, using a key you hold, stored locally. Three are supported, and the panel links you to each one's key page.Groq — recommended Free tier, and genuinely fast The same Whisper model, served quicker Where most people should start OpenAI Whisper Pay-per-use, billed per minute of audio Sensible if you already have a key AssemblyAI Five hours free, good on long files The one that identifies speakers for you After Effects and Premiere Pro. One licence.The same panel loads in both hosts and relabels itself to suit — Comp in After Effects, Sequence in Premiere. Scan the active sequence, render its audio mix, drop the markers on the timeline; in Premiere the transcript also follows live playback. You are not buying the same tool twice. After Effects 2022 to 2026 and Premiere Pro 2022 to 2026 UI labels swap between Comp and Sequence depending on the host Markers, colours and SRT behave the same in both Convert to Captions turns the markers into a native Premiere caption track — no SRT round trip Voices by hand, hidden caption text and single-frame markers all work in Premiere too, and the marker colours match the panel Windows and macOS New in 2.0.4Assigning voices to captions, fixed and made quicker — and voices can now be colour only, with no name. The voice dropdown starts empty every time. It used to arrive pre-loaded with whoever you assigned last, so a fresh run of captions could silently take the previous voice's name Ctrl-drag across captions to select a run of them — drag from an unselected caption to select, from a selected one to deselect, at any speed. Selected captions are shown in bold The clear button drops the selection and leaves the controls in place, instead of closing the whole bar The assign controls stay on screen whenever you have a transcript, with the instructions above them Captions are banded by voice, so who says what is legible at a glance rather than from a 3px stripe A voice can have no name — it colour-codes your markers without putting a name on the captions. Useful for marking up sections and takes Renaming, recolouring and removing unnamed voices works correctly in both After Effects and Premiere Pro One limit worth knowing: an unnamed voice is not recovered by Rebuild from markers, because there is no name in the marker to read back. The markers keep their colours; they simply stop being listed as a voice.New in 2.0.3Long sequences transcribe again — and Premiere gains the voices, colours and declutter tools After Effects already had. Audio is rendered at 16 kHz mono. An 86-minute sequence went from 946 MB to about 165 MB with no loss of accuracy — speaker identification included, since it listens to the audio rather than the channels The upload is streamed rather than held in memory, so length no longer drives memory use Upload errors report the actual HTTP status and the server's reply, and requests time out instead of hanging Files too large for AssemblyAI are refused up front with a clear message After Effects users load the new VoxMark 16k Mono output template once; the panel prompts and walks you through it Premiere: assigning a voice to selected captions now works, and marker colours match the panel Premiere: hide caption text and single-frame markers both work, and a voice can be removed from the Voices list Set preview to caption sets the work area, or in and out points in Premiere, to the captions you have selected Voice colours, marker durations and hidden caption text live in a VoxMark folder beside your project, so they travel with it Rebuild from markers has moved to the Transcribe tab, under Render & Transcribe Minimum supported version is After Effects and Premiere Pro 2022 What it needs, and what it works withHost and platform After Effects 2022, 2023, 2024, 2025, 2026 Premiere Pro 2022, 2023, 2024, 2025, 2026 Windows and macOS Under the hood A CEP panel; the audio mix is rendered by the host itself Transcription runs on your chosen provider with your own API key, stored locally Supported providers: Groq Whisper, OpenAI Whisper, AssemblyAI 99 languages through Whisper — auto-detected, or set per job from the dropdown; markers and SRT are Unicode Audio is rendered at 16 kHz mono and streamed to the provider, so sequence length does not drive memory use In After Effects the panel needs its own VoxMark 16k Mono output template — a one-time load, and the panel walks you through it In and out Comp and sequence markers, ranged or single-frame, with or without caption text SRT import and export; exported files carry frame rate and voice data Premiere Pro caption tracks, built from the markers with Convert to Captions Round-trips to VoxMark for Cinema 4D with colours and names intact Installing and activating Download the zip from your Gumroad library — it contains VoxMark_v2.0.4.zxp and an INSTALL.txt Get the free ZXP/UXP Installer from aescripts.com/learn/zxp-installer/ and drag the .zxp onto it. The certificate is self-signed, so you will see a one-time security prompt — allow it Restart After Effects or Premiere Pro Open the panel from Window ▸ Extensions ▸ VoxMark Paste your Gumroad licence key and click Activate Paste an API key from your chosen provider, click Save, and you are ready to scan Updates. The panel checks for new versions itself. When one is available an Update button appears in its header — click it and the new .zxp downloads to your Downloads folder, ready to drop onto your ZXP installer. Your licence key carries over.Questions, answeredWhich versions does VoxMark support?After Effects 2022 to 2026 and Premiere Pro 2022 to 2026, on Windows and macOS. One licence covers both applications.Do I need a subscription?No. VoxMark is a one-off purchase and there is no VoxMark service to subscribe to. Transcription runs on a provider you choose, with an API key you hold — and Groq's free tier is enough for most work, so for many people the running cost is nothing.Which transcription provider should I use?Start with Groq — free, and the fastest of the three on the same Whisper model. Use OpenAI if you already have a key and would rather keep everything in one account; it bills per minute of audio. Use AssemblyAI when you want speakers identified automatically, or when the file is long — it gives five hours free. The panel links you straight to each provider's key page.Where does my API key go?It is stored locally on your machine and the panel talks directly to the provider. There is no VoxMark server in between.Does it transcribe a file, or my actual comp?Your actual comp. Scan Comp finds the audio layers — including ones inside nested precomps — and Render & Transcribe renders the real mix, music bed and all, then sends that. Nothing to export by hand, and nothing to re-sync afterwards, because the timings come from the timeline itself.Word-level or phrase-level markers?Both, from one dropdown. Segments give you phrase markers, which is what you want for cutting and for subtitles. Words give you a marker per word, which is what kinetic typography needs. Long lines can also be split at a character limit you set.Can it tell speakers apart?Yes, with AssemblyAI selected and Identify speakers ticked — each voice comes back as its own colour, on the markers themselves as well as in the panel. Rename "Speaker A" to a real name and every marker repaints in one undo. On Groq or OpenAI, where there is no diarization, you can select captions in the panel and assign voices by hand; the result is identical.The markers clutter my timeline. Can I calm them down?Yes, two ways, both reversible. One toggle hides the caption text so the markers remain but the words live only in the panel; the other collapses ranged markers to single-frame ones. Neither affects the transcript.Can I clear the markers and get them back later?Save & Clear parks a comp's markers in memory and empties the timeline; Restore puts them back, with the count shown on the button. Each comp keeps its own memory, and that memory survives renaming the comp — so you cannot restore one comp's markers into another by accident.What is different about VoxMark's SRT?It is a standard .srt and works anywhere one does. It also carries the exact frame rate of the project it came from, and the voice each line belongs to — neither of which a plain SRT records. That is what lets a transcript move between applications without the timings drifting.Can I make markers from an SRT I already have?Yes. Import SRT builds markers from any subtitle file with no transcription, no API key and no cost. The imported markers go into marker memory too, so Restore works on them.Does it work with Cinema 4D?VoxMark for Cinema 4D reads VoxMark's SRT natively, frame rate and voices included — so you can transcribe in After Effects and animate to the same markers, in the same colours, in C4D. It is a separate product.Which languages does it handle?Ninety-nine. VoxMark transcribes with Whisper, which covers 99 languages. Leave the language on Auto-detect and it is identified from the audio on every provider, or pick any of the 99 from the dropdown — the fourteen most-used are pinned at the top. Markers and the exported SRT are Unicode throughout, so Japanese, Arabic, Hindi or Russian land on the timeline in their own script.How do I get updates?The panel checks for them and shows an Update button in its header when a new version is out. Clicking it downloads the new .zxp to your Downloads folder; drop that onto your ZXP installer and restart. Your licence key carries over.Stop scrubbing for words. Click the sentence.One click turns your voiceover into frame-accurate, colour-coded markers — in After Effects and Premiere Pro, on the AI provider of your choice, including a free one.