Skip to content
Prism

Subtitles

Transcribe a voiceover and drop timed, styled subtitle text onto your After Effects timeline — a capability AE has no native answer for.

After Effects can place text and read a timeline, but it has no way to listen to a voiceover and turn it into timed captions. Prism adds that. Point Prism at the audio in your composition and it transcribes the speech, then writes a real text layer whose words are timed to the spoken track — native, editable, and yours to restyle. If you only need the timed text, not a caption layer, use Transcribe instead.

The short version
Transcription runs on Prism's server, not in AE.
You get a real, editable text layer with timing aligned to the voiceover.
Audio is processed for the transcript and then discarded.
Costs 5 credits per minute of audio.

What it does

Subtitle generation is one of Prism's AI-native tools. Prism transcribes your spoken track on its server and uses the timestamped result to build a subtitle text layer on the timeline. The result is the same kind of object you would create by hand: a text layer you can re-time, re-word, restyle, or split.

As with every Prism action, the AI reads the live state of your comp before it writes the subtitles as native layers (see how Prism works).

Before you start

  • Your composition contains the voiceover or audio track you want captioned.
  • Prism is installed and activated, and your AI client points at Prism's MCP server. If not, see setup and connect your AI.
  • The comp you want captioned is the active comp in After Effects.

How to use it

Run it from the Tools tab, or ask your AI in plain language — Prism runs the same tool either way.

1
Open the comp
Make the composition with your voiceover the active comp in After Effects.
2
Ask your AI in plain language
Tell your AI client something like "transcribe the voiceover and add subtitles to the timeline." Prism reads the audio and transcribes it.
3
Review the layer
A subtitle text layer appears, timed to the speech. Scrub the timeline to check the alignment against the spoken words.
4
Refine
Ask for restyling — font, size, position, color — or for edits to wording and timing. Prism rewrites the same native layer.

Audio is sent only to produce the transcript, then discarded — your project stays on your machine (how Prism works).

A worked example

You sayPrism does
"Transcribe the voiceover and add captions, one running line, no stroke."Reads the audio in your comp, transcribes it, and writes a single subtitle text layer timed to the speech.
"Now split those captions per word, karaoke style."Rebuilds the same layer with one timed word at a time, using the per-word timing from the transcript.

What you get back: a native subtitle text layer on the timeline, its words aligned to the spoken track and ready to restyle.

If the audio is silent or unreadable: the request surfaces "no speech detected in the audio." Check that the right audio layer is in the active comp and is not muted, then ask again.

Languages and caption format

Prism transcribes a wide range of languages and accents, so most voiceovers — not just English — come back clean. Timing comes back per word, which gives you two ready formats:

FormatLooks likeGood for
One running lineA full caption line on screen at a timeExplainers, talking-head, narration
Per wordOne word revealed at a timePunchy, karaoke-style reads

You can ask for either when you generate the subtitles, or switch between them afterward — the layer is ordinary AE text.

Styling and timing

The subtitle layer is ordinary AE text, so anything you would do by hand still applies — change the font, add an animator, reposition, or split lines. One gotcha to know: if glyphs look clipped, ask for subtitle text without a stroke — see troubleshooting for that and other AE text quirks.

Every Prism action is undoable in one step and protected by a restore point, so experimenting with styling is safe (how Prism works).

Notes and limits

Transcription draws on AI credits — your prepaid wallet for heavy, generative AI — at 5 credits per minute of audio (so a 3-minute track costs about 15 credits). For the monthly credit allowance and how metering works, see usage and limits.

  • Usage and limits — AI-credit allowances and how metering works.
  • Transcribe — the raw transcript and per-word timing, without a caption layer.
  • BPM and beat markers — detect tempo and write beat markers to the timeline.
  • SVG import — bring in SVG as real text and shapes, not outlines.
  • AI-native tools — the family of capabilities AE has no native answer for.

On this page