Transcript viewer
A React component in Beamline's Voice & audio category.
A recording and its words together: the spoken word lights up as it plays and the transcript follows it, any word or paragraph time plays from there, speakers become paragraphs with marks on the scrub bar. Takes word timings from speech-to-text or character alignment from text-to-speech.
Use it for
- Call reviews, meeting recordings and voice notes with timed words.
- Playing generated speech with its text in step.
Not for
- A transcript without audio (a plain text block).
- Live captions of a call in progress (the voice-agent block shows those).
- Audio without words (audio-player).
Anatomy
- player row (play, title, speed)
- transcript (paragraphs: speaker, time button, words)
- active word
- scrub bar with a mark per paragraph
Variants
- input: words (speech-to-text: text, start, end, speaker), alignment (text-to-speech: per-character times)
States
- paused
- playing (active word lit, transcript follows)
- following paused while you scroll (resumes after a few seconds)
- audio tags shown as cues or hidden
- no words (empty state)
- audio failed (the player's error and retry)
Keyboard
- Tab: play, speed, each paragraph's time button, the transcript (scrolls with arrows), the scrub bar
- A paragraph's time button plays from its start (Enter / Space)
- Scrub bar: ← → seek 5 s, Home / End, Space plays and pauses
Motion
The spoken word's tint and underline is one plate that travels word to word as the audio plays (--pui-duration-fast); the transcript scrolls to keep it in view (smoothly, or at once under reduced motion).
Props
| Prop | Type | Default | Description |
|---|---|---|---|
src (required) | string | — | The audio. |
words | TranscriptWord[] | — | { text, start, end, speaker? } from speech-to-text. |
alignment | CharacterAlignment | — | Per-character timing from text-to-speech; turned into words. |
title / subtitle | string | — | |
hideAudioTags | boolean | — | Hide bracketed delivery cues like [laughs] (shown as quiet cues by default). |
follow | boolean | true | Keep the spoken word in view. |
height | number | 240 | Transcript height in px; it scrolls inside. |
onWordClick | (word: TranscriptWord) => void | — | Default: play from the word. |
classNames | TranscriptViewerClassNames | — | |
style / ref | CSSProperties / native Ref | — | |
playerClassNames | AudioPlayerClassNames | — |
Built from
- audio-player
- scrub-bar
- avatar
- button
- empty-state
Dependencies
lucide-react
Import
import { TranscriptViewer } from "@/components/transcript-viewer/transcript-viewer";
Get it
Transcript viewer is part of Beamline: 190 React components and screens your coding agent (Claude Code, Codex, Cursor or any MCP client) installs into your app through Beamline's MCP server, as a ready-built package or as plain React source you can change. $49 one payment (regular $200), no subscription, a year of updates, one licence for your whole team. Get Beamline · Connect your agent
More in Voice & audio
- Audio player: One audio element and its controls: play with a loading state, a scrub bar with time, skip, speed and a playlist whose rows play into the same player…
- Bar visualizer: A row of bars that says what a voice agent is doing: a sweep while connecting, the person's voice while listening, a slow breath while thinking, the…
- Conversation bar: The control bar of a voice call with an agent: one button to call and hang up, the agent's state as bars, a call timer, mute, and typing when talking is…
- Live waveform: Loudness drawn as it happens: a scrolling history of the microphone (or any live source) or its spectrum, a calm processing state while words are…
- Matrix: A dot-matrix display: animated frame sets (loading, pulse, wave, syncing), text and digits in a 5×7 font, or a VU meter of levels or a live source. One…
- Microphone selector: Choose the microphone, test it with a live level and mute it, with every permission and device state spelled out: not yet allowed, blocked (and how to…
- Orb: The agent's presence: a thin ring of light around a lens in the brand's hue that holds still at rest, sends an arc round while thinking and widens with…
- Scrub bar: A seekable timeline for audio or video: elapsed and total time, what has loaded, chapter marks, a time preview under the pointer, and the full keyboard…
- Speech input: A text field you can speak into: the mic dictates through any speech-to-text provider, the phrase being heard shows apart from the finished words, and…
- Voice button: The control that starts and stops talking: click to toggle or hold to talk (pointer, Space, Enter or a global shortcut), with a live level while…
- Voice picker: Choose a text-to-speech voice from a searchable list (name, accent, tone, use) and hear it first: each voice has its own orb, previews play one at a time…
- Waveform: The shape of a recording: one bar per slice of loudness, fitted to any width, with the played part in the accent, unmeasured slices hatched and a loading…