Speech input
A React component in Beamline's Voice & audio category.
A text field you can speak into: the mic dictates through any speech-to-text provider, the phrase being heard shows apart from the finished words, and every failure (blocked, no microphone, connection lost and back) says what to do.
Use it for
- Letting people speak instead of typing: notes, comments, a composer, a search.
- Forms where long answers are easier said than typed.
Not for
- A two-way conversation with an agent (conversation-bar or the voice-agent block).
- Only a record button (voice-button).
- Reviewing a recording's words (transcript-viewer).
Anatomy
- label
- field (text, multi- or single-line)
- mic button (voice-button, icon only)
- cancel (while listening)
- listening strip: live level and the phrase being heard
- status line (reconnecting, microphone fallback)
- failure with fix and retry
- description
Variants
- multiline: true, false
- size: sm, md, lg
States
- idle
- asking for the microphone
- connecting
- listening (partial phrase in the strip)
- reconnecting
- finishing (last words)
- error: blocked / no microphone / in use / connection / service
- without a microphone (the provider runs on its own audio)
- disabled
Keyboard
- Tab reaches the field and the mic button
- Enter or Space on the mic starts and stops
- Escape while listening cancels and drops the phrase in progress
- Typing in the field keeps working while dictating
Motion
The level draws only while listening and on screen; the strip opens and closes without animation under reduced motion (level becomes a number).
Props
| Prop | Type | Default | Description |
|---|---|---|---|
adapter (required) | SpeechToText | — | Your provider, wired through the contract in @/lib/voice. Wiring a provider is one adapter. Browser (no key): browserSpeechToText() from @/lib/voice. Deepgram (live): open new WebSocket('wss://api.deepgram.com/v1/listen?model=nova-3&interim_results=true&smart_format=true', ['token', key]); record the stream with new MediaRecorder(stream, { mimeType: 'audio/webm' }) and send each chunk (ondataavailable, start(250)); on each message parse JSON: if type is 'Results', text = channel.alternatives[0].transcript, call is_final ? on.final(text) : on.partial(text); stop() sends JSON {type: 'CloseStream'} and resolves on close. ElevenLabs realtime speech-to-text: fetch a single-use token from your server, connect with their client SDK or WebSocket, and map its partial transcript events to on.partial and committed transcripts to on.final. OpenAI Realtime transcription: send audio buffers over the session and map transcription delta events to on.partial (accumulated) and completed events to on.final. Report failures as { kind: 'auth' | 'network' | 'quota', message } so the field shows the right fix. |
value / defaultValue / onValueChange | string | — | The text; finished phrases are appended. |
label (required) | ReactNode | — | |
hideLabel / description / placeholder | boolean / ReactNode / string | — | |
multiline | boolean | true | |
rows | number | 3 | |
language | string | — | BCP 47, passed to the adapter. |
microphone | Microphone | — | Share the input choice with a MicSelector. |
shortcut | string[] | — | A page-wide key to start and stop (VoiceButton's). |
onPartial / onFinal | (text: string) => void | — | |
onError | (error: VoiceError) => void | — | |
size / disabled / name / id | "sm" | "md" | "lg" / boolean / string / string | — | |
classNames | SpeechInputClassNames | — | Classes for documented visual parts; internal refs and event handlers are retained. |
style / ref | CSSProperties / Ref | — | Visual root style; ref targets native textarea or input, according to multiline. |
Built from
- voice-button
- live-waveform
- button
- tooltip
Dependencies
lucide-react
Import
import { SpeechInput } from "@/components/speech-input/speech-input";
Get it
Speech input is part of Beamline: 190 React components and screens your coding agent (Claude Code, Codex, Cursor or any MCP client) installs into your app through Beamline's MCP server, as a ready-built package or as plain React source you can change. $49 one payment (regular $200), no subscription, a year of updates, one licence for your whole team. Get Beamline · Connect your agent
More in Voice & audio
- Audio player: One audio element and its controls: play with a loading state, a scrub bar with time, skip, speed and a playlist whose rows play into the same player…
- Bar visualizer: A row of bars that says what a voice agent is doing: a sweep while connecting, the person's voice while listening, a slow breath while thinking, the…
- Conversation bar: The control bar of a voice call with an agent: one button to call and hang up, the agent's state as bars, a call timer, mute, and typing when talking is…
- Live waveform: Loudness drawn as it happens: a scrolling history of the microphone (or any live source) or its spectrum, a calm processing state while words are…
- Matrix: A dot-matrix display: animated frame sets (loading, pulse, wave, syncing), text and digits in a 5×7 font, or a VU meter of levels or a live source. One…
- Microphone selector: Choose the microphone, test it with a live level and mute it, with every permission and device state spelled out: not yet allowed, blocked (and how to…
- Orb: The agent's presence: a thin ring of light around a lens in the brand's hue that holds still at rest, sends an arc round while thinking and widens with…
- Scrub bar: A seekable timeline for audio or video: elapsed and total time, what has loaded, chapter marks, a time preview under the pointer, and the full keyboard…
- Transcript viewer: A recording and its words together: the spoken word lights up as it plays and the transcript follows it, any word or paragraph time plays from there…
- Voice button: The control that starts and stops talking: click to toggle or hold to talk (pointer, Space, Enter or a global shortcut), with a live level while…
- Voice picker: Choose a text-to-speech voice from a searchable list (name, accent, tone, use) and hear it first: each voice has its own orb, previews play one at a time…
- Waveform: The shape of a recording: one bar per slice of loudness, fitted to any width, with the played part in the accent, unmeasured slices hatched and a loading…