Speech input

A React component in Beamline's Voice & audio category.

A text field you can speak into: the mic dictates through any speech-to-text provider, the phrase being heard shows apart from the finished words, and every failure (blocked, no microphone, connection lost and back) says what to do.

Use it for

  • Letting people speak instead of typing: notes, comments, a composer, a search.
  • Forms where long answers are easier said than typed.

Not for

Anatomy

  • label
  • field (text, multi- or single-line)
  • mic button (voice-button, icon only)
  • cancel (while listening)
  • listening strip: live level and the phrase being heard
  • status line (reconnecting, microphone fallback)
  • failure with fix and retry
  • description

Variants

  • multiline: true, false
  • size: sm, md, lg

States

  • idle
  • asking for the microphone
  • connecting
  • listening (partial phrase in the strip)
  • reconnecting
  • finishing (last words)
  • error: blocked / no microphone / in use / connection / service
  • without a microphone (the provider runs on its own audio)
  • disabled

Keyboard

  • Tab reaches the field and the mic button
  • Enter or Space on the mic starts and stops
  • Escape while listening cancels and drops the phrase in progress
  • Typing in the field keeps working while dictating

Motion

The level draws only while listening and on screen; the strip opens and closes without animation under reduced motion (level becomes a number).

Props

PropTypeDefaultDescription
adapter (required)SpeechToText—Your provider, wired through the contract in @/lib/voice. Wiring a provider is one adapter. Browser (no key): browserSpeechToText() from @/lib/voice. Deepgram (live): open new WebSocket('wss://api.deepgram.com/v1/listen?model=nova-3&interim_results=true&smart_format=true', ['token', key]); record the stream with new MediaRecorder(stream, { mimeType: 'audio/webm' }) and send each chunk (ondataavailable, start(250)); on each message parse JSON: if type is 'Results', text = channel.alternatives[0].transcript, call is_final ? on.final(text) : on.partial(text); stop() sends JSON {type: 'CloseStream'} and resolves on close. ElevenLabs realtime speech-to-text: fetch a single-use token from your server, connect with their client SDK or WebSocket, and map its partial transcript events to on.partial and committed transcripts to on.final. OpenAI Realtime transcription: send audio buffers over the session and map transcription delta events to on.partial (accumulated) and completed events to on.final. Report failures as { kind: 'auth' | 'network' | 'quota', message } so the field shows the right fix.
value / defaultValue / onValueChangestring—The text; finished phrases are appended.
label (required)ReactNode—
hideLabel / description / placeholderboolean / ReactNode / string—
multilinebooleantrue
rowsnumber3
languagestring—BCP 47, passed to the adapter.
microphoneMicrophone—Share the input choice with a MicSelector.
shortcutstring[]—A page-wide key to start and stop (VoiceButton's).
onPartial / onFinal(text: string) => void—
onError(error: VoiceError) => void—
size / disabled / name / id"sm" | "md" | "lg" / boolean / string / string—
classNamesSpeechInputClassNames—Classes for documented visual parts; internal refs and event handlers are retained.
style / refCSSProperties / Ref—Visual root style; ref targets native textarea or input, according to multiline.

Built from

Dependencies

lucide-react

Import

import { SpeechInput } from "@/components/speech-input/speech-input";

Get it

Speech input is part of Beamline: 190 React components and screens your coding agent (Claude Code, Codex, Cursor or any MCP client) installs into your app through Beamline's MCP server, as a ready-built package or as plain React source you can change. $49 one payment (regular $200), no subscription, a year of updates, one licence for your whole team. Get Beamline · Connect your agent

More in Voice & audio

  • Audio player: One audio element and its controls: play with a loading state, a scrub bar with time, skip, speed and a playlist whose rows play into the same player…
  • Bar visualizer: A row of bars that says what a voice agent is doing: a sweep while connecting, the person's voice while listening, a slow breath while thinking, the…
  • Conversation bar: The control bar of a voice call with an agent: one button to call and hang up, the agent's state as bars, a call timer, mute, and typing when talking is…
  • Live waveform: Loudness drawn as it happens: a scrolling history of the microphone (or any live source) or its spectrum, a calm processing state while words are…
  • Matrix: A dot-matrix display: animated frame sets (loading, pulse, wave, syncing), text and digits in a 5×7 font, or a VU meter of levels or a live source. One…
  • Microphone selector: Choose the microphone, test it with a live level and mute it, with every permission and device state spelled out: not yet allowed, blocked (and how to…
  • Orb: The agent's presence: a thin ring of light around a lens in the brand's hue that holds still at rest, sends an arc round while thinking and widens with…
  • Scrub bar: A seekable timeline for audio or video: elapsed and total time, what has loaded, chapter marks, a time preview under the pointer, and the full keyboard…
  • Transcript viewer: A recording and its words together: the spoken word lights up as it plays and the transcript follows it, any word or paragraph time plays from there…
  • Voice button: The control that starts and stops talking: click to toggle or hold to talk (pointer, Space, Enter or a global shortcut), with a live level while…
  • Voice picker: Choose a text-to-speech voice from a searchable list (name, accent, tone, use) and hear it first: each voice has its own orb, previews play one at a time…
  • Waveform: The shape of a recording: one bar per slice of loudness, fitted to any width, with the played part in the accent, unmeasured slices hatched and a loading…