
# Build a voice agent UI in React

By Danylo Pravda, 2026-10-08

A voice interface has more states than a chat: the microphone is not yet allowed, or blocked, or missing; the agent
is listening, thinking or talking; the connection drops and comes back; the person would rather type for a moment.
Most voice demos handle the happy path and freeze on the rest. Beamline has thirteen voice parts and a
[voice agent](https://beamline.io/components/voice-agent) screen that puts them together on any provider. This guide covers how they fit.

![The Beamline voice agent: the agent's orb, the live transcript and the conversation bar](https://beamline.io/components/og/voice-agent.png "The voice agent block, on demo data")

## Three layouts, one session

The voice agent block draws the same session three ways:

- **Panel:** a support widget where people type or talk to the agent.
- **Call:** a phone-style screen with the agent's [orb](https://beamline.io/components/orb) and one call button.
- **Workspace:** voice as the main way to work, with an empty state, the live transcript and the
  [conversation bar](https://beamline.io/components/conversation-bar).

All three run on one session: the same adapter, or a session you share with the rest of the screen.

## One adapter for any provider

The block does not know which speech service you use. You give it an adapter with one method, `connect`, and it hands
you back handlers to report what happens:

```ts
import { VoiceAgent } from "@beamline/voice-agent";
import type { VoiceAgent as VoiceAdapter } from "@beamline/lib-voice";

export const myAgent: VoiceAdapter = {
  async connect({ stream, mode, signal }, on) {
    const call = await provider.start({ audio: stream, textOnly: mode === "text", signal });
    call.onState((s) => on.status(s)); // "listening" | "thinking" | "talking"
    call.onText((t) => on.transcript({ id: t.id, role: t.speaker, text: t.text, final: t.done }));
    call.onDrop(() => on.connection?.("reconnecting"));
    call.onError((e) => on.error(e));
    return {
      sendText: (text) => call.send(text),
      setMuted: (muted) => call.mute(muted),
      outputLevel: () => call.agentVolume(), // 0–1, drives the orb
      disconnect: () => call.end(),
    };
  },
};

<VoiceAgent layout="workspace" adapter={myAgent} />
```

A transcript line grows in place: send the same `id` again while the words arrive (`final: false`) and once more when
it is finished. The orb widens with `outputLevel` while the agent talks. Optional methods let a typed session turn into
a call without losing the reply in flight, and switch microphones mid-call. For dictation, the library also ships
speech-to-text from the browser's own recognition, with no dependency.

## The parts on their own

| part | what it is for |
|---|---|
| [Orb](https://beamline.io/components/orb) | the agent's presence: still at rest, an arc while thinking, widening with the voice; from a 16 px mark to a 240 px call screen |
| [Conversation bar](https://beamline.io/components/conversation-bar) | the bottom bar of a call: call and hang up, the agent's state as bars, a timer, mute, and typing when talking is not an option |
| [Voice button](https://beamline.io/components/voice-button) | start and stop talking: click to toggle or hold to talk, with a keyboard shortcut |
| [Mic selector](https://beamline.io/components/mic-selector) | choose and test the microphone, with every permission and device state spelled out |
| [Transcript viewer](https://beamline.io/components/transcript-viewer) | a recording and its words together: the spoken word lights up as it plays |
| [Speech input](https://beamline.io/components/speech-input) | a text field you can speak into |
| [Live waveform](https://beamline.io/components/live-waveform) | loudness drawn as it happens |

The rest of the set: [audio player](https://beamline.io/components/audio-player), [bar visualizer](https://beamline.io/components/bar-visualizer),
[waveform](https://beamline.io/components/waveform), [scrub bar](https://beamline.io/components/scrub-bar), [voice picker](https://beamline.io/components/voice-picker) and
[matrix](https://beamline.io/components/matrix).

## The states that make it trustworthy

- **The microphone tells the truth.** Not yet allowed, blocked (and how to unblock it), no microphone, unplugged: each
  has its own words and its own fix, instead of a button that silently does nothing.
- **State comes from your operation.** The voice button shows recording, finishing and done from what your code reports,
  never from a timer, so it cannot claim words were heard when they were not.
- **Typing is always there.** The conversation bar opens a text field when talking is not an option, and Escape closes
  it and returns focus.
- **The keyboard works the call.** Call, mute and the voice button are buttons; hold-to-talk works with Space, Enter or
  a shortcut you choose.

## When not to use it

A text-only assistant is the [agent workspace](https://beamline.io/components/agent-workspace) or a
[conversation](https://beamline.io/components/conversation) with [messages](https://beamline.io/components/message)
([the AI interface guide](https://beamline.io/guides/ai-agent-interface)). Dictating into a form is [speech input](https://beamline.io/components/speech-input).
People talking to people is the [chat thread](https://beamline.io/components/chat-thread).

## Ask your agent for it

With Beamline connected ([setup for your agent](https://beamline.io/connect)): "a voice support panel on our speech provider that also
takes typed questions, with a live transcript". The agent installs the block and writes the adapter around your
provider's SDK.
