Skip to main content
Kupe Realtime is an OpenAI-shaped voice WebSocket. You mint an ephemeral session over HTTP, then attach to wss://x.kupe.in/agents/v1/realtime (same agents host as telephony media — not the LiveKit root on x.kupe.in). The session always runs Kupe STT, TTS, and LLM. Voices are addressed by sanitized name (priya) or voice id (copy it from the voice library). Tools attached to the agent run server-side. This path is web-only — it does not write telephony minutes.

Run it in a terminal

Same snippets live on the Python and TypeScript SDK tabs.

Mint

POST /v1/realtime/sessions Identify the agent with either name or agent_id (copy the id from the agent editor). If name does not exist in the project, Kupe creates the agent with the prompt, greeting, voice, and tools you pass.
string
Agent name. Reuses a live agent with this name, or creates one.
string
Existing agent id. Alias: id. Optional when name is set.
string
Voice name or voice id. Alias: voice_id. Overrides the agent’s voice for this session.
string
Alias for voice — library voice_id or row id.
string
System prompt. Alias: instructions. Saved on a newly created agent; overlays this session for an existing one.
string
Spoken first message. Alias: greetings.
array
OpenAI-style function tools, or webhook/MCP tools (http_url, type: "mcp").
object
MCP server: { url, headers?, tools? } where tools are names or full defs. Same as passing MCP items in tools.
object
Values for {{placeholders}} in the prompt and greeting.
string
Optional when using an API key (GET /v1/me already has it).
string
Optional when using an API key.
The response includes client_secret.value (single-use ticket), websocket_url, and the resolved agent_id.
You can pass OpenAI-style tools (or MCP) on the same mint call. Webhook/MCP tools are saved on a newly created agent; function-only tools run for this session.

Connect

You can also send Authorization: Bearer {secret} on the WebSocket handshake instead of the query string. ticket is an alias for client_secret. The SDK helper does this for you:
On accept, the server sends session.created with instructions, voice, tools, and pcm16 in/out formats.

Client → server

Text-turn example (what send_text encodes):
Mic path: send input_audio_buffer.append frames; server VAD emits speech start/stop and runs the turn.
Do not send the agent’s own audio back. If you play response.output_audio.delta through open speakers next to the mic, your input_audio_buffer.append frames will contain the agent’s voice. Server VAD treats that as the caller speaking, so the agent interrupts and answers itself, and its own lines appear as user transcripts.Use input that already has acoustic echo cancellation (a headset, a phone line, or getUserMedia({ audio: { echoCancellation: true } }) in the browser), or send silence in input_audio_buffer.append while the agent’s audio is still playing. Track that from the bytes you have queued for playback, not from the arrival of the last delta — deltas arrive faster than realtime.Keep sending frames: server VAD and STT run over a continuous stream, so simply stopping the appends stalls turn detection and the agent stops hearing the caller even after it finishes speaking. Muting this way also means the caller cannot barge in.Both SDKs implement this: echo_suppression="half_duplex" in Python, echoSuppression: "half_duplex" in TypeScript.

Server → client

Every event has event_id and type. Audio deltas are base64 PCM16. The socket may also send dashboard-shaped { "kind": "transcript" | "latency" | "tool_call", ... } frames on the same connection.

Close codes

LiveKit web sessions

For the in-app web tester (room + JWT, not this WS), use POST /v1/sessions with channel: "web". That returns a LiveKit ws_url and participant token. Realtime mint + /v1/realtime is the API you want from Python, TypeScript, or cURL.