
Sep 28, 2026
July 12, 2026
A concise tutorial for building a realtime voice agent that uses GPT-Realtime-2.1 for conversation and Parallel Search Turbo when a question needs current web context.
Voice agents are most useful when they can answer questions about what is happening now. The hard part is adding web search without turning every turn into a slow, opaque research job.
This demo uses GPT-Realtime-2.1 for conversation and Parallel Search Turbo for current web context. The model decides whether a turn needs research. If it does, the app runs one Turbo search, shows the returned sources, and asks Realtime to answer from that evidence. Greetings and clarification questions can skip search entirely.
Try asking about tomorrow’s weather, the latest box office results, or planning a summer trip.
### GPT-Realtime-2.1 voice research demo
Interactive GPT-Realtime-2.1 voice research demo. In non-interactive or machine-readable contexts, try it at https://turbo-voice-research.vercel.app/embed.
[Open interactive demo](https://turbo-voice-research.vercel.app/embed)
Realtime handles the conversation, Turbo searches the web, and the app controls the handoff between them.
`search_web` creates a self-contained question and a few targeted search queries. The server sends one request to Parallel Search Turbo.The two-pass design keeps the first decision silent, prevents the model from speaking before research finishes, and avoids unnecessary searches for conversational turns.
Reference docs:
Create the Realtime client secret on your server. The standard OpenAI API key should never reach the browser.
1234567891011121314151617181920212223242526272829303132333435363738394041424344export async function POST(request: Request) {
const response = await fetch(
"https://api.openai.com/v1/realtime/client_secrets",
{
method: "POST",
signal: AbortSignal.timeout(10_000),
headers: {
Authorization: "Bearer " + process.env.OPENAI_API_KEY,
"Content-Type": "application/json",
"OpenAI-Safety-Identifier": hashUser(request),
},
body: JSON.stringify({
session: {
type: "realtime",
model: "gpt-realtime-2.1",
instructions:
"Wait for the client session configuration before responding.",
reasoning: { effort: "low" },
tools: [],
tool_choice: "auto",
audio: {
input: {
noise_reduction: { type: "far_field" },
transcription: {
language: "en",
model: "gpt-realtime-whisper",
},
turn_detection: {
type: "semantic_vad",
eagerness: "medium",
interrupt_response: true,
create_response: true,
},
},
output: { voice: "marin" },
},
},
}),
},
);
const data = await response.json();
return Response.json({ value: data.value }, { status: response.status });
}``` export async function POST(request: Request) { const response = await fetch( "https://api.openai.com/v1/realtime/client_secrets", { method: "POST", signal: AbortSignal.timeout(10_000), headers: { Authorization: "Bearer " + process.env.OPENAI_API_KEY, "Content-Type": "application/json", "OpenAI-Safety-Identifier": hashUser(request), }, body: JSON.stringify({ session: { type: "realtime", model: "gpt-realtime-2.1", instructions: "Wait for the client session configuration before responding.", reasoning: { effort: "low" }, tools: [], tool_choice: "auto", audio: { input: { noise_reduction: { type: "far_field" }, transcription: { language: "en", model: "gpt-realtime-whisper", }, turn_detection: { type: "semantic_vad", eagerness: "medium", interrupt_response: true, create_response: true, }, }, output: { voice: "marin" }, }, }, }), }, ); const data = await response.json(); return Response.json({ value: data.value }, { status: response.status });}``` The browser gives the agent and tool config after it receives this credential. The first automatic response is configured as a required, text-only routing pass, so Realtime can choose what to do without speaking too early.
In the browser, request microphone access, create the Agents SDK WebRTC transport, and connect with the short-lived token.
12345678910111213141516171819202122232425262728293031323334353637383940414243import {
OpenAIRealtimeWebRTC,
RealtimeAgent,
RealtimeSession,
} from "@openai/agents/realtime";
async function connectRealtime() {
const media = await navigator.mediaDevices.getUserMedia({ audio: true });
const token = await fetch("/api/realtime-token", { method: "POST" })
.then((response) => response.json());
const audio = document.createElement("audio");
audio.autoplay = true;
const transport = new OpenAIRealtimeWebRTC({
mediaStream: media,
audioElement: audio,
});
const agent = new RealtimeAgent({
name: "Parallel voice assistant",
instructions: voiceAgentInstructions(),
tools: [searchWebTool, respondDirectlyTool, waitForUserTool],
});
const session = new RealtimeSession(agent, {
model: "gpt-realtime-2.1",
transport,
config: {
toolChoice: "required",
parallelToolCalls: false,
outputModalities: ["text"],
reasoning: { effort: "low" },
},
});
await session.connect({
apiKey: token.value,
model: "gpt-realtime-2.1",
});
return session;
}``` import { OpenAIRealtimeWebRTC, RealtimeAgent, RealtimeSession,} from "@openai/agents/realtime"; async function connectRealtime() { const media = await navigator.mediaDevices.getUserMedia({ audio: true }); const token = await fetch("/api/realtime-token", { method: "POST" }) .then((response) => response.json()); const audio = document.createElement("audio"); audio.autoplay = true; const transport = new OpenAIRealtimeWebRTC({ mediaStream: media, audioElement: audio, }); const agent = new RealtimeAgent({ name: "Parallel voice assistant", instructions: voiceAgentInstructions(), tools: [searchWebTool, respondDirectlyTool, waitForUserTool], }); const session = new RealtimeSession(agent, { model: "gpt-realtime-2.1", transport, config: { toolChoice: "required", parallelToolCalls: false, outputModalities: ["text"], reasoning: { effort: "low" }, }, }); await session.connect({ apiKey: token.value, model: "gpt-realtime-2.1", }); return session;}``` Only `search_web` starts web research. The other routes handle simple conversational responses or ignore audio that was not directed at the assistant.
The search route turns the user’s request into a self-contained question and one to three searches. The server includes the resolved question in the final query list, adds the current date and time zone to the objective, and calls the Search API in turbo mode.
12345678910111213141516171819202122232425262728293031async function searchParallelTurbo(
question: string,
generatedQueries: string[],
) {
const searchQueries = unique([
question,
...generatedQueries,
]).slice(0, 3);
const response = await fetch("https://api.parallel.ai/v1/search", {
method: "POST",
signal: AbortSignal.timeout(8_000),
headers: {
"Content-Type": "application/json",
"x-api-key": process.env.PARALLEL_API_KEY!,
},
body: JSON.stringify({
objective: buildSearchObjective(question),
search_queries: searchQueries,
mode: "turbo",
max_chars_total: 10_000,
advanced_settings: {
max_results: 20,
excerpt_settings: { max_chars_per_result: 1_000 },
},
}),
});
if (!response.ok) throw new Error("Parallel Search failed.");
return response.json();
}``` async function searchParallelTurbo( question: string, generatedQueries: string[],) { const searchQueries = unique([ question, ...generatedQueries, ]).slice(0, 3); const response = await fetch("https://api.parallel.ai/v1/search", { method: "POST", signal: AbortSignal.timeout(8_000), headers: { "Content-Type": "application/json", "x-api-key": process.env.PARALLEL_API_KEY!, }, body: JSON.stringify({ objective: buildSearchObjective(question), search_queries: searchQueries, mode: "turbo", max_chars_total: 10_000, advanced_settings: { max_results: 20, excerpt_settings: { max_chars_per_result: 1_000 }, }, }), }); if (!response.ok) throw new Error("Parallel Search failed."); return response.json();}``` The demo retrieves broadly, then passes a smaller, bounded evidence packet to Realtime. That gives the model enough context to answer without stuffing the full search response into the latency-sensitive audio turn.
Conversation history helps create better follow-ups. If someone asks about San Francisco weather and then says “what about tomorrow?”, Realtime carries the location forward and resolves the date before searching.
When the search tool finishes, the app explicitly requests an audio response with tools disabled. The immediately preceding Turbo result is the evidence for factual claims in that answer.
1234567891011121314151617181920session.on(
"agent_tool_end",
(_context, _agent, toolDefinition) => {
if (toolDefinition.name !== "search_web") return;
session.transport.requestResponse({
tools: [],
tool_choice: "none",
output_modalities: ["audio"],
max_output_tokens: 800,
reasoning: { effort: "low" },
instructions: [
"Answer directly from the preceding Parallel Search output.",
"Use one or two complete spoken sentences under 60 words.",
"Do not mention tools, snippets, JSON, citations, or URLs.",
"Start with the answer.",
].join("\n"),
});
},
);``` session.on( "agent_tool_end", (_context, _agent, toolDefinition) => { if (toolDefinition.name !== "search_web") return; session.transport.requestResponse({ tools: [], tool_choice: "none", output_modalities: ["audio"], max_output_tokens: 800, reasoning: { effort: "low" }, instructions: [ "Answer directly from the preceding Parallel Search output.", "Use one or two complete spoken sentences under 60 words.", "Do not mention tools, snippets, JSON, citations, or URLs.", "Start with the answer.", ].join("\n"), }); },);``` The sources appear in the UI as soon as search completes. Realtime then speaks the answer; the demo does not render a separate written answer transcript.
We don’t label the whole voice interaction as search latency. The full turn also includes voice activity detection, routing, answer generation, and audio playback.
The app records more detailed timings for debugging, but keeping these two numbers separate is enough to explain where users are waiting.
This tech works well when a voice experience benefits from fast, sourced web context: current events, explainers, weather, travel planning, product research, and conversational follow-ups.
Turbo is a fast grounding step, not a full research pipeline. Use a specialized feed for authoritative live prices or scores, and a deeper research mode when the question needs broad retrieval or multi-hop reasoning.
The pattern is simple: Realtime decides how to handle the turn, Turbo grounds factual answers in current web sources, and the app controls the boundary between search and speech.
Sign up for free. No credit card required.
By George Pickett
July 12, 2026

Sep 28, 2026

Sep 28, 2026

Sep 24, 2026

Sep 23, 2026

Sep 18, 2026

Sep 15, 2026

Sep 10, 2026

Aug 31, 2026

Aug 29, 2026

Aug 25, 2026

Aug 21, 2026

Aug 19, 2026

Aug 13, 2026

Jul 30, 2026

Jul 21, 2026

Jul 20, 2026

Jul 16, 2026

Jul 15, 2026

Jul 13, 2026

Jul 10, 2026

Jul 8, 2026

Jun 9, 2026

Jun 5, 2026

Jun 4, 2026

May 20, 2026

May 19, 2026

May 7, 2026

May 5, 2026

Apr 29, 2026

Apr 29, 2026

Apr 24, 2026

Apr 23, 2026

Apr 21, 2026

Apr 20, 2026

Apr 8, 2026

Apr 8, 2026

Apr 7, 2026

Mar 30, 2026

Mar 25, 2026

Mar 19, 2026

Mar 18, 2026

Mar 17, 2026

Mar 10, 2026

Mar 4, 2026

Mar 2, 2026

Feb 23, 2026

Feb 4, 2026

Jan 28, 2026

Jan 21, 2026

Jan 15, 2026

Jan 8, 2026

Dec 17, 2025

Dec 16, 2025

Dec 11, 2025

Dec 10, 2025

Nov 20, 2025

Nov 18, 2025

Nov 13, 2025

Nov 12, 2025

Nov 11, 2025

Nov 6, 2025

Nov 3, 2025

Oct 30, 2025

Oct 23, 2025

Oct 22, 2025

Oct 17, 2025

Oct 16, 2025

Oct 9, 2025

Oct 8, 2025

Oct 7, 2025

Oct 6, 2025

Sep 30, 2025

Sep 16, 2025

Sep 12, 2025

Sep 11, 2025

Sep 9, 2025

Sep 5, 2025

Aug 21, 2025

Aug 14, 2025

Aug 7, 2025

Aug 5, 2025

Aug 4, 2025

Jul 31, 2025

Jul 31, 2025

Jul 28, 2025

Jul 14, 2025

Jul 8, 2025

Jul 2, 2025

Jun 17, 2025

Jun 10, 2025

May 30, 2025

May 16, 2025

Apr 24, 2025