Realtime voice agents with GPT-Realtime-2.1 and Parallel | Parallel

Introducing Parallel Search Turbo: the fastest and lowest-cost search for AI. Learn more.Learn more.

HumanMachine

July 12, 2026

\# Building a realtime voice agent with GPT-Realtime-2.1 and Parallel Search Turbo

A concise tutorial for building a realtime voice agent that uses GPT-Realtime-2.1 for conversation and Parallel Search Turbo when a question needs current web context.

Tags:Developers

Reading time: 4 min

Voice agents are most useful when they can answer questions about what is happening now. The hard part is adding web search without turning every turn into a slow, opaque research job.

This demo uses GPT-Realtime-2.1 for conversation and Parallel Search Turbo for current web context. The model decides whether a turn needs research. If it does, the app runs one Turbo search, shows the returned sources, and asks Realtime to answer from that evidence. Greetings and clarification questions can skip search entirely.

Try asking about tomorrow’s weather, the latest box office results, or planning a summer trip.

Parallel Voice Search Demo

Ask the web.

Tap the mic and ask a question.

Try asking

### GPT-Realtime-2.1 voice research demo

Interactive GPT-Realtime-2.1 voice research demo. In non-interactive or machine-readable contexts, try it at https://turbo-voice-research.vercel.app/embed.

Open interactive demo

\## How the handoff works

Realtime handles the conversation, Turbo searches the web, and the app controls the handoff between them.

  1. The user speaks. Semantic voice activity detection decides when the turn is complete.
  2. Realtime makes a routing decision: search the web, respond directly, or wait for clearer input.
  3. For a substantive factual question, search_web creates a self-contained question and a few targeted search queries. The server sends one request to Parallel Search Turbo.
  4. The app shows the returned sources and passes what Turbo found back into the same Realtime session.
  5. The app turns tools off and requests one short audio answer grounded in that evidence.

The two-pass design keeps the first decision silent, prevents the model from speaking before research finishes, and avoids unnecessary searches for conversational turns.

Reference docs:

\## 1. Create a short-lived Realtime session

Create the Realtime client secret on your server. The standard OpenAI API key should never reach the browser.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44

export async function POST(request: Request) {
  const response = await fetch(
    "https://api.openai.com/v1/realtime/client_secrets",
    {
      method: "POST",
      signal: AbortSignal.timeout(10_000),
      headers: {
        Authorization: "Bearer " + process.env.OPENAI_API_KEY,
        "Content-Type": "application/json",
        "OpenAI-Safety-Identifier": hashUser(request),
      },
      body: JSON.stringify({
        session: {
          type: "realtime",
          model: "gpt-realtime-2.1",
          instructions:
            "Wait for the client session configuration before responding.",
          reasoning: { effort: "low" },
          tools: [],
          tool_choice: "auto",
          audio: {
            input: {
              noise_reduction: { type: "far_field" },
              transcription: {
                language: "en",
                model: "gpt-realtime-whisper",
              },
              turn_detection: {
                type: "semantic_vad",
                eagerness: "medium",
                interrupt_response: true,
                create_response: true,
              },
            },
            output: { voice: "marin" },
          },
        },
      }),
    },
  );

const data = await response.json();
  return Response.json({ value: data.value }, { status: response.status });
}``` export async function POST(request: Request) {
  const response = await fetch(
    "https://api.openai.com/v1/realtime/client_secrets",
    {
      method: "POST",
      signal: AbortSignal.timeout(10_000),
      headers: {
        Authorization: "Bearer " + process.env.OPENAI_API_KEY,
        "Content-Type": "application/json",
        "OpenAI-Safety-Identifier": hashUser(request),
      },
      body: JSON.stringify({
        session: {
          type: "realtime",
          model: "gpt-realtime-2.1",
          instructions:
            "Wait for the client session configuration before responding.",
          reasoning: { effort: "low" },
          tools: [],
          tool_choice: "auto",
          audio: {
            input: {
              noise_reduction: { type: "far_field" },
              transcription: {
                language: "en",
                model: "gpt-realtime-whisper",
              },
              turn_detection: {
                type: "semantic_vad",
                eagerness: "medium",
                interrupt_response: true,
                create_response: true,
              },
            },
            output: { voice: "marin" },
          },
        },
      }),
    },
  );

const data = await response.json();
  return Response.json({ value: data.value }, { status: response.status });
}
```

The browser gives the agent and tool config after it receives this credential. The first automatic response is configured as a required, text-only routing pass, so Realtime can choose what to do without speaking too early.

\## 2. Connect with WebRTC

In the browser, request microphone access, create the Agents SDK WebRTC transport, and connect with the short-lived token.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43

import {
OpenAIRealtimeWebRTC,
RealtimeAgent,
RealtimeSession,
} from "@openai/agents/realtime";

async function connectRealtime() {
const media = await navigator.mediaDevices.getUserMedia({ audio: true });
const token = await fetch("/api/realtime-token", { method: "POST" })
    .then((response) => response.json());

const audio = document.createElement("audio");
audio.autoplay = true;

const transport = new OpenAIRealtimeWebRTC({
    mediaStream: media,
    audioElement: audio,
});

const agent = new RealtimeAgent({
    name: "Parallel voice assistant",
    instructions: voiceAgentInstructions(),
    tools: [searchWebTool, respondDirectlyTool, waitForUserTool],
});

const session = new RealtimeSession(agent, {
    model: "gpt-realtime-2.1",
    transport,
    config: {
      toolChoice: "required",
      parallelToolCalls: false,
      outputModalities: ["text"],
      reasoning: { effort: "low" },
    },
});

await session.connect({
    apiKey: token.value,
    model: "gpt-realtime-2.1",
});

return session;
}``` import {
OpenAIRealtimeWebRTC,
RealtimeAgent,
RealtimeSession,
} from "@openai/agents/realtime";

const audio = document.createElement("audio");
audio.autoplay = true;

const transport = new OpenAIRealtimeWebRTC({
    mediaStream: media,
    audioElement: audio,
});

await session.connect({
    apiKey: token.value,
    model: "gpt-realtime-2.1",
});

return session;
}
```

Only search_web starts web research. The other routes handle simple conversational responses or ignore audio that was not directed at the assistant.

\## 3. Search with Parallel Turbo

The search route turns the user’s request into a self-contained question and one to three searches. The server includes the resolved question in the final query list, adds the current date and time zone to the objective, and calls the Search API in turbo mode.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31

async function searchParallelTurbo(
  question: string,
  generatedQueries: string[],
) {
  const searchQueries = unique([\
    question,\
    ...generatedQueries,\
  ]).slice(0, 3);

const response = await fetch("https://api.parallel.ai/v1/search", {
    method: "POST",
    signal: AbortSignal.timeout(8_000),
    headers: {
      "Content-Type": "application/json",
      "x-api-key": process.env.PARALLEL_API_KEY!,
    },
    body: JSON.stringify({
      objective: buildSearchObjective(question),
      search_queries: searchQueries,
      mode: "turbo",
      max_chars_total: 10_000,
      advanced_settings: {
        max_results: 20,
        excerpt_settings: { max_chars_per_result: 1_000 },
      },
    }),
  });

if (!response.ok) throw new Error("Parallel Search failed.");
  return response.json();
}``` async function searchParallelTurbo(
  question: string,
  generatedQueries: string[],
) {
  const searchQueries = unique([\
    question,\
    ...generatedQueries,\
  ]).slice(0, 3);

if (!response.ok) throw new Error("Parallel Search failed.");
  return response.json();
}
```

The demo retrieves broadly, then passes a smaller, bounded evidence packet to Realtime. That gives the model enough context to answer without stuffing the full search response into the latency-sensitive audio turn.

Conversation history helps create better follow-ups. If someone asks about San Francisco weather and then says “what about tomorrow?”, Realtime carries the location forward and resolves the date before searching.

\## 4. Speak the grounded answer

When the search tool finishes, the app explicitly requests an audio response with tools disabled. The immediately preceding Turbo result is the evidence for factual claims in that answer.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20

session.on(
"agent_tool_end",
(_context, _agent, toolDefinition) => {
    if (toolDefinition.name !== "search_web") return;

session.transport.requestResponse({
      tools: [],
      tool_choice: "none",
      output_modalities: ["audio"],
      max_output_tokens: 800,
      reasoning: { effort: "low" },
      instructions: [\
        "Answer directly from the preceding Parallel Search output.",\
        "Use one or two complete spoken sentences under 60 words.",\
        "Do not mention tools, snippets, JSON, citations, or URLs.",\
        "Start with the answer.",\
      ].join("\n"),
    });
},
);``` session.on(
"agent_tool_end",
(_context, _agent, toolDefinition) => {
    if (toolDefinition.name !== "search_web") return;

session.transport.requestResponse({
      tools: [],
      tool_choice: "none",
      output_modalities: ["audio"],
      max_output_tokens: 800,
      reasoning: { effort: "low" },
      instructions: [\
        "Answer directly from the preceding Parallel Search output.",\
        "Use one or two complete spoken sentences under 60 words.",\
        "Do not mention tools, snippets, JSON, citations, or URLs.",\
        "Start with the answer.",\
      ].join("\n"),
    });
},
);
```

The sources appear in the UI as soon as search completes. Realtime then speaks the answer; the demo does not render a separate written answer transcript.

\## Latency notes

We don’t label the whole voice interaction as search latency. The full turn also includes voice activity detection, routing, answer generation, and audio playback.

The app records more detailed timings for debugging, but keeping these two numbers separate is enough to explain where users are waiting.

\## When to use this stack

This tech works well when a voice experience benefits from fast, sourced web context: current events, explainers, weather, travel planning, product research, and conversational follow-ups.

Turbo is a fast grounding step, not a full research pipeline. Use a specialized feed for authoritative live prices or scores, and a deeper research mode when the question needs broad retrieval or multi-hop reasoning.

\## Before you ship

The pattern is simple: Realtime decides how to handle the turn, Turbo grounds factual answers in current web sources, and the app controls the boundary between search and speech.

\## Ready to get started?

Sign up for free. No credit card required.

Try ParallelTry Parallel Contact salesContact sales

Are you an agent? Read this to onboard ParallelAre you an agent? Read this to onboard Parallel

By George Pickett

July 12, 2026

\## Related Posts80

Jul 30, 2026\ - Building an always-on background agent to proactively support customers

Tags: Developers

Author: By Khushi Shelat

Jul 21, 2026\ - Introducing the Parallel Responses API

Tags: Product

Author: By Parallel

Jul 20, 2026\ - Building a vendor intelligence system with Parallel

Tags: Developers

Author: By Sahith Jagarlamudi

Jul 16, 2026\ - Parallel and Google Cloud Announce Partnership for Agentic Web Search on Gemini Enterprise Agent Platform

Tags: Product

Author: By Parallel

Jul 15, 2026\ - $5 in free Parallel credits, every month

Tags: Product

Author: By Parallel

\ \ Jul 13, 2026\ \ - [Introducing Parallel Search Turbo](/content/blog/parallel-search-turbo/index.html) \ \ Author: By Parallel](/content/blog/parallel-search-turbo/index.html)

Jul 10, 2026\ - How Nooks cut web search costs 70.5% by switching to Parallel

Tags: Customers

Author: By Parallel

Jul 8, 2026\ - How Build created live geofenced alerts powered by Parallel for institutional real estate

Tags: Customers

Author: By Parallel

Jun 9, 2026\ - OpenClaw now has free, LLM-optimized web search by default powered by Parallel

Tags: Company

Author: By Parallel

Jun 5, 2026\ - Introducing real-time Entity Search

Tags: Product

Author: By Parallel

Jun 3, 2026\ - How we enrich & triage inbound leads using the Parallel Task API

Tags: Developers

Author: By Khushi Shelat

May 20, 2026\ - How AirOps creates citation-worthy content at scale, powered by Parallel

Tags: Customers

Author: By Parallel

May 18, 2026\ - Introducing Index by Parallel

Tags: Product

Author: By Parallel

May 7, 2026\ - Parallel Monitor API: New processor tiers, snapshots and event streams, and Basis on every event

Tags: Product

Author: By Parallel

May 4, 2026\ - How we built parallelmpp.dev

Tags: Developers

Author: By Son Do

Apr 29, 2026\ - How Actively's Per Account Agents use Parallel to turn the entire web into a proactive sales intelligence layer

Tags: Customers

Author: By Parallel

Apr 28, 2026\ - Parallel Raises at $2 Billion Valuation to Scale Web Infrastructure for Agents

Tags: Company

Author: By Parallel

Apr 24, 2026\ - Building a free CLI agent with Pi, Ollama, Gemma 4, and Parallel

Tags: Developers

Author: By Matt Harris

Apr 23, 2026\ - Parallel Search is now free for agents via MCP

Tags: Product

Author: By Parallel

Apr 21, 2026\ - Upgrades to the Parallel Search & Extract APIs

Tags: Benchmarks

Author: By Parallel

Apr 20, 2026\ - How Finch is scaling plaintiff law with AI agents that research like associates

Tags: Customers

Author: By Parallel

Apr 8, 2026\ - Genpact and Parallel Web Systems Partner to Drive Tangible Efficiency from AI Systems

Tags: Company

Author: By Parallel

Apr 8, 2026\ - How Genpact helps top US insurers cut contents claims processing times in half with Parallel

Tags: Customers

Author: By Parallel

Apr 7, 2026\ - A new deep research frontier on DeepSearchQA with the Task API Harness

Tags: Benchmarks

Author: By Parallel

Mar 30, 2026\ - How Modal saves tens of thousands annually by building in-house GTM pipelines with Parallel

Tags: Customers

Author: By Parallel

Mar 25, 2026\ - How Opendoor uses Parallel as the enterprise grade web research layer powering its AI-native real estate operations

Tags: Customers

Author: By Parallel

Mar 19, 2026\ - Introducing stateful web research agents with multi-turn conversations

Tags: Product

Author: By Parallel

Mar 18, 2026\ - Parallel is live on Tempo, now available natively to agents with the Machine Payments Protocol

Tags: Company

Author: By Parallel

Mar 17, 2026\ - How Parallel helped Kepler build AI that finance professionals can actually trust

Tags: Customers

Author: By Parallel

Mar 10, 2026\ - Introducing the Parallel CLI

Tags: Product

Author: By Parallel

Mar 4, 2026\ - How Profound helps brands win AI Search with high-quality web research and content creation powered by Parallel

Tags: Customers

Author: By Parallel

Mar 2, 2026\ - How Harvey is expanding legal AI internationally with Parallel

Tags: Customers

Author: By Parallel

Feb 23, 2026\ - How Tabstack by Mozilla enables agents to navigate the web with Parallel’s best-in-class web search

Tags: Customers

Author: By Parallel

Feb 4, 2026\ - Parallel Web Tools and Agents now available across Vercel AI Gateway, AI SDK, and Marketplace

Tags: Product

Author: By Parallel

Jan 28, 2026\ - Authenticated page access for the Parallel Task API

Tags: Product

Author: By Parallel

Jan 21, 2026\ - Introducing structured outputs for the Monitor API

Tags: Product

Author: By Parallel

Jan 15, 2026\ - Introducing research models with Basis for the Parallel Chat API

Tags: Product

Author: By Parallel

Jan 8, 2026\ - Build a real-time fact checker with Parallel and Cerebras

Tags: Developers

Author: By Parallel

Dec 17, 2025\ - Parallel Task API achieves state-of-the-art accuracy on DeepSearchQA

Tags: Benchmarks

Author: By Parallel

Dec 16, 2025\ - Introducing Granular Basis for the Task API

Tags: Product

Author: By Parallel

Dec 11, 2025\ - How Amp’s coding agents build better software with Parallel Search

Tags: Customers

Author: By Parallel

Dec 10, 2025\ - Latency improvements on the Parallel Task API

Tags: Product

Author: By Parallel

Nov 20, 2025\ - Introducing Parallel Extract

Tags: Product

Author: By Parallel

Nov 18, 2025\ - Introducing Parallel FindAll

Tags: Product,Benchmarks

Author: By Parallel

Nov 13, 2025\ - Introducing Parallel Monitor

Tags: Product

Author: By Parallel

Nov 12, 2025\ - Parallel raises $100M Series A to build web infrastructure for agents

Tags: Company

Author: By Parallel

Nov 11, 2025\ - How Macroscope reduced code review false positives with Parallel

Tags: Customers

Author: By Parallel

Nov 6, 2025\ - Introducing Parallel Search

Tags: Benchmarks

Author: By Parallel

Nov 3, 2025\ - Parallel processors set new price-performance standard on SealQA benchmark

Tags: Benchmarks

Author: By Parallel

Oct 30, 2025\ - Introducing LLMTEXT, an open source toolkit for the llms.txt standard

Tags: Product

Author: By Parallel

Oct 23, 2025\ - How Starbridge powers public sector GTM with state-of-the-art web research

Tags: Customers

Author: By Parallel

Oct 22, 2025\ - Building a market research platform with Parallel Deep Research

Tags: Developers

Author: By Parallel

Oct 17, 2025\ - How Lindy brings state-of-the-art web research to automation flows

Tags: Customers

Author: By Parallel

Oct 16, 2025\ - Introducing the Parallel Task MCP Server

Tags: Product

Author: By Parallel

Oct 9, 2025\ - Introducing the Core2x Processor for improved compute control on the Task API

Tags: Product

Author: By Parallel

Oct 8, 2025\ - How Day AI merges private and public data for business intelligence

Tags: Customers

Author: By Parallel

Oct 7, 2025\ - Full Basis framework for all Task API Processors

Tags: Product

Author: By Parallel

Oct 6, 2025\ - Building a real-time streaming task manager with Parallel

Tags: Developers

Author: By Parallel

Sep 30, 2025\ - How Gumloop built a new AI automation framework with web intelligence as a core node

Tags: Customers

Author: By Parallel

Sep 16, 2025\ - Introducing the TypeScript SDK

Tags: Product

Author: By Parallel

Sep 12, 2025\ - Building a serverless competitive intelligence platform with MCP + Task API

Tags: Developers

Author: By Parallel

Sep 11, 2025\ - Introducing Parallel Deep Research reports

Tags: Product

Author: By Parallel

Sep 9, 2025\ - A new pareto-frontier for Deep Research price-performance

Tags: Benchmarks

Author: By Parallel

Sep 5, 2025\ - Building a Full-Stack Search Agent with Parallel and Cerebras

Tags: Developers

Author: By Parallel

Aug 21, 2025\ - Webhooks for the Parallel Task API

Tags: Product

Author: By Parallel

Aug 14, 2025\ - Introducing Parallel: Web Search Infrastructure for AIs

Tags: Benchmarks,Product

Author: By Parallel

Aug 7, 2025\ - Introducing SSE for Task Runs

Tags: Product

Author: By Parallel

Aug 5, 2025\ - A new line of advanced Processors: Ultra2x, Ultra4x, and Ultra8x

Tags: Product

Author: By Parallel

Aug 4, 2025\ - Introducing Auto Mode for the Parallel Task API

Tags: Product

Author: By Parallel

Jul 31, 2025\ - A state-of-the-art search API purpose-built for agents

Tags: Benchmarks

Author: By Parallel

Jul 31, 2025\ - Parallel Search MCP Server in Devin

Tags: Product

Author: By Parallel

Jul 28, 2025\ - Introducing Tool Calling via MCP Servers

Tags: Product

Author: By Parallel

Jul 14, 2025\ - Introducing the Parallel Search MCP Server

Tags: Product

Author: By Parallel

Jul 8, 2025\ - Introducing Source Policy

Tags: Product

Author: By Parallel

Jul 2, 2025\ - The Parallel Task Group API

Tags: Product

Author: By Parallel

Jun 17, 2025\ - State of the Art Deep Research APIs

Tags: Benchmarks

Author: By Parallel

Jun 10, 2025\ - Parallel Search API is now available in alpha

Tags: Product

Author: By Parallel

May 29, 2025\ - Introducing the Parallel Chat API

Tags: Product

Author: By Parallel

May 16, 2025\ - Introducing Basis with Calibrated Confidences

Tags: Product

Author: By Parallel

Apr 24, 2025\ - Introducing the Parallel Task API

Tags: Product,Benchmarks

Author: By Parallel