A new deep research frontier on DeepSearchQA with the Task API Harness | Parallel

Introducing Parallel Search Turbo: the fastest and lowest-cost search for AI. Learn more.Learn more.

HumanMachine

April 7, 2026

\# A new deep research frontier on DeepSearchQA with the Task API Harness

Tags:Benchmarks

Reading time: 7 min

The Parallel Task API is the most powerful deep research agent on the market. It’s used in production agents by teams at Opendoor, Attio, Modal, Starbridge, Profound, and more.

Parallel’s ProcessorProcessor architecture allows for flexibility and fine-tuned application of different compute budgets depending on the complexity of a research task. The “Ultra” range of Task API Processors is state-of-the-art on DeepSearchQA:

This blog details some of the techniques we’ve employed to achieve this state of the art research quality.

\## Results

We evaluated Parallel's Task API Processors against GPT 5.4, Opus 4-6, Gemini 3.1 Pro, Exa Search Deep Reasoning, and Perplexity Sonar Pro.

Provider Model Cost (CPM) Accuracy (%)
Parallel Ultra 8x $2400 82
Parallel Ultra 4x $1200 81
Parallel Ultra 2x $600 77
Parallel Ultra $300 70
OpenAI GPT 5.4 with code execution $701 63
Google Gemini 3.1 Pro with code execution $707 62
Anthropic Opus 4-6 with PTC $36,231* 58
Perplexity Sonar Pro $883 28
Exa Search Deep Reasoning $15 18

_CPM: USD per 1,000 requests. Cost is shown on a log scale._

_*The cost of Opus 4-6 is higher than expected due to Anthropic’s potential billing issue, where prompt caching savings for PTC are not passed on to the user._

\## About DeepSearchQA

DeepSearchQA is a 900-question evaluation from Google designed to test agents on multi-step information-seeking tasks across 17 fields of expertise. Each question is a causal chain: you can't answer the second part without resolving the first, and you can't resolve the first without searching the web, reading the results, and reasoning about what to search next.

A typical query might ask: _identify every researcher who co-authored papers with a specific professor at three different institutions over a decade, then determine which of those co-authors later joined a federal advisory committee._ Simple web retrieval won't cut it. The agent needs to plan a research strategy, execute multiple searches, cross-reference results across sources, and synthesize a precise answer with zero false positives.

Here, accuracy on DeepSearchQA means "fully correct": the response must be semantically identical to the ground-truth set. The agent has to find all correct answers while including none that are wrong.

_*We use DeepSearchQA instead of BrowseComp to ensure more reliable evaluation, as some models have begun to memorize portions of the BrowseComp dataset, potentially inflating performance._

\## Inside the Task API Harness

Visualization">https://cdn.sanity.io/images/5hzduz3y/production/af4801e37c89acb2b3dfa0002ef6ead13dea4c4c-900x600.gif)Visualization of the Parallel Task API Harness

Our results stem from combining several techniques.

Together, these techniques allow us to scale reasoning depth and reliability while maintaining efficiency, leading to state-of-the-art results on DeepSearchQA. Instead of having the model orchestrate tools through tool calling, we gave it the ability to write and execute code.

The orchestrating model generates Python that calls research tools as ordinary functions. This code runs in a sandboxed interpreter. Only the final output of each code block re-enters the model's context. Intermediate data stays in the interpreter's variable state, not in the conversation history.

Task">https://cdn.sanity.io/images/5hzduz3y/production/dfb06ffe55ea24e567f370fa1ce0b26be8b4d772-1920x1004.png)Task API high-level architecture

1
2
3
4
5

Step 1: search("company X revenue 2024") → [5 results in context]
Step 2: extract(url_1) → [full page content in context]
Step 3: extract(url_2) → [full page content in context]
Step 4: search("company X revenue 2023") → [5 more results in context]
...context grows with every step``` Step 1: search("company X revenue 2024") → [5 results in context]
Step 2: extract(url_1) → [full page content in context]
Step 3: extract(url_2) → [full page content in context]
Step 4: search("company X revenue 2023") → [5 more results in context]
...context grows with every step
```

Each step inflates the context window. By step 8, the model spends most of its capacity re-reading old results.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19

# First, use one search step to discover the URL pattern.
results_2024 = search("Company X 2024 annual report")
report_2024 = results_2024[0].url
# e.g. https://investors.companyx.com/financials/annual-reports/2024-annual-report.pdf

# Infer a reusable template from the 2024 URL and a parse function parse_fn.
url_template = "https://investors.companyx.com/financials/annual-reports/{year}-annual-report.pdf"

years = [2022, 2023, 2024]
reports = {}

for year in years:
    url = url_template.format(year=year)
    reports[year] = parse_fn(extract(
        question=f"What was Company X's reported revenue in {year}? Return the value and supporting quote.",
        urls=[url],
    ))

return reports``` # First, use one search step to discover the URL pattern.
results_2024 = search("Company X 2024 annual report")
report_2024 = results_2024[0].url
# e.g. https://investors.companyx.com/financials/annual-reports/2024-annual-report.pdf

years = [2022, 2023, 2024]
reports = {}

return reports
```

Multiple searches, extractions, and analyses happen in a single execution step. The full page content of both extracted pages (potentially tens of thousands of tokens) never enters the research agent's context. Only the extracted revenue figures flow back.

This compounds. A 20-step research task that would fill a 128K context window under tool calling stays under 30K tokens because intermediate data lives in the interpreter, not the conversation.

\### Persistent state as working memory

Variables created in one code execution step survive to the next. When the model writes _findings["report"] = report_2024_ in iteration 3, that variable is still accessible in iteration 7 when the model needs to cross-reference revenue against employee headcount.

This creates a separation that matters: the conversation history captures the model's high-level reasoning trajectory, while the interpreter's variable state captures the raw data it has gathered. The conversation can be compacted without losing granular data.

\### Inside the sandbox

Running LLM-generated code in production requires strong isolation. Our interpreter is a sandboxed Python runtime built on Rust with no access to the network, filesystem, or operating system. The only way code interacts with the outside world is through explicitly injected functions: search, extract and a handful of state management utilities.

Sandbox">https://cdn.sanity.io/images/5hzduz3y/production/614565599ea2bfe5f962306c615fca92bc6e82a5-2080x1144.png)Sandbox architecture

This boundary gives us two things. First, the model can write complex data processing logic (filtering, aggregation, string manipulation, conditional branching) without risk of side effects.

\### Budget-aware execution

One of the attractive features of Parallel’s Task API is fixed and predictable costs rather than per-token pricing. We have built our agent harness to support budgets as a first-class concept. This budget isn’t just a fixed limit on the number of steps. A fixed step limit penalizes simple and complex queries equally. We track cumulative cost across iterations instead. The system monitors token spend across all LLM calls, both the orchestrating model and any sub-model invocations, and injects budget warnings when remaining spend drops below a threshold. When the budget is nearly exhausted, the model synthesizes its findings and produces a final answer with whatever it has gathered.

A simple factual query might terminate in 2 iterations and cost a few cents. A complex multi-source comparison might run for 15 iterations across a larger budget. The architecture adapts to the difficulty of the question rather than imposing a uniform ceiling. This is also how we offer the full range of Processors from Ultra ($300 CPM) through Ultra8x ($2,400 CPM): higher-tier Processors get a larger budget, which lets the agent pursue more research paths before synthesizing.

\### Context compaction

Even with code execution keeping intermediate data out of context, long research sessions accumulate history: the model's reasoning, code blocks, and execution summaries. When this history approaches context limits, we trigger compaction, a summarization pass that condenses earlier conversation turns while preserving key findings and the current research trajectory.

The persistent variable state in the interpreter is unaffected by compaction (it lives in the interpreter, not the conversation). This means the model can sustain research across many more iterations than its raw context window would suggest.

\## Why we picked this architecture

Most deep research systems follow a familiar loop: an LLM generates a plan, calls a tool, reads the result, and repeats. This works for simple queries but degrades as research complexity grows. We evaluated several architectural patterns for complex research workflows, each with different trade-offs across context efficiency, granularity of information access, and adaptability.

Architecture Cost efficiency Fine-grained Research Control Adaptability Reason
Naive Agent Loop (LLM → tool → result → LLM → …) Maximally flexible, but context grows at every step. The model ends up spending too much capacity re-reading intermediate results.
Agent Loop with context compression 🟠 Keeps context manageable, but important details are often compressed away. This makes it harder to revisit evidence or pursue subtle lines of inquiry.
Static Plan + sub-agents 🟠 Works well for predictable workflows, but cannot adapt cleanly once the plan is set.
Agent Loop with sub-agents 🟠 🟠 Better runtime flexibility than static planning, but still suffers from coordination and memory fragmentation across agents.
Task API Harness (Parallel) Preserves detailed evidence access while keeping the intermediate state out of the model context. It can adapt mid-run, branch when needed, and scale without bloating the prompt.
1
2
3
4
5
6
7
8
9
10
11
12
13
14

import parallel

client = parallel.Client(api_key="your-api-key")

task = client.task_runs.create(
    objective="Identify every researcher who co-authored papers "
              "with Dr. Maria Chen at Stanford, MIT, and Caltech "
              "between 2010 and 2020, then determine which of those "
              "co-authors later joined the NIH Advisory Committee "
              "to the Director.",
    processor="ultra8x",
)

print(task.output)``` import parallel

client = parallel.Client(api_key="your-api-key")

print(task.output)
```

\## About the Parallel Task API

The Task API is a general-purpose web research agent API. Define what you need in natural language or structured JSON, and it handles research, synthesis, and structured output with citations and confidence levels. Processors range from Lite (basic lookups, $5/1K) through Ultra8x (the hardest deep research, $2,400/1K).

\## About Parallel Web Systems

Parallel builds web infrastructure for AI. Our APIs, including Search, Extract, Task, FindAll, and Monitor, give AI agents structured, grounded access to the open web, powered by a rapidly growing proprietary index of the global internet.

Parallel turns human workflows that took days into agentic workflows that take seconds. Fortune 100s and leading frontier AI companies including Harvey, Manus, Starbridge, and Profound rely on Parallel for legal grounding, fact-checking, contract monitoring, and high-quality content generation.

Get started at platform.parallel.aiplatform.parallel.ai

\## Ready to get started?

Sign up for free. No credit card required.

Try ParallelTry Parallel Contact salesContact sales

Are you an agent? Read this to onboard ParallelAre you an agent? Read this to onboard Parallel

By Parallel

April 7, 2026

\## Related Posts80

Jul 30, 2026\ - Building an always-on background agent to proactively support customers

Tags: Developers

Author: By Khushi Shelat

Jul 21, 2026\ - Introducing the Parallel Responses API

Tags: Product

Author: By Parallel

Jul 20, 2026\ - Building a vendor intelligence system with Parallel

Tags: Developers

Author: By Sahith Jagarlamudi

Jul 16, 2026\ - Parallel and Google Cloud Announce Partnership for Agentic Web Search on Gemini Enterprise Agent Platform

Tags: Product

Author: By Parallel

Jul 15, 2026\ - $5 in free Parallel credits, every month

Tags: Product

Author: By Parallel

\ \ Jul 13, 2026\ \ - [Introducing Parallel Search Turbo](/content/blog/parallel-search-turbo/index.html) \ \ Author: By Parallel](/content/blog/parallel-search-turbo/index.html)

Jul 12, 2026\ - Building a realtime voice agent with GPT-Realtime-2.1 and Parallel Search Turbo

Tags: Developers

Author: By George Pickett

Jul 10, 2026\ - How Nooks cut web search costs 70.5% by switching to Parallel

Tags: Customers

Author: By Parallel

Jul 8, 2026\ - How Build created live geofenced alerts powered by Parallel for institutional real estate

Tags: Customers

Author: By Parallel

Jun 9, 2026\ - OpenClaw now has free, LLM-optimized web search by default powered by Parallel

Tags: Company

Author: By Parallel

Jun 5, 2026\ - Introducing real-time Entity Search

Tags: Product

Author: By Parallel

Jun 3, 2026\ - How we enrich & triage inbound leads using the Parallel Task API

Tags: Developers

Author: By Khushi Shelat

May 20, 2026\ - How AirOps creates citation-worthy content at scale, powered by Parallel

Tags: Customers

Author: By Parallel

May 18, 2026\ - Introducing Index by Parallel

Tags: Product

Author: By Parallel

May 7, 2026\ - Parallel Monitor API: New processor tiers, snapshots and event streams, and Basis on every event

Tags: Product

Author: By Parallel

May 4, 2026\ - How we built parallelmpp.dev

Tags: Developers

Author: By Son Do

Apr 29, 2026\ - How Actively's Per Account Agents use Parallel to turn the entire web into a proactive sales intelligence layer

Tags: Customers

Author: By Parallel

Apr 28, 2026\ - Parallel Raises at $2 Billion Valuation to Scale Web Infrastructure for Agents

Tags: Company

Author: By Parallel

Apr 24, 2026\ - Building a free CLI agent with Pi, Ollama, Gemma 4, and Parallel

Tags: Developers

Author: By Matt Harris

Apr 23, 2026\ - Parallel Search is now free for agents via MCP

Tags: Product

Author: By Parallel

Apr 21, 2026\ - Upgrades to the Parallel Search & Extract APIs

Tags: Benchmarks

Author: By Parallel

Apr 20, 2026\ - How Finch is scaling plaintiff law with AI agents that research like associates

Tags: Customers

Author: By Parallel

Apr 8, 2026\ - Genpact and Parallel Web Systems Partner to Drive Tangible Efficiency from AI Systems

Tags: Company

Author: By Parallel

Apr 8, 2026\ - How Genpact helps top US insurers cut contents claims processing times in half with Parallel

Tags: Customers

Author: By Parallel

Mar 30, 2026\ - How Modal saves tens of thousands annually by building in-house GTM pipelines with Parallel

Tags: Customers

Author: By Parallel

Mar 25, 2026\ - How Opendoor uses Parallel as the enterprise grade web research layer powering its AI-native real estate operations

Tags: Customers

Author: By Parallel

Mar 19, 2026\ - Introducing stateful web research agents with multi-turn conversations

Tags: Product

Author: By Parallel

Mar 18, 2026\ - Parallel is live on Tempo, now available natively to agents with the Machine Payments Protocol

Tags: Company

Author: By Parallel

Mar 17, 2026\ - How Parallel helped Kepler build AI that finance professionals can actually trust

Tags: Customers

Author: By Parallel

Mar 10, 2026\ - Introducing the Parallel CLI

Tags: Product

Author: By Parallel

Mar 4, 2026\ - How Profound helps brands win AI Search with high-quality web research and content creation powered by Parallel

Tags: Customers

Author: By Parallel

Mar 2, 2026\ - How Harvey is expanding legal AI internationally with Parallel

Tags: Customers

Author: By Parallel

Feb 23, 2026\ - How Tabstack by Mozilla enables agents to navigate the web with Parallel’s best-in-class web search

Tags: Customers

Author: By Parallel

Feb 4, 2026\ - Parallel Web Tools and Agents now available across Vercel AI Gateway, AI SDK, and Marketplace

Tags: Product

Author: By Parallel

Jan 28, 2026\ - Authenticated page access for the Parallel Task API

Tags: Product

Author: By Parallel

Jan 21, 2026\ - Introducing structured outputs for the Monitor API

Tags: Product

Author: By Parallel

Jan 15, 2026\ - Introducing research models with Basis for the Parallel Chat API

Tags: Product

Author: By Parallel

Jan 8, 2026\ - Build a real-time fact checker with Parallel and Cerebras

Tags: Developers

Author: By Parallel

Dec 17, 2025\ - Parallel Task API achieves state-of-the-art accuracy on DeepSearchQA

Tags: Benchmarks

Author: By Parallel

Dec 16, 2025\ - Introducing Granular Basis for the Task API

Tags: Product

Author: By Parallel

Dec 11, 2025\ - How Amp’s coding agents build better software with Parallel Search

Tags: Customers

Author: By Parallel

Dec 10, 2025\ - Latency improvements on the Parallel Task API

Tags: Product

Author: By Parallel

Nov 20, 2025\ - Introducing Parallel Extract

Tags: Product

Author: By Parallel

Nov 18, 2025\ - Introducing Parallel FindAll

Tags: Product,Benchmarks

Author: By Parallel

Nov 13, 2025\ - Introducing Parallel Monitor

Tags: Product

Author: By Parallel

Nov 12, 2025\ - Parallel raises $100M Series A to build web infrastructure for agents

Tags: Company

Author: By Parallel

Nov 11, 2025\ - How Macroscope reduced code review false positives with Parallel

Tags: Customers

Author: By Parallel

Nov 6, 2025\ - Introducing Parallel Search

Tags: Benchmarks

Author: By Parallel

Nov 3, 2025\ - Parallel processors set new price-performance standard on SealQA benchmark

Tags: Benchmarks

Author: By Parallel

Oct 30, 2025\ - Introducing LLMTEXT, an open source toolkit for the llms.txt standard

Tags: Product

Author: By Parallel

Oct 23, 2025\ - How Starbridge powers public sector GTM with state-of-the-art web research

Tags: Customers

Author: By Parallel

Oct 22, 2025\ - Building a market research platform with Parallel Deep Research

Tags: Developers

Author: By Parallel

Oct 17, 2025\ - How Lindy brings state-of-the-art web research to automation flows

Tags: Customers

Author: By Parallel

Oct 16, 2025\ - Introducing the Parallel Task MCP Server

Tags: Product

Author: By Parallel

Oct 9, 2025\ - Introducing the Core2x Processor for improved compute control on the Task API

Tags: Product

Author: By Parallel

Oct 8, 2025\ - How Day AI merges private and public data for business intelligence

Tags: Customers

Author: By Parallel

Oct 7, 2025\ - Full Basis framework for all Task API Processors

Tags: Product

Author: By Parallel

Oct 6, 2025\ - Building a real-time streaming task manager with Parallel

Tags: Developers

Author: By Parallel

Sep 30, 2025\ - How Gumloop built a new AI automation framework with web intelligence as a core node

Tags: Customers

Author: By Parallel

Sep 16, 2025\ - Introducing the TypeScript SDK

Tags: Product

Author: By Parallel

Sep 12, 2025\ - Building a serverless competitive intelligence platform with MCP + Task API

Tags: Developers

Author: By Parallel

Sep 11, 2025\ - Introducing Parallel Deep Research reports

Tags: Product

Author: By Parallel

Sep 9, 2025\ - A new pareto-frontier for Deep Research price-performance

Tags: Benchmarks

Author: By Parallel

Sep 5, 2025\ - Building a Full-Stack Search Agent with Parallel and Cerebras

Tags: Developers

Author: By Parallel

Aug 21, 2025\ - Webhooks for the Parallel Task API

Tags: Product

Author: By Parallel

Aug 14, 2025\ - Introducing Parallel: Web Search Infrastructure for AIs

Tags: Benchmarks,Product

Author: By Parallel

Aug 7, 2025\ - Introducing SSE for Task Runs

Tags: Product

Author: By Parallel

Aug 5, 2025\ - A new line of advanced Processors: Ultra2x, Ultra4x, and Ultra8x

Tags: Product

Author: By Parallel

Aug 4, 2025\ - Introducing Auto Mode for the Parallel Task API

Tags: Product

Author: By Parallel

Jul 31, 2025\ - A state-of-the-art search API purpose-built for agents

Tags: Benchmarks

Author: By Parallel

Jul 31, 2025\ - Parallel Search MCP Server in Devin

Tags: Product

Author: By Parallel

Jul 28, 2025\ - Introducing Tool Calling via MCP Servers

Tags: Product

Author: By Parallel

Jul 14, 2025\ - Introducing the Parallel Search MCP Server

Tags: Product

Author: By Parallel

Jul 8, 2025\ - Introducing Source Policy

Tags: Product

Author: By Parallel

Jul 2, 2025\ - The Parallel Task Group API

Tags: Product

Author: By Parallel

Jun 17, 2025\ - State of the Art Deep Research APIs

Tags: Benchmarks

Author: By Parallel

Jun 10, 2025\ - Parallel Search API is now available in alpha

Tags: Product

Author: By Parallel

May 29, 2025\ - Introducing the Parallel Chat API

Tags: Product

Author: By Parallel

May 16, 2025\ - Introducing Basis with Calibrated Confidences

Tags: Product

Author: By Parallel

Apr 24, 2025\ - Introducing the Parallel Task API

Tags: Product,Benchmarks

Author: By Parallel