Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
All benchmarks

WideSearch

WideSearch asks the model to fill in an entire table: given a query and a required column schema, it has to find every entity and every attribute, like all launches of a rocket family with dates and outcomes. Its 200 tasks (100 English, 100 Chinese) are graded strictly — a task counts only when every row and every cell matches the reference table. We run it with the model held fixed and vary the search engine, request format, and maximum search budget. Where BrowseComp goes deep on one hidden fact, WideSearch goes wide across many facts and attributes.

Last benchmark run Aug 11, 2026, 11:09 AM UTC

PaperProjectDataset
Search providers
Favicon for Perplexity
Favicon for openai
Favicon for Exa
Favicon for Parallel
4 search providers
compared across configurations
Models
Winning resultQuality & completionPrice & speedSearch budgetAll configurationsWhy we run itWhat scores tell youHow tasks are scoredMethodology
01

Winning result

The configuration that scored highest, the breakdown behind that score, and the strongest value and speed alternatives.

Highest quality

Favicon for Perplexity
Perplexity
·25-turn
Favicon for openai
GPT-5.6 Sol · high
84.0%answer-item accuracy

Winning result in detail

Favicon for openai
Favicon for openai
Favicon for anthropic
Favicon for deepseek
4 models
GPT-5.6 Luna, GPT-5.6 Sol, Claude Opus 5, DeepSeek V4 Flash 0731
Test configurations
1 / 5 / 25
search budgets and plugin mode
Complete tables
21.0% · 21 of 100
Incomplete tables
79 of 100
Cost / question
$0.83
Latency / question
2.3m
Best value
$0.063 / question
Favicon for Perplexity
Perplexity
Favicon for openai
GPT-5.6 Luna · xhigh
5-turn search budget · 80.0% answer-item accuracy
Fastest strong result
1.9m / question
Favicon for Perplexity
Perplexity
Favicon for openai
GPT-5.6 Sol · high
1-turn search budget · 79.8% answer-item accuracy
02

Quality and completion

Answer-item accuracy is primary; complete-table success is shown beneath.

200 tasks · 100 English · 100 Chinese
Favicon for Perplexity
Perplexity
GPT-5.6 Sol · high · 25-turn
84.0%21.0% complete tables
Favicon for openai
OpenAI Native
GPT-5.6 Sol · high · 25-turn
81.8%16.0% complete tables
Favicon for Parallel
Parallel
GPT-5.6 Sol · high · 25-turn
80.8%15.0% complete tables
Favicon for Exa
Exa
GPT-5.6 Luna · xhigh · 25-turn
78.0%16.0% complete tables
Answer-item accuracy Complete tables
03

Price and speed

Compare answer-item accuracy with cost and latency; complete tables and Pareto points appear in each chart.

Price versus quality

Average cost per question on a logarithmic scale. The line shows the best quality available at each price level.

  • Perplexity
  • OpenAI Native
  • Exa
  • Parallel
  • Perplexity, 25-turn: GPT-5.6 Luna · xhigh, 83.4% answer-item accuracy, 22.0% complete tables, $0.17, on the efficiency and quality frontier
  • Perplexity, 25-turn: GPT-5.6 Sol · high, 84.0% answer-item accuracy, 21.0% complete tables, $0.83, on the efficiency and quality frontier
  • Perplexity, 25-turn: Claude Opus 5 · high, 68.2% answer-item accuracy, 20.2% complete tables, $2.14
  • Perplexity, 25-turn: DeepSeek V4 Flash 0731 · high, 77.1% answer-item accuracy, 20.0% complete tables, $0.10
  • Perplexity, 5-turn: GPT-5.6 Luna · xhigh, 80.0% answer-item accuracy, 18.0% complete tables, $0.063, on the efficiency and quality frontier
  • Perplexity, 5-turn: GPT-5.6 Sol · high, 81.4% answer-item accuracy, 18.0% complete tables, $0.45
  • Perplexity, 5-turn: DeepSeek V4 Flash 0731 · high, 72.1% answer-item accuracy, 17.3% complete tables, $0.050, on the efficiency and quality frontier
  • OpenAI Native, 5-turn: GPT-5.6 Sol · high, 78.4% answer-item accuracy, 17.0% complete tables, $0.53
  • OpenAI Native, 25-turn: GPT-5.6 Sol · high, 81.8% answer-item accuracy, 16.0% complete tables, $0.98
  • Exa, 25-turn: GPT-5.6 Luna · xhigh, 78.0% answer-item accuracy, 16.0% complete tables, $0.22
  • Parallel, 25-turn: GPT-5.6 Sol · high, 80.8% answer-item accuracy, 15.0% complete tables, $1.74
  • Exa, 25-turn: DeepSeek V4 Flash 0731 · high, 74.5% answer-item accuracy, 15.0% complete tables, $0.15
  • Parallel, 25-turn: Claude Opus 5 · high, 64.6% answer-item accuracy, 14.1% complete tables, $7.52
  • Parallel, 5-turn: GPT-5.6 Sol · high, 77.2% answer-item accuracy, 14.0% complete tables, $0.68
  • Parallel, 25-turn: GPT-5.6 Luna · xhigh, 78.5% answer-item accuracy, 14.0% complete tables, $0.17
  • Perplexity, 5-turn: Claude Opus 5 · high, 67.8% answer-item accuracy, 13.1% complete tables, $0.74
  • Perplexity, 1-turn: GPT-5.6 Sol · high, 79.8% answer-item accuracy, 13.0% complete tables, $0.26
  • OpenAI Native, 1-turn: GPT-5.6 Sol · high, 76.9% answer-item accuracy, 13.0% complete tables, $0.35
  • Parallel, 1-turn: GPT-5.6 Sol · high, 77.1% answer-item accuracy, 13.0% complete tables, $0.30
  • Parallel, 5-turn: GPT-5.6 Luna · xhigh, 74.1% answer-item accuracy, 13.0% complete tables, $0.063
  • Exa, 5-turn: Claude Opus 5 · high, 65.8% answer-item accuracy, 12.1% complete tables, $0.62
  • Exa, 25-turn: Claude Opus 5 · high, 67.0% answer-item accuracy, 12.1% complete tables, $1.85
  • Parallel, 5-turn: Claude Opus 5 · high, 64.8% answer-item accuracy, 12.0% complete tables, $1.28
  • Exa, 5-turn: GPT-5.6 Luna · xhigh, 75.7% answer-item accuracy, 12.0% complete tables, $0.075
  • Exa, 25-turn: GPT-5.6 Sol · high, 77.4% answer-item accuracy, 11.6% complete tables, $0.86
  • Exa, 5-turn: GPT-5.6 Sol · high, 74.4% answer-item accuracy, 11.6% complete tables, $0.47
  • Perplexity, 1-turn: GPT-5.6 Luna · xhigh, 71.1% answer-item accuracy, 11.0% complete tables, $0.037, on the efficiency and quality frontier
  • Perplexity, 1-turn: Claude Opus 5 · high, 65.8% answer-item accuracy, 10.0% complete tables, $0.23
  • Parallel, 5-turn: DeepSeek V4 Flash 0731 · high, 68.8% answer-item accuracy, 10.0% complete tables, $0.054
  • Exa, 5-turn: DeepSeek V4 Flash 0731 · high, 66.3% answer-item accuracy, 9.2% complete tables, $0.059
  • Parallel, 25-turn: DeepSeek V4 Flash 0731 · high, 75.3% answer-item accuracy, 9.0% complete tables, $0.13
  • Exa, 1-turn: GPT-5.6 Sol · high, 68.2% answer-item accuracy, 8.5% complete tables, $0.29
  • Parallel, 1-turn: GPT-5.6 Luna · xhigh, 68.5% answer-item accuracy, 8.0% complete tables, $0.036, on the efficiency and quality frontier
  • Exa, 1-turn: GPT-5.6 Luna · xhigh, 69.1% answer-item accuracy, 8.0% complete tables, $0.038
  • Parallel, 1-turn: Claude Opus 5 · high, 60.3% answer-item accuracy, 7.0% complete tables, $0.26
  • Exa, 1-turn: Claude Opus 5 · high, 59.6% answer-item accuracy, 5.5% complete tables, $0.26

Latency versus quality

Typical run-level average time per question. The line shows the best quality available at each latency level.

  • Perplexity
  • OpenAI Native
  • Exa
  • Parallel
  • Perplexity, 25-turn: GPT-5.6 Luna · xhigh, 83.4% answer-item accuracy, 22.0% complete tables, 2.7m
  • Perplexity, 25-turn: GPT-5.6 Sol · high, 84.0% answer-item accuracy, 21.0% complete tables, 2.3m, on the efficiency and quality frontier
  • Perplexity, 25-turn: Claude Opus 5 · high, 68.2% answer-item accuracy, 20.2% complete tables, 2.5m
  • Perplexity, 25-turn: DeepSeek V4 Flash 0731 · high, 77.1% answer-item accuracy, 20.0% complete tables, 2.0m
  • Perplexity, 5-turn: GPT-5.6 Luna · xhigh, 80.0% answer-item accuracy, 18.0% complete tables, 2.4m
  • Perplexity, 5-turn: GPT-5.6 Sol · high, 81.4% answer-item accuracy, 18.0% complete tables, 2.2m, on the efficiency and quality frontier
  • Perplexity, 5-turn: DeepSeek V4 Flash 0731 · high, 72.1% answer-item accuracy, 17.3% complete tables, 2.6m
  • OpenAI Native, 5-turn: GPT-5.6 Sol · high, 78.4% answer-item accuracy, 17.0% complete tables, 2.7m
  • OpenAI Native, 25-turn: GPT-5.6 Sol · high, 81.8% answer-item accuracy, 16.0% complete tables, 3.9m
  • Exa, 25-turn: GPT-5.6 Luna · xhigh, 78.0% answer-item accuracy, 16.0% complete tables, 2.9m
  • Parallel, 25-turn: GPT-5.6 Sol · high, 80.8% answer-item accuracy, 15.0% complete tables, 2.7m
  • Exa, 25-turn: DeepSeek V4 Flash 0731 · high, 74.5% answer-item accuracy, 15.0% complete tables, 2.3m
  • Parallel, 25-turn: Claude Opus 5 · high, 64.6% answer-item accuracy, 14.1% complete tables, 3.4m
  • Parallel, 5-turn: GPT-5.6 Sol · high, 77.2% answer-item accuracy, 14.0% complete tables, 2.4m
  • Parallel, 25-turn: GPT-5.6 Luna · xhigh, 78.5% answer-item accuracy, 14.0% complete tables, 2.8m
  • Perplexity, 5-turn: Claude Opus 5 · high, 67.8% answer-item accuracy, 13.1% complete tables, 1.8m, on the efficiency and quality frontier
  • Perplexity, 1-turn: GPT-5.6 Sol · high, 79.8% answer-item accuracy, 13.0% complete tables, 1.9m, on the efficiency and quality frontier
  • OpenAI Native, 1-turn: GPT-5.6 Sol · high, 76.9% answer-item accuracy, 13.0% complete tables, 2.7m
  • Parallel, 1-turn: GPT-5.6 Sol · high, 77.1% answer-item accuracy, 13.0% complete tables, 2.0m
  • Parallel, 5-turn: GPT-5.6 Luna · xhigh, 74.1% answer-item accuracy, 13.0% complete tables, 2.2m
  • Exa, 5-turn: Claude Opus 5 · high, 65.8% answer-item accuracy, 12.1% complete tables, 2.0m
  • Exa, 25-turn: Claude Opus 5 · high, 67.0% answer-item accuracy, 12.1% complete tables, 2.8m
  • Parallel, 5-turn: Claude Opus 5 · high, 64.8% answer-item accuracy, 12.0% complete tables, 1.9m
  • Exa, 5-turn: GPT-5.6 Luna · xhigh, 75.7% answer-item accuracy, 12.0% complete tables, 2.4m
  • Exa, 25-turn: GPT-5.6 Sol · high, 77.4% answer-item accuracy, 11.6% complete tables, 3.4m
  • Exa, 5-turn: GPT-5.6 Sol · high, 74.4% answer-item accuracy, 11.6% complete tables, 3.1m
  • Perplexity, 1-turn: GPT-5.6 Luna · xhigh, 71.1% answer-item accuracy, 11.0% complete tables, 2.4m
  • Perplexity, 1-turn: Claude Opus 5 · high, 65.8% answer-item accuracy, 10.0% complete tables, 70s, on the efficiency and quality frontier
  • Parallel, 5-turn: DeepSeek V4 Flash 0731 · high, 68.8% answer-item accuracy, 10.0% complete tables, 2.7m
  • Exa, 5-turn: DeepSeek V4 Flash 0731 · high, 66.3% answer-item accuracy, 9.2% complete tables, 2.8m
  • Parallel, 25-turn: DeepSeek V4 Flash 0731 · high, 75.3% answer-item accuracy, 9.0% complete tables, 2.3m
  • Exa, 1-turn: GPT-5.6 Sol · high, 68.2% answer-item accuracy, 8.5% complete tables, 2.8m
  • Parallel, 1-turn: GPT-5.6 Luna · xhigh, 68.5% answer-item accuracy, 8.0% complete tables, 2.3m
  • Exa, 1-turn: GPT-5.6 Luna · xhigh, 69.1% answer-item accuracy, 8.0% complete tables, 2.4m
  • Parallel, 1-turn: Claude Opus 5 · high, 60.3% answer-item accuracy, 7.0% complete tables, 74s
  • Exa, 1-turn: Claude Opus 5 · high, 59.6% answer-item accuracy, 5.5% complete tables, 90s
04

How does search budget affect answer-item accuracy?

See how answer-item accuracy changes with search budget; complete-table success appears beneath. Missing cells were not run.

Search provider1-turn5-turn25-turn
Favicon for Perplexity
Perplexity
79.8%13.0% complete tablesGPT-5.6 Sol · high81.4%18.0% complete tablesGPT-5.6 Sol · high84.0%21.0% complete tablesGPT-5.6 Sol · high
Favicon for openai
OpenAI Native
76.9%13.0% complete tablesGPT-5.6 Sol · high78.4%17.0% complete tablesGPT-5.6 Sol · high81.8%16.0% complete tablesGPT-5.6 Sol · high
Favicon for Exa
Exa
69.1%8.0% complete tablesGPT-5.6 Luna · xhigh75.7%12.0% complete tablesGPT-5.6 Luna · xhigh78.0%16.0% complete tablesGPT-5.6 Luna · xhigh
Favicon for Parallel
Parallel
77.1%13.0% complete tablesGPT-5.6 Sol · high77.2%14.0% complete tablesGPT-5.6 Sol · high80.8%15.0% complete tablesGPT-5.6 Sol · high
05

All search configurations

Verified configurations for all models, ranked by answer-item accuracy; complete-table success and intervals are shown too.

#ModelSearch providerSearch budgetReasoning effortAnswer-item accuracyComplete tablesCost / questionTypical time / questionQuestions
1
Favicon for openai
GPT-5.6 Sol
Favicon for Perplexity
Perplexity
25-turnhigh84.0%21.0%$0.832.3m100
2
Favicon for openai
GPT-5.6 Luna
Favicon for Perplexity
Perplexity
25-turnxhigh83.4%22.0%$0.172.7m100
3
Favicon for openai
GPT-5.6 Sol
Favicon for openai
OpenAI Native
25-turnhigh81.8%16.0%$0.983.9m100
4
Favicon for openai
GPT-5.6 Sol
Favicon for Perplexity
Perplexity
5-turnhigh81.4%18.0%$0.452.2m100
5
Favicon for openai
GPT-5.6 Sol
Favicon for Parallel
Parallel
25-turnhigh80.8%15.0%$1.742.7m100
6
Favicon for openai
GPT-5.6 Luna
Favicon for Perplexity
Perplexity
5-turnxhigh80.0%18.0%$0.0632.4m100
7
Favicon for openai
GPT-5.6 Sol
Favicon for Perplexity
Perplexity
1-turnhigh79.8%13.0%$0.261.9m100
8
Favicon for openai
GPT-5.6 Luna
Favicon for Parallel
Parallel
25-turnxhigh78.5%14.0%$0.172.8m100
9
Favicon for openai
GPT-5.6 Sol
Favicon for openai
OpenAI Native
5-turnhigh78.4%17.0%$0.532.7m100
10
Favicon for openai
GPT-5.6 Luna
Favicon for Exa
Exa
25-turnxhigh78.0%16.0%$0.222.9m100

Why we run this benchmark

WideSearch measures breadth. The agent has to enumerate a full entity set, chase down each attribute, and emit a structured table — closer to real research workflows (market scans, literature surveys, competitive tables) than single-answer trivia. Half the tasks are in Chinese, so it also exercises engines outside English-language results.

We run it with the model held fixed because that isolates the variables OpenRouter users actually control. Those are which engine handles the searches, whether search runs as a server tool or a plugin, and how many agent turns the loop is allowed. Those knobs are exactly what you can set on a request today.

What the scores can and can't tell you

These scores compare search configurations, not agent products. The model reads search result excerpts only, with no full-page fetching and no code tools, so absolute numbers sit below published agent leaderboards, which allow both. Compare configurations rather than raw levels.

The headline score is answer-item accuracy, which gives partial credit for matched table items so near-misses remain visible. Strict WideSearch Success Rate remains the secondary measure and requires every row and cell in the table to be correct. Even frontier models score low upstream — the paper reports about 4.5% single-agent success for OpenAI o3 (Avg@4), with answer-item accuracy around 52–58% for leading single-agent systems. Engine gaps smaller than the confidence ranges are unresolved, and a configuration's cost should be read alongside its quality.

The questions are public, which means memorization can inflate absolute scores over time; differences between search configurations remain meaningful because every configuration shares the same model.

How tasks are scored

Each task is one query with a reference table. The model's table is parsed, its rows and columns are aligned to the reference by a judge model, and each cell is compared — exact match where possible, judge-graded equivalence otherwise. A task succeeds only when every reference row is found and every cell matches, with no missing or extra rows.

reward = all_rows_aligned ∧ every_cell_correct   // ∈ {0, 1}
judge  = gpt-4.1 at temperature 0 for alignment and fuzzy cells
unparseable or empty tables score 0

A search configuration is one model, search engine, request format, and maximum search budget. Scores are weighted by the number of questions in each verified run; small smoke tests are excluded.

EX

Real run example

Fixed historical runs, independent of the model filter above. Long results and operational metadata are trimmed; selected queries may be edited or omitted.

Task

Build a verified clinic table for the Fresno area

Make a comprehensive, verified list of abortion clinics operating in California during 2024 within 65 miles of E Fedora Ave in Fresno. Omit facilities that closed during 2024. Return clinic, county, full address, phone number, and Saturday hours in one Markdown table.

This is a historical 2024 benchmark snapshot, not current healthcare-service guidance.

Reference

Six rows: Planned Parenthood Fulton Street, FPA Women's Health Fresno, Planned Parenthood Family First, Planned Parenthood Madera, Planned Parenthood Merced, and FPA Women's Health Tulare. The reference marks only Fulton Street as open Saturday.

Answer-item accuracy 62.2%
Favicon for Exa
Exa
Claude Opus 4.8 · 1 turn

Final answer

ClinicCountyAddressPhoneSaturday hours
Planned Parenthood - Fulton StreetFresno650 N. Fulton Street, Fresno, CA 93728(559) 488-49008:00 am - 4:30 pm
Planned Parenthood - Family FirstFresno6095 N. 1st Street, Fresno, CA 93710(559) 446-15158:00 am - 4:30 pm
FPA Women's Health - FresnoFresno165 N. Clark Street, Fresno, CA 93701(559) 233-8657Not available

What happened

The one-turn run found three of six clinics. Two rows were fully correct; Family First had the wrong Saturday hours. Answer-item accuracy was 62.2%, and the complete-table verdict was incorrect.

Methodology

Runs execute on OpenRouter's native benchmark harness against production endpoints. Server-tool configurations set a request-level tool-call budget on the web search server tool. Plugin configurations make one pre-inference search of the question with the web search plugin. Engines use the same default configurations that serve production traffic.

Every run persists its exact model, engine, request format, search budget, cost, and available timing telemetry. Missing configurations stay missing in the comparison table, and absent or zero timing telemetry is not treated as instantaneous performance.