Agentic Search Tooling Benchmarks for Leading Local Models

Last updated: August 31, 2026
Models tested: DeepSeek V4 Flash 0731 (2x DGX Spark) GLM 5.3 Flash (2x DGX Spark) Qwen 3.8 27B (1x DGX Spark) Qwen 3.8 Flash Next (2x DGX Spark)
Search providers tested: BraveExaParallelSerperTavily
Best overall combination
Qwen logo+Parallel logo Qwen 3.8 27B + Parallel turbo mode
This combination got 50 out of 50 questions correct, at 5.7 avg seconds per answer, and an avg cost of $1.26 per 1,000 questions. Six other pairings also aced the test, but none with this combination of speed and cost. Runner-up: DeepSeek V4 Flash + Serper, at 7.2 seconds per answer and $1.30 per 1,000 questions.

Search provider leaderboard

Accuracy
Parallel turbo
197/200 across all models
Speed
Parallel turbo
7.1s avg per answer (0.48s per search call)
Cost
Parallel turbo
$1.26 avg per 1,000 answers
Each award pools all four models' graded answers on the same questions.
Local Search Bench: Accuracy vs Cost vs Speed
Every search provider, pooled across all four local models. Up and to the left wins: accuracy on the vertical, effective search fees per 1,000 answers on the horizontal (log scale), bigger dot = faster answers. Cost is not the provider's list price: it is computed from this benchmark's measured usage, the actual number of searches each model chose to run on each question (typically 1 to 2), billed at the provider's per-search list rate. It answers "what would 1,000 answers actually have cost me," which is why every provider plots above its raw per-search price. Model-only baseline (5.5% accuracy, $0) is off the chart.
Credit: x.com/ScottSanchez
Accuracy vs cost vs speed scatter chart: Parallel turbo leads at 98.5% accuracy, 7.1s per answer, $1.26 per 1,000 answers

Sort providers by click the chips to set your own priority order · default: accuracy, then cost, then speed
accuracy = correct answers out of 50 time = full wall clock: searching + page reading + model inference tokens: big number = input read; small line = output written search cost = list-price fees for the searches actually used, scaled to 1,000 answers

Results by question

Show the full question-by-question grid50 questions, every provider, all four models. Click to expand
correct wrong answer didn't answer G = GLM 5.3 Flash (2x DGX Spark) · Q = Qwen 3.8 27B (1x DGX Spark) · F = Qwen 3.8 Flash Next (2x DGX Spark) · D = DeepSeek V4 Flash 0731 (2x DGX Spark)

Wall time per task (all 50 questions)

GLM 5.3 Flash (2x DGX Spark)Qwen 3.8 27B (1x DGX Spark)Qwen 3.8 Flash Next (2x DGX Spark)DeepSeek V4 Flash 0731 (2x DGX Spark)

Key observations: models and provider combinations.

How this benchmark works.
Have feedback or suggestions?
Please reach out to me on X: @ScottSanchez