OpenbenchmarksOpenbenchmarks

Backed by Y Combinator

Independent benchmarks for build-vs-buy decisions. Verifiable, reproducible and open source.

Last run on 21 Aug 2026

Benchmarks

Company Lookalikes

Tests how well company-search and lookalike APIs turn a seed company domain or description into a ranked list of genuinely similar businesses. Use it to compare tools for account discovery, prospecting, market mapping, and TAM expansion.

Precision@100 · higher is better · scale 0–100%

Extruct61%

Ocean.io60%

Parallel57%

Exa49%

Discolike35%

open benchmark→

open data + code · github.com/openbenchmarks-labs/lookalikes

Company Firmographic Enrichment

282 company domains × 8 APIs, normalized into the same seven-field active scoring contract. Compare enrichment success rate, field accuracy, coverage, company match rate, latency, and cost by workflow.

Correct field yield · higher is better · scale 0–100%

People Data Labs89.0%

Apollo88.6%

Parallel86.0%

Predict Leads80.6%

Explorium70.5%

+3 more

open benchmark→

open data + code · github.com/openbenchmarks-labs/company-enrichment

Company Funding Data Enrichment

16 funding vendors, judged on latest funding stage against human labelled ground truth.

Correct stage yield · higher is better · scale 0–100%

Firecrawl92.3%

Parallel90.0%

Parallel Responses API (medium)89.0%

Exa88.7%

Exa88.6%

+15 more

open benchmark→

open data + code · github.com/openbenchmarks-labs/company-funding

Voice Agents Latency Benchmark

TTFAB — how long a caller waits before a voice AI agent starts speaking — measured over real phone calls from the call's own audio, never platform-reported timestamps.

Telnyx1296 ms

ElevenLabs1424 ms

Bland AI1520 ms

Vapi1558 ms

Retell AI1740 ms

p50 p90 · p95 p992,078 turns measured

open benchmark→

open data + code · github.com/openbenchmarks-labs/voice-agent-latency

Text-to-Speech

28 text-to-speech models on Time to First Audio (TTFA) + Word Error Rate, measured by Coval under production-realistic conditions. Mirrored with attribution.

Median TTFA · lower is better

vui103 ms

palabra-tts-v1104 ms

inworld-tts-2-flash115 ms

inworld-tts-1.5-mini143 ms

inworld-tts-2160 ms

+23 more

open benchmark→

Live Speech-to-Text

29 live speech-to-text models on Word Error Rate and Time to Final Segment (TTFS), measured by Coval under production-realistic conditions. Rolling 7-day window, mirrored with attribution. Ranked here by WER, not median TTFS.

Word Error Rate · lower is better · scale 0–100%

universal-3.5-pro3.3%

realtime3.6%

chirp_34.1%

enhanced4.3%

velma-2-stt-streaming-english-v24.7%

+24 more

open benchmark→

Structured Speech-to-Text

300 human-recorded workplace utterances × 17 ASR systems, checked for exact recovery of 1,482 structured values — emails, phone numbers, CLI flags, file paths, IDs. Cell value is Task Success Rate — recordings with every value correct.

Task Success Rate · higher is better · scale 0–100%

Deepgram Nova-371%

ElevenLabs Scribe v268%

Deepgram Nova-3 (streaming)61%

ElevenLabs Scribe v2 Realtime (streaming)59%

Google Cloud Chirp 358%

+12 more

open benchmark→