omni-bench · 13 measurements · 11 engines · august 2026

Two Rankings, One Language

Every engine here transcribes Russian. Ranked on ordinary Russian speech they are within a factor of 1.4 of each other. Ranked on the English terminology inside that speech they span a factor of 71.

For Russian engineering speech: ElevenLabs Scribe v2. It carries 71 of 94 technical terms and holds a competitive word error rate while doing it — the only engine measured that does not force a choice between the two.

For Russian itself: GigaAM-v3. On spontaneous speech it reaches 6.1% word error where the next best is 10.1% — at a fifth of the parameters. It also carries 19 terms of 94, second worst on the board.

Yandex SpeechKit carries one term of 94 while transcribing Russian perfectly acceptably. The Russian-market default is unusable for this workload, and no general benchmark would tell you.


The thesis

The two axes disagree

Horizontal: word error on spontaneous Russian carrying no English at all — 59 control recordings, so it measures Russian and nothing else. Vertical: technical terms carried, of 94. A single-metric leaderboard reads only one edge of this field.

6.1% ← better Russianword error, spontaneous Russianworse → 19.6%

GigaAM-v3 sits alone at the left edge: the best Russian on the board by a wide margin, and near the bottom for terminology. Yandex sits at the floor. Nothing in the horizontal position predicts the vertical one.


The board

Thirteen measurements

87 recordings of a Russian software-engineering podcast, 33 minutes, 94 marked term occurrences. Every run complete over the exact sample set.

EngineWhereParams Terms /94Term error WER allWER Russian-only
ElevenLabs scribe_v2hosted 7124.5%10.6%10.2%
OpenAI gpt-4o-transcribehosted 6530.9%18.4%19.6%
whisper large-v3local1.54B 6036.2%11.0%10.1%
AssemblyAI universal-2hosted 5244.7%11.5%10.3%
Deepgram nova-3 · multihosted 5046.8%11.5%10.4%
Qwen3-ASRlocal2.04B 4453.2%12.1%11.0%
parakeet-tdt-0.6b-v3local0.63B 3958.5%13.0%11.6%
Canary 1B v2local0.96B 3661.7%13.9%12.4%
Deepgram nova-3 · ruhosted 3463.8%12.8%
Nemotron-3.5-asrlocal0.64B 2573.4%16.1%
GigaAM-v3local0.22B 1979.8%8.5%6.1%
Yandex SpeechKithosted 198.9%16.4%13.3%
GigaAM-Multilinguallocal0.59B 0100%11.9%

The same engine, two modes, sixteen terms apart

Deepgram's Nova-3 offers language=ru and a language=multi mode documented for code switching. Nothing else changes — same model, same audio, same request. The mode is worth 34 terms against 50, a wider gap than separates most different engines on this board. Configuration is not a footnote here; it is the second-largest effect measured.


The control

Why a standard benchmark would have misled us

The obvious control is FLEURS, the usual multilingual read-speech benchmark. It gives a reassuring, useless answer.

EngineFLEURS ru (read) Podlodka controls (spontaneous)Change
GigaAM-v38.7%6.1%0.70×
whisper large-v38.7%10.1%1.16×
ElevenLabs scribe_v29.0%10.2%1.14×
Yandex SpeechKit10.2%13.3%1.30×
OpenAI gpt-4o-transcribe7.2%19.6%2.71×

On FLEURS, whisper and GigaAM tie to six decimal places — 105 word errors out of 1204 each. Their transcripts differ, so the tie is a coincidence of counts rather than identical output, but the conclusion a reader would draw from that column is that the two are equivalent. On spontaneous speech GigaAM is ahead by four-tenths of the other's error rate.

OpenAI's profile is the sharpest: best on read speech, worst on spontaneous, degrading by a factor of 2.7 where nobody else exceeds 1.3. Two explanations fit — FLEURS is a public benchmark that plausibly sits in training mixes, and read speech is simply easier — and these measurements cannot separate them. Either way the practical lesson is the same: a read-speech benchmark did not rank these engines the way the workload does.


Ruled out

What does not explain the terminology gap


Interventions

What moved the number, and what did not

InterventionEngineEffectCost
code-switch modeDeepgram nova-334 → 50 termsone parameter
oracle glossarywhisper large-v324 → 30 of 44knowing the answers
oracle glossaryGigaAM-v3 CTC5 → 12 of 44a custom rescorer
realistic glossarywhisper large-v324 → 24 of 446,100 mined terms
LoRA on synthetic dataGigaAM-v3 CTC14 → 14 of 94a day, and WER worsened

A glossary only helps if it already contains the words

An oracle glossary — the actual answer key — lifts whisper by six terms. A realistic glossary of 6,100 terms, mined by frequency from a thousand episodes of a different Russian IT podcast, lifts it by none. Coverage explains it: that glossary contains 58 of the 94 occurrences, and observability appears in none of the thousand episodes. Podlodka, Crew and Datadog are speaker-context proper nouns no external list will ever hold.

The deployable conclusion is narrower than "use a glossary": the glossary has to be your own vocabulary — service names, internal tools, team names out of your wiki and repositories. A generic IT word list is dominated by Google, Apple and iPhone, and shares almost nothing with what an engineering standup actually says.

Fine-tuning was tried properly and failed. Synthetic code-switched audio was built from two monolingual FLEURS corpora by phrase-level splicing, a LoRA was trained on the encoder's feed-forward layers, and the holdout loss fell from 2.33 to 0.60 — the model learned the synthetic task. It transferred nothing: 14 terms before, 14 after, and word error rose from 9.7% to 11.0%. At 772 utterances against the thousand hours per language pair the published recipe uses, this tests whether crumbs transfer, not whether the method works.


Limits

What this does not establish


Provenance

Reproducing this

Every figure comes from a Result document complete over the exact sample set, scored by the same pinned scorer, with the dataset, protocol, model revision and hardware bound into the identity.

Pinned identity

task   asr.code_switch.ru_tech.v1 · sha256:ed04da300adc95567b55fbe27e8279ddcb3b38d4cb2e2d21759e10a295228474
control   asr.fleurs.ru.quick.v1 — 64 recordings, 12.4 min
dataset   bond005/podlodka_speech @ 4a9161485a70ea849ccaa54a72faa9bfc0a60627 · splits train + test
excluded   validation — byte-identical duplicate of test, same LFS sha256 and git object
audio   16 kHz mono PCM16, downmixed and resampled with libsoxr VHQ
profile   audio_transcription.batch_single.v1 · batch · concurrency 1 · warmup 0
local hardware   Intel Core i9-13900K · NVIDIA GeForce RTX 4090 · 24 GB
hosted   measured from an Apple M1 Max; compute is the vendor's and unpinnable

whisper   openai/whisper-large-v3 · float16 · faster-whisper.ctranslate2 · beam=5
qwen3   Qwen/Qwen3-ASR-1.7B-hf @ bcd2b5b7 · bfloat16 · 2,038,052,480 params
parakeet   nvidia/parakeet-tdt-0.6b-v3 @ 541d1f99 · float32
nemotron   nvidia/nemotron-3.5-asr-streaming-0.6b @ 1c8deaec · lookahead=6
canary   nvidia/canary-1b-v2 · NeMo EncDecMultiTaskModel · pnc=yes
gigaam-v3   ai-sage/GigaAM-v3 @ 7655ad71 · e2e_rnnt · 222,509,313 params
gigaam-ml   ai-sage/GigaAM-Multilingual @ 3905cd51 · large_ctc
elevenlabs   scribe_v2 · v1/speech-to-text
openai   gpt-4o-transcribe · v1/audio/transcriptions · temperature 0
assemblyai   universal-2 · v2/transcript
deepgram   nova-3 · v1/listen · smart_format · language ru and multi
yandex   speechkit v1 · lpcm · 28 s chunks cut at the quietest frame