ASR Leaderboard — Arabic + English short-form

Loading leaderboard.tsv

Evaluation normalization

All reported WER/CER scores use the same bilingual Arabic–English text normalizer before scoring. The goal is to reduce formatting noise while preserving Arabic words, Latin words, digits, selected symbols, decimals, thousands separators, and contractions.

Show normalization steps
  1. Remove Arabic diacritics.
  2. Keep Arabic letters, Latin letters, digits, @, %, and whitespace.
  3. Keep . or , only when they appear between two digits, such as decimal or thousands separators.
  4. Keep apostrophes only when they appear between two non-space characters, such as contractions.
  5. Replace all other characters with a space.
  6. Normalize Alef variants to bare Alef: ا.
  7. Collapse repeated whitespace into a single space.
  8. Lowercase Latin letters.
  9. Strip leading and trailing whitespace.

Columns & datasets

Pick which metrics and datasets appear in the table. Datasets are grouped by cluster — toggle a whole cluster with its header checkbox. The Average columns recompute over the datasets you have selected.

Choose columns
Loading…

Models (rows)

Check the models you want to compare. Use the search box to find models fast, then bulk-check the matches.

Choose models
Loading…
Leaderboard
Loading leaderboard data…