WorldBet is James L. Hieronymus's ASCII encoding of the IPA — plus extra broad-phonetic symbols — designed to cover all the world's languages for speech-database labelling. It was defined in an AT&T Bell Laboratories technical memorandum, "ASCII Phonetic Symbols for the World's Languages: Worldbet" (c. 1993). Where earlier ASCII schemes were built around European languages, WorldBet's stated motivation was that those schemes "left out many of the sounds of the other languages" — so it aimed to be genuinely multilingual from the start.1
The core design principle
WorldBet's founding rule is that "any spectrally and temporally distinct speech sound (not including pitch) which is phonemic in some language should have a separate base symbol." That is a deliberate contrast with more graphemic label sets: as the memo puts it, "each type of /r/ will have its separate IPA-like designation, rather than the more graphemic r used in some label sets." This is the root of WorldBet's central teaching point — it splits apart sounds that SAMPA merges.1
How the codes are built
- Fixed two-character baseforms: every base symbol is two ASCII characters (which may differ in case, and one may be a space) — around 2,900 possibilities, of which roughly 299 are in use. The fixed length makes strings uniquely parseable by a computer.1
- Vowels use orthographic vowel letters (a, A, o, O) plus the numerals 2–8 (0 and 1 are avoided as confusable with 'o' and 'l'); length is marked with a colon; diphthongs are written by concatenating their endpoint-vowel symbols.1
- Phonemic-in-many-languages features — nasalization, rhotization, aspiration — are written explicitly in the baseform; finer allophonic detail is added with diacritics; and a linking character '-' builds double articulations. Verified diacritics include aspirated h, ejective ', implosive <, palatalized j, labialized w, nasalized ~, and voiceless -.1
- Beyond the IPA: WorldBet adds broad-phonetic symbols the IPA of the day did not have, particularly for plosive bursts and hyper-aspiration.1
Why WorldBet ≠ SAMPA / X-SAMPA
The memo explicitly positions WorldBet against its predecessors — Klatt's phonbet, PHONASCII, ARPABET, TIMITBET, SAMPA, MRPA, Maddieson's UCLA set, and the sci.lang scheme (Kirshenbaum). Its central critique of SAMPA is that, because SAMPA assumed within-language use, it reuses "the same symbol for quite different sounds, most notably for /r/." WorldBet deliberately does the opposite. That means WorldBet codes and X-SAMPA/SAMPA codes are not interchangeable — for example, WorldBet's Q is ɢ (voiced uvular plosive), while in X-SAMPA Q is the vowel ɒ.12
Status: mostly a legacy-reading skill
WorldBet was motivated by corpus-driven speech research, phonological inventories, automatic language identification, and multilingual recognition and synthesis; its symbol coverage was checked against Maddieson's UPSID inventory database, and it appears in some 1990s multilingual corpora (the memo names TIMIT, SCRIBE, BDSONS, PHONDAT). It never displaced SAMPA/X-SAMPA in general use, so today it is chiefly a legacy-reading skill and a source of rows for cross-system comparison tables.1
Errors to avoid
- Conflating WorldBet codes with X-SAMPA/SAMPA — the assignments differ, and WorldBet deliberately splits /r/-type sounds that SAMPA merges.12
- Assuming IPA-completeness in either direction — WorldBet both exceeds the IPA (extra broad-phonetic symbols) and is a labelling set, not a narrow-transcription standard.1
- Reading a two-character code as two separate segments — the baseform is a fixed two-character unit (one character may be a space).1
Learn more
- Hieronymus, "ASCII Phonetic Symbols for the World's Languages: Worldbet" (AT&T Bell Labs) — The defining memo — design principles, symbol tables, and comparison with earlier schemes.
- X-SAMPA — our companion guide — The universal ASCII-IPA scheme WorldBet is most confused with (different assignments).
- ARPABET — our companion guide — One of the predecessor English schemes the memo distinguishes from WorldBet.
- The IPA — our companion guide — The alphabet WorldBet encodes and then extends.
Notes & Bibliography
- Hieronymus, James L. "ASCII Phonetic Symbols for the World's Languages: Worldbet." AT&T Bell Laboratories, Murray Hill, NJ, c. 1993 (date derivable: the memo notes the IPA's '105 years' since 1888 = 1993). Read Run 5 (body + Tables 1–2). Anchors: ASCII IPA + broad-phonetic symbols for all languages; one base symbol per spectrally/temporally distinct phonemic sound; fixed two-character baseforms (~2,900 possible, ~299 used, one char may be a space); vowels = letters + numerals 2–8; diphthongs = concatenated endpoints; linking '-' for double articulations; explicit critique of SAMPA reusing symbols for different /r/; corpora TIMIT/SCRIBE/BDSONS/PHONDAT; coverage checked vs UPSID/Maddieson. [source] ↩
- Cross-system divergence (LinguaCommons research, primary-sourced columns): WorldBet Q = ɢ (voiced uvular plosive) vs X-SAMPA Q = ɒ (open back rounded vowel) — one illustration that WorldBet and X-SAMPA/SAMPA assignments are not interchangeable. Built from Hieronymus Table 1 [W01] and Wells's X-SAMPA. [source] ↩