Skip to content
LinguaCommons
← Notation Systems

SAMPA — the Speech Assessment Methods Phonetic Alphabet

SAMPA — the Speech Assessment Methods Phonetic Alphabet — is a machine-readable phonetic alphabet that maps IPA symbols onto printable 7-bit ASCII, in the range 33–127. It exists because 1980s and 1990s computing could not render IPA glyphs: e-mail, databases and speech technology systems all needed phonetic transcription that would survive a 7-bit channel intact.1 Unicode has since removed that constraint, but SAMPA has not gone away, because an enormous quantity of speech-technology data was transcribed in it and still has to be read.

The single most important thing to understand about SAMPA is what it is not. It is not one universal alphabet. It is a mapping plus per-language transcription guidelines — in the words of the official UCL page, “associated with the coding (mapping) are guidelines for the transcription of the languages to which SAMPA has been applied”.1 The same ASCII character can therefore mean different things in the SAMPA table for Dutch and the SAMPA table for Thai. This is the difference between SAMPA and X-SAMPA, and it is the source of most of the confusion around both.

1. Origins — with an honest asterisk

The official UCL page states that SAMPA was “originally developed under the ESPRIT project 1541, SAM (Speech Assessment Methods) in 1987-89 by an international group of phoneticians”.1 Wells's own 1995 X-SAMPA paper instead dates the work to 1988–1991 and cites ESPRIT project 2589.2 These two primary sources disagree, and this guide does not resolve the discrepancy — it records it. If you need to cite the project number, cite the source you are following and say which one.

What both sources agree on is that SAMPA was collaborative. The UCL page draws the contrast explicitly: “Unlike other proposals for mapping the IPA onto ASCII, SAMPA is not one single author's scheme, but represents the outcome of collaboration and consultation among speech researchers in many different countries,” developed “by or in consultation with native speakers of every language to which they have been applied”.1 X-SAMPA and Kirshenbaum are single-author proposals; SAMPA is a committee product, and its per-language structure follows directly from that.

Language coverage grew in waves.1 The initial six were in place by 1989 — Danish, Dutch, English, French, German and Italian; Norwegian and Swedish followed by 1992; Greek, Portuguese and Spanish in 1993. The BABEL project added Bulgarian, Estonian, Hungarian, Polish and Romanian in 1996, and OrienTel added Arabic, Hebrew and Turkish. Cantonese, Croatian, Czech, Russian, Slovenian and Thai were also covered.

SAMPA language coverage by project wave

WaveYearLanguages
Initial sixby 1989Danish, Dutch, English, French, German, Italian
Nordicby 1992Norwegian, Swedish
Southern Europe1993Greek, Portuguese, Spanish
BABEL1996Bulgarian, Estonian, Hungarian, Polish, Romanian
OrienTelArabic, Hebrew, Turkish
FurtherCantonese, Croatian, Czech, Russian, Slovenian, Thai

2. The design in one rule

The base rule is simple enough to state in a sentence: “All IPA symbols that coincide with lower-case letters of the Latin alphabet remain the same; all other symbols are recoded within the ASCII range 37..126.”1 Two consequences follow. SAMPA is case-sensitive — upper and lower case are different symbols, so a and A are different vowels. And it is uniquely parsable: a string of SAMPA symbols can be read unambiguously with no spaces between them.

The vowels of the basic table

Vowels: SAMPA basic table, with the X-SAMPA value alongside

IPANameSAMPAX-SAMPA
iclose front unroundedii
yclose front roundedyy
ɨclose central unrounded11
ʉclose central rounded}}
ɯclose back unroundedMM
uclose back roundeduu
ɪnear-close near-front unroundedII
ʏnear-close near-front roundedYY
ʊnear-close near-back roundedUU
eclose-mid front unroundedee
øclose-mid front rounded22
oclose-mid back roundedoo
əmid central schwa@@
ɛopen-mid front unroundedEE
œopen-mid front rounded99
ɜopen-mid central unrounded33
ʌopen-mid back unroundedVV
ɔopen-mid back roundedOO
ænear-open front unrounded{{
ɐnear-open central66
aopen front unroundedaa
ɶopen front rounded&&
ɑopen back unroundedAA
ɒopen back roundedQQ

Consonants that are not simply themselves

Consonants whose IPA symbol is already a lower-case Latin letter keep that letter. The ones worth learning are the rest:

Recoded consonants in the SAMPA basic table

IPANameSAMPA
ɡvoiced velar plosiveg
ʔglottal stop?
ɱlabiodental nasalF
ɲpalatal nasalJ
ŋvelar nasalN
βvoiced bilabial fricativeB
θvoiceless dental fricativeT
ðvoiced dental fricativeD
ʃvoiceless postalveolar fricativeS
ʒvoiced postalveolar fricativeZ
çvoiceless palatal fricativeC
ɣvoiced velar fricativeG
χvoiceless uvular fricativeX
ʁvoiced uvular fricativeR
ʋlabiodental approximantP
ʎpalatal lateral approximantL
ʍvoiceless labial-velar fricativeW
ɥlabial-palatal approximantH

Beyond the segments, length is : for ː, primary stress is " and secondary stress is %.1

3. Two obsolescence notes that bite

The official page carries two notes that matter to anyone reading older material.1

  • The tone marks are obsolete. SAMPA's original tone marks — backtick for falling, apostrophe for rising — were based on the pre-1990 IPA and have been superseded by SAMPROSA, the companion prosodic transcription scheme.3 X-SAMPA later reassigned both of those characters to other jobs entirely, which is a good way to misread a file.
  • The diacritic order was reversed. SAMPA originally placed the syllabicity diacritic before the base character — =n for syllabic n. ISO and Unicode require diacritics to follow the base, and the official page directs that new work follow the base. Old corpora will still show the old order.

Neither of these is a matter of taste: they are documented changes, and a file transcribed before them will not parse the way a modern reader expects.

4. Where SAMPA stands today

SAMPA remains embedded in legacy speech-technology resources — EUROM 1, BABEL, Onomastica and OrienTel among them — has been used by Oxford University Press, and is listed by the Linguistic Data Consortium.1 For new work, Unicode IPA has removed the original motivation and X-SAMPA covers the universal case. SAMPA literacy is now a reading skill: you need it to open a 1990s lexicon, not to write a 2026 one.

The canonical printed citation is Wells's chapter in the Handbook of Standards and Resources for Spoken Language Systems (Mouton de Gruyter, 1997), Part IV, section B; the UCL site is the canonical online reference and was last revised on 25 October 2005.14

5. Errors to avoid

  • Treating SAMPA as one universal alphabet. It is a family of per-language tables, and the same ASCII character can map differently between them.1
  • Conflating SAMPA with X-SAMPA. Per-language versus universal — see the sibling guide.2
  • Using the obsolete tone marks, or the pre-Unicode diacritic order.1
  • Expecting prosody coverage. That is SAMPROSA's job, not SAMPA's.3
  • Assuming any ASCII phonetic string you meet is SAMPA. Kirshenbaum, WorldBet and ARPABET all look similar and assign different values — see the comparison table in the X-SAMPA guide.

6. Sources

Notes & Bibliography

  1. Wells, John C., maintainer. “SAMPA computer readable phonetic alphabet.” University College London, site last revised 25 October 2005. The primary source for the ASCII 37–127 design rule, case sensitivity and unique parsability, the basic six-language symbol table, the ESPRIT 1541 / 1987–89 origin statement, the collaborative-development statement, the language timeline, the legacy-resource list, and both official obsolescence notes. [source]
  2. Wells, John C. “Computer-coding the IPA: a proposed extension of SAMPA.” Revised draft, 28 April 1995. University College London. Cited here for the conflicting origin statement — 1988–1991 under ESPRIT project 2589 — and for the SAMPA/X-SAMPA distinction. [source]
  3. University College London. “SAMPROSA — SAM Prosodic Transcription.” The companion scheme that supersedes SAMPA's original tone marks. [source]
  4. Wells, John C. “SAMPA computer readable phonetic alphabet.” In Handbook of Standards and Resources for Spoken Language Systems, edited by Dafydd Gibbon, Roger Moore and Richard Winski, part IV, section B. Berlin and New York: Mouton de Gruyter, 1997. The canonical printed citation, as given by the official site. [source]