X-SAMPA is John C. Wells's 1995 proposal for a keyboard-compatible ASCII coding of the entire IPA — in his words, “everything on the 1993 IPA Chart, including diacritics and tone marks” — designed so that IPA-transcribed material could be sent through e-mail and other 7-bit channels without loss.1 Where SAMPA is a family of per-language tables, X-SAMPA is one universal mapping.12 That difference is the whole point of the scheme, and it is why X-SAMPA, not SAMPA, became the ASCII-IPA that people actually use for general phonetic work.
The defining document is a roughly 7,000-word paper, the fetched version marked “Revised draft 1995 04 28”, hosted at UCL as a PDF, with an IPA-free HTML summary page that was last revised on 3 May 2000 when Unicode code points were added to the symbol listing.13 Note that Wells framed the paper explicitly as a proposal soliciting colleagues' reactions. X-SAMPA's status as a de-facto standard comes from adoption, not from any ratifying body.1
1. The base layer
IPA symbols that are also lower-case Latin letters a–z keep their IPA values unchanged. Everything else is recoded within ASCII 33–126.1 Case is significant: upper and lower case are different symbols. And the scheme's defining property, stated in §26 of the paper, is that symbol strings remain uniquely parsable even with no spaces between successive characters.1
A sample of consonants X-SAMPA codes that the SAMPA basic table does not
| IPA | Name | X-SAMPA |
|---|---|---|
| ʈ | voiceless retroflex plosive | t` |
| ɖ | voiced retroflex plosive | d` |
| c | voiceless palatal plosive | c |
| ɟ | voiced palatal plosive | J\ |
| q | voiceless uvular plosive | q |
| ɢ | voiced uvular plosive | G\ |
| ɳ | retroflex nasal | n` |
| ɴ | uvular nasal | N\ |
| ʙ | bilabial trill | B\ |
| ʀ | uvular trill | R\ |
| ɾ | alveolar tap | 4 |
| ɽ | retroflex flap | r` |
| ɸ | voiceless bilabial fricative | p\ |
| ʂ | voiceless retroflex fricative | s` |
| ʐ | voiced retroflex fricative | z` |
| ʝ | voiced palatal fricative | j\ |
| ħ | voiceless pharyngeal fricative | X\ |
| ʕ | voiced pharyngeal fricative | ?\ |
| ɦ | voiced glottal fricative | h\ |
| ɬ | voiceless alveolar lateral fricative | K |
| ɮ | voiced alveolar lateral fricative | K\ |
| ɹ | alveolar approximant | r\ |
2. The two special characters
Two ASCII characters do structural work, and understanding them is most of understanding X-SAMPA.
The backslash — the “universal diacritic”
Section 12 of the paper defines \ as meaning “the preceding character is to be interpreted in a special way”, which roughly doubles the available symbol space by giving two-place codes.1 The syntax is fixed: the backslash attaches to the single immediately preceding character, and is written after it. Attaching it to more than one character is a syntax error, not a stylistic variant.
The backslash in action: base symbol versus backslash form
| Code | Value | Code | Value |
|---|---|---|---|
| G | ɣ — voiced velar fricative | G\ | ɢ — voiced uvular plosive |
| R | ʁ — voiced uvular fricative/approximant | R\ | ʀ — voiced uvular trill |
| ? | ʔ — glottal stop | ?\ | ʕ — voiced pharyngeal fricative |
| r | r — alveolar trill | r\ | ɹ — alveolar approximant |
| @ | ə — schwa | @\ | ɘ — close-mid central unrounded |
| 3 | ɜ — open-mid central unrounded | 3\ | ɞ — open-mid central rounded |
The underscore — the diacritic introducer
Section 13 assigns _ (ASCII 95) the role “interpret the following character as a diacritic”.1 So t_h is aspirated t, t_w is labialized t, and d_n is d with nasal release. Ejectives are _> (alternatively _?) and implosives are _<.1
Common underscore diacritics
| Code | Meaning | Code | Meaning |
|---|---|---|---|
| _h | aspirated | _> | ejective |
| _w | labialized | _< | implosive |
| _n | nasal release | _j | palatalized |
| _T | extra-high tone | _B | extra-low tone |
| _R | rising tone | _F | falling tone |
The precedence rule. Where \ and _ co-occur, the backslash takes priority. t_?\ parses as t plus pharyngealization, because ?\ — the pharyngeal-fricative code — binds first.1 Reversing this is one of the classic X-SAMPA errors.
3. Two SAMPA marks that X-SAMPA redefined
This is where reading old SAMPA files with X-SAMPA eyes goes wrong.1
- The apostrophe (ASCII 39) becomes the palatalization diacritic. In SAMPA it had denoted rising tone — a use Wells notes was unpopular.
- The grave accent / backtick (ASCII 96) becomes rhoticity and retroflexion: t` is retroflex t, @` is r-coloured schwa. In SAMPA it had denoted falling tone.
Note in passing that the retroflex approximant is a composite of the two mechanisms: r\` is the alveolar approximant r\ plus rhoticity.13
4. Prosody
Tone and prosody can be handled two ways.1 Either as underscore diacritics — _T extra-high through _B extra-low, _R rising, _F falling — or on a separate tier delimited by angle brackets < >, an approach that draws on the companion SAMPROSA proposals.14
5. The trap: ASCII schemes that look alike and are not
X-SAMPA is not the only ASCII coding of the IPA, and the others are not compatible with it. Kirshenbaum (used by eSpeak) and WorldBet assign different ASCII characters to the same IPA symbols, and — much worse — assign the same ASCII characters to different IPA symbols. The table below is drawn from this project's cross-system divergence dataset and shows only the rows where the schemes actually disagree.5
Where the ASCII schemes diverge — the rows to check before trusting a file
| IPA | Name | X-SAMPA | SAMPA | Kirshenbaum | WorldBet |
|---|---|---|---|---|---|
| ʈ | voiceless retroflex plosive | t` | — | t. | tr |
| ɖ | voiced retroflex plosive | d` | — | d. | dr |
| ɟ | voiced palatal plosive | J\ | — | J | J |
| ɢ | voiced uvular plosive | G\ | — | G | Q |
| ɱ | labiodental nasal | F | F | M | — |
| ɳ | retroflex nasal | n` | — | n. | nr |
| ɲ | palatal nasal | J | J | n^ | — |
| ɴ | uvular nasal | N\ | — | n" | — |
| ʙ | bilabial trill | B\ | — | b<trl> | — |
| r | alveolar trill | r | r | r<trl> | r |
| ʀ | uvular trill | R\ | — | r" | — |
| ɾ | alveolar tap | 4 | — | * | — |
| ɽ | retroflex flap | r` | — | *. | — |
| ɸ | voiceless bilabial fricative | p\ | — | P | F |
| β | voiced bilabial fricative | B | B | B | V |
| ʂ | voiceless retroflex fricative | s` | — | s. | sr |
| ʐ | voiced retroflex fricative | z` | — | z. | zr |
| ʝ | voiced palatal fricative | j\ | — | C<vcd> | j^ |
| ɣ | voiced velar fricative | G | G | Q | G |
| ʁ | voiced uvular fricative | R | R | g" | — |
| ħ | voiceless pharyngeal fricative | X\ | — | H | H |
| ʕ | voiced pharyngeal fricative | ?\ | — | H<vcd> | — |
| ɦ | voiced glottal fricative | h\ | — | h<?> | hv |
| ɬ | voiceless alveolar lateral fricative | K | — | s<lat> | — |
| ɮ | voiced alveolar lateral fricative | K\ | — | z<lat> | — |
| ʋ | labiodental approximant | P | P | r<lbd> | — |
The practical consequence: an ASCII phonetic string is not self-identifying. Before parsing one, establish which scheme produced it. A file that renders as sense under X-SAMPA and as nonsense under Kirshenbaum has told you what it is; a file that renders as plausible but different sense under both has not.
6. Errors to avoid
- Mixing X-SAMPA with per-language SAMPA values. SAMPA is a set of language-specific tables; X-SAMPA is one universal mapping. The same ASCII glyph can differ between them.12
- Assuming any ASCII-IPA you meet online is X-SAMPA. Kirshenbaum and WorldBet are distinct schemes with conflicting assignments — see the table above.5
- Citing “CXS” (Conlang X-SAMPA) as standardized. CXS is community lore. It does not appear anywhere in Wells's defining document, and no authoritative specification exists. Label it explicitly as an unstandardized community variant, or leave it out.1
- Reversing the _ / \ precedence, or attaching a backslash to more than the single preceding character. Both violate §12–13 syntax.1
- Assuming eSpeak speaks X-SAMPA. It uses Kirshenbaum. “ASCII IPA tool” does not imply “X-SAMPA”.
7. Sources
- Wells, “Computer-coding the IPA: a proposed extension of SAMPA” (1995) — The defining document. Sections 1–27 are the source for the base layer, the backslash and underscore mechanisms, the precedence rule, the redefined SAMPA marks and the unique-parsability property.
- X-SAMPA — UCL summary page — The IPA-free HTML symbol listing, last revised 3 May 2000 when Unicode code points were added.
- SAMPA — official home page, University College London — The per-language scheme X-SAMPA generalizes.
- SAMPROSA — SAM Prosodic Transcription — The prosodic proposals X-SAMPA's angle-bracket tier draws on.
- Wikipedia — X-SAMPA — Secondary overview with a full symbol table.
Notes & Bibliography
- Wells, John C. “Computer-coding the IPA: a proposed extension of SAMPA.” Revised draft, 28 April 1995. University College London. Read in full for the LinguaCommons Priority-1 research folder; claims here are anchored to its numbered sections — §1 scope and proposal status, §2 and §4 the relation to SAMPA, §3 and §27 the base layer, §5 the apostrophe, §6 the backtick, §12 the backslash, §13 the underscore and the precedence rule, §14 and §18 ejectives and implosives, §20–24 prosody, §26 unique parsability. [source] ↩
- Wells, John C., maintainer. “SAMPA computer readable phonetic alphabet.” University College London, site last revised 25 October 2005. Cited for the per-language design that X-SAMPA generalizes. [source] ↩
- University College London. “X-SAMPA.” The IPA-free HTML summary and symbol listing, last revised 3 May 2000, when Unicode code points were added. [source] ↩
- University College London. “SAMPROSA — SAM Prosodic Transcription.” The prosodic transcription proposals underlying X-SAMPA's separate tone tier. [source] ↩
- LinguaCommons Cross-System Divergence Table, RESEARCH/TRANSCRIPTION-LINGUISTIC-DEVICES/Cross-System-Divergence-Table/divergence-table.csv — 91 rows aligning IPA against X-SAMPA, SAMPA, Kirshenbaum, WorldBet and ARPABET, compiled from each scheme's defining document. The comparison table in section 5 is generated from it. [source] ↩