Skip to content
LinguaCommons
Back

Notation Systems

How linguists write sound down. These guides cover the International Phonetic Alphabet and its clinical extensions, the ASCII encodings built for early computing and speech technology (SAMPA, X-SAMPA, Kirshenbaum, WorldBet, ARPABET), the scholarly transcription traditions behind the documentation of whole language families (Americanist, Uralic, Slavicist, Teuthonista), the prosody-annotation framework ToBI, the signed-language notation HamNoSys, and the Praat program's input conventions — what each system is for, who maintains it, and the errors to avoid when reading or citing it.

Fundamentals

The International Phonetic Alphabet and its official extension for disordered and atypical speech. Read these first — every other system on this page is defined by reference to the IPA, or was built to encode it.

Other Notation Systems

ASCII encodings, scholarly transcription traditions, prosody and signed-language notation, and toolchain-specific conventions.

Americanist Phonetic Notation (NAPA)

Americanist Phonetic Notation — also called NAPA, the North American Phonetic Alphabet — is not a single chartered alphabet but a family of transcription conventions developed by North American anthropological linguists for documenting the Indigenous languages of the Americas. Its lineage runs from the Powell-era Bureau of American Ethnology alphabets through Boas to the famous 1934 "Some orthographic recommendations" (Herzog, Newman, Sapir, the Swadeshes & Voegelin, American Anthropologist 36(4): 629–631) — and it lives on, modernized, in Americanist journals and in many practical orthographies today.[^1][^2]

ARPABET

ARPABET (also spelled ARPAbet) is a set of ASCII codes for the phonemes of General American English. It was developed by ARPA in the 1970s as part of the Speech Understanding Research (SUR) project, and it does one specific job well: it lets you write the sounds of English words using only plain keyboard characters — AA, AE, IY, UW, CH, JH, with digits for stress. The single most important thing to understand about it is what it is not: it is a phoneme code for one language, not a general phonetic alphabet like the IPA.[^1][^2][^3]

HamNoSys (Hamburg Sign Language Notation System)

HamNoSys — the Hamburg Sign Language Notation System (German Hamburger Notationssystem für Gebärdensprachen) — is a notation system for the form of signed-language signs. The team that maintains it, the Institute of German Sign Language at the University of Hamburg, describes it as a "phonetic" transcription system in the tradition of Stokoe notation, expanded so that it is not tied to the assumptions of American Sign Language and can be applied to signed languages internationally.[^1][^2]

Kirshenbaum (ASCII-IPA)

Kirshenbaum — also called ASCII-IPA or erkIPA — is an ASCII encoding of the International Phonetic Alphabet: a way of writing IPA using only the characters on a plain keyboard. It was developed by Evan Kirshenbaum and collaborators between about 1991 and 1993 on Usenet (the newsgroups sci.lang and alt.usage.english), so that people could discuss pronunciation in email and newsgroup posts before Unicode made real IPA glyphs easy to type. The defining document is Kirshenbaum's spec circulated in that period; a widely-referenced "second draft" appeared in sci.lang in January 1993.[^1][^2]

Praat Phonetic Encoding (Backslash Trigraphs)

Praat's "phonetic encoding" is different in kind from the other systems in this hub. It is not a transcription standard and not a rival alphabet — it is the input and rendering convention of the Praat program, the dominant free phonetics software, written by Paul Boersma and David Weenink at the University of Amsterdam. The right way to describe it is simply: this is how you type IPA into Praat.[^1][^2]

SAMPA — the Speech Assessment Methods Phonetic Alphabet

SAMPA — the Speech Assessment Methods Phonetic Alphabet — is a machine-readable phonetic alphabet that maps IPA symbols onto printable 7-bit ASCII, in the range 33–127. It exists because 1980s and 1990s computing could not render IPA glyphs: e-mail, databases and speech technology systems all needed phonetic transcription that would survive a 7-bit channel intact.[^1] Unicode has since removed that constraint, but SAMPA has not gone away, because an enormous quantity of speech-technology data was transcribed in it and still has to be read.

Shriberg & Kent (S-K) Clinical Transcription

The Shriberg & Kent (S-K) system is the transcription framework taught in Clinical Phonetics, the standard American textbook for training speech-language pathologists in phonetic transcription. It is not a rival alphabet: it is the IPA plus a set of clinical conventions and — just as importantly — a structured, audio-based method for teaching people to transcribe disordered speech reliably. Its identity is as much a training tradition (the famous transcription drills) as a symbol set.[^1][^2]

Teuthonista (German Dialectology Transcription)

Teuthonista is the standard phonetic transcription system of German dialectology — and of adjacent Germanic and Romance dialect work. It takes its name from the journal Teuthonista (Zeitschrift für deutsche Dialektforschung und Sprachgeschichte), in which the system was presented in 1924/25. Where the IPA aims at one symbol per sound, Teuthonista does the opposite: it keeps familiar Latin letters and loads fine articulatory detail onto them with a dense stack of diacritics.[^1][^2]

The Slavicist Phonetic Alphabet (Slavic Transcription Tradition)

Unlike the IPA or SAMPA, the "Slavic phonetic alphabet" is not a single chartered specification with one authority and one document. The name is used in practice for the Slavicist transcription tradition: the Latin-based phonetic notation that Slavic dialectologists, etymologists, and comparative grammarians have used for over a century. Anyone reading Slavic historical-comparative literature meets it constantly, so cross-reading it with the IPA is a genuine skill — but it is a tradition with variants, not a fixed standard, and honest guides say so.[^1][^2]

ToBI (Tones and Break Indices)

ToBI — Tones and Break Indices — is not a phonetic alphabet. It is a system for transcribing intonation and prosody: the melody of speech, and how tightly words are joined together. It was created by speech scientists in the early 1990s to give prosody the kind of shared standard that the IPA gives to segments, so that prosodically labelled speech databases could be exchanged between research sites. The original design target was English, specifically Mainstream American English.[^1][^2]

Uralic Phonetic Alphabet (UPA / Finno-Ugric Transcription)

The Uralic Phonetic Alphabet — in Finnish suomalais-ugrilainen tarkekirjoitus, "Finno-Ugric fine transcription" — is the standard transcription system of Uralic studies. It was first published by E. N. Setälä in 1901 and, well over a century later, is still the notation of Finno-Ugrist journals, etymological dictionaries, and dialect archives. If the IPA is the alphabet of general phonetics, the UPA is the house notation of one deeply documented language family.[^1][^2]

VoQS — Voice Quality Symbols

VoQS (Voice Quality Symbols) is the standard notation for voice quality in transcription — the long-domain settings that colour whole stretches of speech rather than single segments: whisper, creak, falsetto, breathiness, harshness, laryngeal and supralaryngeal settings (labialized or nasalized speech settings), and special airstreams such as œsophageal and electrolarynx speech.[^1] Where the IPA describes segments and extIPA describes atypical segments, VoQS answers the third question a clinician or phonetician must ask: what is the voice doing over this whole stretch?

WorldBet

WorldBet is James L. Hieronymus's ASCII encoding of the IPA — plus extra broad-phonetic symbols — designed to cover all the world's languages for speech-database labelling. It was defined in an AT&T Bell Laboratories technical memorandum, "ASCII Phonetic Symbols for the World's Languages: Worldbet" (c. 1993). Where earlier ASCII schemes were built around European languages, WorldBet's stated motivation was that those schemes "left out many of the sounds of the other languages" — so it aimed to be genuinely multilingual from the start.[^1]

X-SAMPA — Wells's ASCII encoding of the full IPA

X-SAMPA is John C. Wells's 1995 proposal for a keyboard-compatible ASCII coding of the entire IPA — in his words, “everything on the 1993 IPA Chart, including diacritics and tone marks” — designed so that IPA-transcribed material could be sent through e-mail and other 7-bit channels without loss.[^1] Where SAMPA is a family of per-language tables, X-SAMPA is one universal mapping.[^1][^2] That difference is the whole point of the scheme, and it is why X-SAMPA, not SAMPA, became the ASCII-IPA that people actually use for general phonetic work.

Notation is a way of writing distinctions down; the terminology section explains what those distinctions are.