CALLHOME American English Lexicon (PRONLEX) Second Edition

Full Official Name: CALLHOME American English Lexicon (PRONLEX) Second Edition
Submission date: July 1, 2026, 6:44 p.m.

**Introduction** CALLHOME American English Lexicon (PRONLEX) Second Edition (LDC2026L05) was developed by the Linguistic Data Consortium (LDC) and contains 90,988 English words with citation-form pronunciations. This second edition updates file formats, directory structure and documentation. The first edition is available as CALLHOME American English Lexicon (PRONLEX) (LDC97L20). The CALLHOME series consists of telephone conversations, transcripts and lexicons developed by LDC and Rutgers, The State University of New Jersey, in support of research in speaker identification, language identification and related technologies. Languages in the series include American English, Egyptian Arabic, German, Japanese, Mandarin Chinese, and Spanish. **Data** The words in the lexicon were derived from Wall Street Journal text used in the continuous speech recognition publication series CSR-1 WSJ0 Complete (LDC93S6A), transcripts from the Switchboard telephone collection (LDC97S62), and transcripts representing unscripted telephone conversations between native American English speakers contained in CALLHOME American English Second Edition (LDC2026S08). PRONLEX transcription is a phonemic transcription system designed to support speech recognition by providing a consistent and simplified representation of how words are pronounced in standard American English that allows variation to be generated later to avoid listing many pronunciation variations for each word. This single systematic base form can be expanded through rules or modeling. The transcription was created using a modified ARPABET phoneme set. The lexicon contains three tab-separated information fields: (1) word: orthographic representation of word; (2) pron: transcribed citation-form pronunciations using modified ARPABET phoneme set; and (3) comments: (OPTIONAL) comment on the entry; The lexicon is presented as a tab-delimited TSV file encoded in UTF-8 format. This release also includes a pronunciation dictionary derived from the lexicon in UTF-8 encoded CMUdict format.

Creator(s)
Distributor(s)
Right Holder(s)