Applications and Techniques of Corpus
Work in pairs.
Your mini report should answer:
You have about 30 minutes to search and prepare your report.
Ask for help any time.
Each pair: 5 minutes.
Listen to the other reports: did anyone’s findings connect to yours?
As a class:
A speech corpus is a structured collection of spoken language recordings, often accompanied by metadata and annotations.
Speech production is influenced by a large number of factors:
→ Big data is needed to model all these sources of variation.
Theory-oriented
Applied science-oriented
| Dimension | Options | Examples |
|---|---|---|
| Speech type | read vs. spontaneous | TIMIT, LibriSpeech vs. Switchboard, CALLHOME |
| Environment | lab-recorded vs. in-the-wild | TIMIT vs. CHiME, VoxCeleb |
| Modality | audio-only vs. multimodal | WSJ vs. AVICAR, IEMOCAP |
| Dimension | Options | Examples |
|---|---|---|
| Language scope | monolingual vs. multilingual/dialectal | BNC vs. BABEL, Common Voice |
| Demographics | age-specific; pathological speech | CHILDES, ElderSpeech; UA-Speech, Torgo |
| Domain | task dialogues; archives | ATIS (flight booking), AMI (meetings); BBC archives |
The type of speech corpus you use depends on the research question or application.
Two more design variables:
Get to know speech corpora: Mozilla Common Voice.
Sound familiar?
| Available corpus | Self-recorded corpus | |
|---|---|---|
| Pros | time-saving; large amounts of real data; searchable | high-quality recordings; user-specific; controlled data and documentation |
| Cons | variable recording quality; need to work on a subset; mainly mainstream languages; limited documentation; biased (gender, age, ethnicity) | large time and work effort; fieldwork conditions; small amount of data |
1 hour of conversation can take several days to transcribe manually.
You want to build a small corpus of Macau students’ casual English conversations.
Annotation = assigning labels to recorded speech:
The hard part: where does each unit start and end?
Forced alignment: take an orthographic transcription of an audio file and generate a time-aligned version, using a pronunciation dictionary to look up the phones of words. — Montreal Forced Aligner
Needs three inputs: transcript + pronunciation dictionary + audio signal.
Popular tools: Montreal Forced Aligner (MFA), DARLA, WebMAUS.
Go to WebMAUS and try aligning a short audio file with its transcript.
A classic phonetic finding: high vowels (/i, u/) tend to have shorter duration than low vowels (/a/) — The rhyme in tea is shorter than tar.
Testable on VoxCommunis: a phonetic corpus derived from Mozilla Common Voice with word- and phone-level alignments (via MFA) — 75 languages, 16,000+ hours.
Spoken corpora let us study real speech:
Now that you know what corpora are and how to use them, imagine what could have changed in your English learning?
| Type | Composition | Use |
|---|---|---|
| Monolingual | texts in one language (source or target) | translator training, terminology |
| Comparable | similarly designed corpora of two+ languages | what is specific to translated text |
| Parallel | originals + their translations (uni- or bidirectional) | bilingual lexicography, machine translation, training |
Corpus methods applied to:
Corpora are built around social variables — age, gender, class, region:
Corpora have shown that intuitions about language are often faulty — including intuitions behind textbooks:
“Every student is Sherlock Holmes.” — Johns (2002: 108)
For authentic corpus data:
Against / cautions:
What is your biggest fear/concern in using English more often?
How would you help yourself improve using corpus methods (both text and spoken corpora)?
Why should you aim for (non-)native like English?
Discuss with your partner and write down your ideas.
Write down both the advantages and disadvantages of targeting native-like English.
Prodromou’s proposal: judge learners as expert users, not failed native speakers.
The goal of learning is to be a good human communicator — with native speakers or with anyone.
Macau context: students will mostly use English with other non-native speakers (mainland Chinese, Korean, Japanese, European partners).
Advanced English II