Advanced English II (BFE19013)

Applications and Techniques of Corpus

Instructor: Miao Zhang 張淼

Today’s content

  1. Part 1: Your own corpus mini project
  2. Part 2: Spoken corpora
  3. Part 3: What corpora have changed

Part 1: Mini research

The task

Work in pairs.

  1. Choose one topic from the list (or propose your own)
  2. Run your searches in AntConc using different tools
  3. Take notes and screenshots of your key results
  4. Prepare a 5-minute report for the class

What to report

Your mini report should answer:

  1. Question — what did you want to find out?
  2. Method — what did you search? Which tools/settings did you use?
  3. Findings — what did the corpus show? (numbers + 1–2 example lines)
  4. Reflection — was the result surprising? What is the limitation of your search?

Topic list: words and phrases

  1. big vs. large — what kinds of nouns does each appear with?
  2. make vs. do — what phrase patterns does each verb form?
  3. important vs. significant vs. vital — compare the company each word keeps. Do they mean the same?
  4. -ful vs. -less words — which ending is more common in your corpus?

Topic list: patterns

  1. Frequent 2- or 3-word sequences — what formulaic sequences dominate the corpus?
  2. Distribution — pick a key term from this week’s reading. Is it spread evenly across texts, or concentrated?
  3. A word’s behavior — pick a verb and find the words it most strongly associates with.
  4. Regular vs. irregular verbs — which is more common in your corpus? Are there any patterns in the irregulars?

Tips before you start

  • Write down your exact search strings — your study should be repeatable
  • Look at type vs. token counts, and check range (how many files an item appears in)
  • Read actual concordance lines — numbers alone can mislead you
  • A surprising result is a good result — report what you found, not what you expected. Just ask “Why”.

Research time

You have about 30 minutes to search and prepare your report.

Ask for help any time.

Share your findings

Each pair: 5 minutes.

  1. Your question
  2. Your searches
  3. Your findings (show a screenshot if you can)
  4. One thing that surprised you

Listen to the other reports: did anyone’s findings connect to yours?

Debrief

As a class:

  • Which question produced the most convincing evidence? Why?
  • What made a search unrepeatable or hard to interpret?
  • What would you do differently with more time or a bigger corpus?

☕ Break

Part 2: Spoken corpora

What is a speech corpus?

A speech corpus is a structured collection of spoken language recordings, often accompanied by metadata and annotations.

  • used in linguistics, sociolinguistics, speech technology, and clinical research
  • may contain read or spontaneous speech, typical and atypical speech patterns

Why collect speech?

Speech production is influenced by a large number of factors:

  • language variety, accent, culture, speaking style
  • phonological, lexical, syntactic, discourse and conversational structures
  • individual characteristics and attitudes

→ Big data is needed to model all these sources of variation.

What are speech corpora for?

Theory-oriented

  • linguistic and sociolinguistic analysis, language diversity
  • speaker identity, spoken language acquisition, conversation analysis

What are speech corpora for?

Applied science-oriented

  • speech therapy / rehabilitation / clinical applications
  • automatic speech recognition (ASR), text-to-speech synthesis
  • speaker and emotion recognition, speech-to-speech translation

Types of speech corpora (1): speech and setting

Dimension Options Examples
Speech type read vs. spontaneous TIMIT, LibriSpeech vs. Switchboard, CALLHOME
Environment lab-recorded vs. in-the-wild TIMIT vs. CHiME, VoxCeleb
Modality audio-only vs. multimodal WSJ vs. AVICAR, IEMOCAP

Types of speech corpora (2): coverage

Dimension Options Examples
Language scope monolingual vs. multilingual/dialectal BNC vs. BABEL, Common Voice
Demographics age-specific; pathological speech CHILDES, ElderSpeech; UA-Speech, Torgo
Domain task dialogues; archives ATIS (flight booking), AMI (meetings); BBC archives

The corpus follows the question

The type of speech corpus you use depends on the research question or application.

Two more design variables:

  • annotation depth — transcription only (LibriSpeech) vs. detailed phone/intonation labels (TIMIT, Switchboard)
  • size / collection method — large crowdsourced (Common Voice) vs. small and niche (Buckeye)

Quick check

Get to know speech corpora: Mozilla Common Voice.

Good documentation

  • purpose, speaker information, speech type, spoken text, size and duration, file format, metadata, recording set-up, etc.

Sound familiar?

Record your own, or use an available corpus?

Available corpus Self-recorded corpus
Pros time-saving; large amounts of real data; searchable high-quality recordings; user-specific; controlled data and documentation
Cons variable recording quality; need to work on a subset; mainly mainstream languages; limited documentation; biased (gender, age, ethnicity) large time and work effort; fieldwork conditions; small amount of data

1 hour of conversation can take several days to transcribe manually.

Discuss

You want to build a small corpus of Macau students’ casual English conversations.

  • Record your own, or look for an available corpus?
  • Which two cons worry you most on either side?

Open-access speech corpora

  • LibriSpeech — 1,000 hours of American English, read speech (audiobooks)
  • Buckeye — 40 hours of American English, spontaneous interviews
  • CHILDES / PhonBank — collections of child speech
  • Mozilla Common Voice — 100+ languages, crowdsourced, read speech
  • DoReCo — ~50 languages, spontaneous fieldwork interviews
  • FLEURS — 102 languages, parallel read speech

From recording to usable data: annotation

Annotation = assigning labels to recorded speech:

  • utterance: Wie heisst du?
  • word: Wie, heisst, du
  • phones/segments: v, iː, h, aɪ, s, t, d, uː
  • speaker demographics, speech act, intonation (e.g. L+H*)

The hard part: where does each unit start and end?

Manual vs. automatic annotation

  • Manual — slow, time-consuming; needs inter-annotator consistency checks
  • Automatic — fast, efficient, consistent; but can produce errors
  • Semi-automatic — automatic first, then human correction where necessary

Forced alignment: transcript → time stamps

Forced alignment: take an orthographic transcription of an audio file and generate a time-aligned version, using a pronunciation dictionary to look up the phones of words. — Montreal Forced Aligner

Needs three inputs: transcript + pronunciation dictionary + audio signal.

Popular tools: Montreal Forced Aligner (MFA), DARLA, WebMAUS.

Let’s try forced alignment

Go to WebMAUS and try aligning a short audio file with its transcript.

Research showcase: vowel intrinsic pitch

A classic phonetic finding: high vowels (/i, u/) tend to have shorter duration than low vowels (/a/) — The rhyme in tea is shorter than tar.

Testable on VoxCommunis: a phonetic corpus derived from Mozilla Common Voice with word- and phone-level alignments (via MFA) — 75 languages, 16,000+ hours.

The results

Why spoken corpora matter for us

Spoken corpora let us study real speech:

  • hesitations, false starts, overlap, vague language
  • discourse markers and response tokens
  • the grammar of talk — not just the grammar of writing

☕ Break

Part 3: What corpora have changed

Use your imagination

Now that you know what corpora are and how to use them, imagine what could have changed in your English learning?

  • Lexicography (documenting words and their meanings/uses)
  • Grammar (documenting patterns of usage)
  • Stylistics and translation (documenting literary style and cross-linguistic patterns)

Lexicography

From Johnson to COBUILD

  • 1750s: Samuel Johnson manually collated samples of usage (1560–1660) for the first comprehensive English dictionary.
  • 1980: the COBUILD project (Birmingham, John Sinclair) — the first large corpus-based dictionary (Collins COBUILD English Dictionary, 1987)
  • Today’s publisher corpora: billions of words, constantly updated (e.g. Cambridge International Corpus (CIC) > 1 billion)

Grammar

  • COBUILD: “pattern grammar” — the interface of lexis and grammar (Hunston & Francis 2000)
  • Biber et al. (1999): grammar of conversation, fiction, news, academic prose (40M words, 7 years)
  • Carter & McCarthy (2006): separates spoken from written grammar; 7,000+ examples with sound clips

Stylistics and translation

Stylistics

  • Corpus methods + close reading: Louw’s (1993) collocational analysis of literary texts
  • Semantic prosody as a tool for studying creativity and irony in fiction

Translation: three corpus types (Aston 1999)

Type Composition Use
Monolingual texts in one language (source or target) translator training, terminology
Comparable similarly designed corpora of two+ languages what is specific to translated text
Parallel originals + their translations (uni- or bidirectional) bilingual lexicography, machine translation, training

Forensic linguistics and sociolinguistics

Forensic linguistics

Corpus methods applied to:

  • authorship identification and plagiarism
  • genuineness of confessions, suicide notes, threat letters
  • courtroom discourse (Cotterill’s corpus of the entire O. J. Simpson trial)
  • Boucher (2005): truthful vs. deceptive recounts differed significantly in hesitation, lexical repetition, utterance length

Sociolinguistics

Corpora are built around social variables — age, gender, class, region:

  • COLT (Bergen Corpus of London Teenage Language): like as discourse marker, tags, taboo language
  • SCOTS: Scottish English, Scots, Gaelic and community languages
  • gender studies: apologies (Aijmer 1995), linguistic sexism (Holmes 2001)

Corpora and language teaching

Textbooks vs. real data

Corpora have shown that intuitions about language are often faulty — including intuitions behind textbooks:

  • Holmes (1988): ESL textbooks over-teach modal verbs, under-teach alternative epistemic strategies
  • Carter (1998): textbook dialogues lack discourse markers, vague language, ellipsis, hedges
  • Gilmore (2004): scripted dialogues differ from real interaction in turn length, false starts, pausing, overlap, hesitation

Learner corpora and DDL

  • ICLE (International Corpus of Learner English): 2M+ words, 19 L1 backgrounds, error-coded
  • Even advanced learners struggle with high-frequency verbs (make) — over-/under-use relative to native norms (Altenberg & Granger 2001)
  • Data-driven learning (DDL), Computer Assisted Language Learning (CALL): learners explore corpora to discover patterns and test hypotheses

“Every student is Sherlock Holmes.” — Johns (2002: 108)

Issues and debates

Authenticity: a real debate

For authentic corpus data:

  • real language, real motivation

Against / cautions:

  • corpus examples are decontextualised — wrenched from their original audience and setting
  • may be culturally opaque to learners
  • contrived examples are easier to grade by level

Discuss

What is your biggest fear/concern in using English more often?

  • Lack of fluency? Lack of accuracy? Lack of vocabulary? Lack of confidence? Lack of cultural knowledge?
  • Lack of people to talk to? Lack of opportunities to practice?

How would you help yourself improve using corpus methods (both text and spoken corpora)?

Discuss

Why should you aim for (non-)native like English?

Discuss with your partner and write down your ideas.

Write down both the advantages and disadvantages of targeting native-like English.

The native speaker in the classroom

  • Learners often target native-speaker English (Timmis 2002) — but which native variety? (AmE? BrE? Irish? Singaporean?)
  • Non-native users now outnumber native speakers (Graddol 1998)
  • ELF features that teachers usually correct — dropped 3rd-person -s, missing articles, discuss about — are typically unproblematic for ELF communication

SUEs: Successful Users of English

Prodromou’s proposal: judge learners as expert users, not failed native speakers.

  • Idiomatic chunks come free to native speakers but block even proficient L2 users
    • “Got up the wrong side o’ the bed”
  • Successful L2 communicators strategically use resources differently from native speakers

The goal of learning is to be a good human communicator — with native speakers or with anyone.

Final discussion

Macau context: students will mostly use English with other non-native speakers (mainland Chinese, Korean, Japanese, European partners).

  1. Should our classrooms target native-speaker models, ELF, or SUE-style expert use?
  2. What would each choice change about materials, assessment, and pronunciation teaching?