Advanced English II (BFE19013)

COCA Hands-on Practice

Instructor: Miao Zhang 張淼

Today’s plan

Three parts

  1. Part 1 (~40 min): Review — what we learned in Weeks 1–2
  2. Short break
  3. Part 2 (~70 min): Step-by-step tutorial — searching the COCA corpus
  4. Short break
  5. Part 3 (~40 min): Your own mini research with COCA

Bring your laptop. You will need internet access.

What you will be able to do today

By the end of this session, you can:

  • explain what a corpus is and why linguists use it
  • search a real corpus (COCA) by yourself
  • use corpus evidence to answer a small research question

Part 1: Review

Quick warm-up

Without looking at your notes:

What is a corpus? How is it different from just “a lot of texts”?

Discuss with your neighbor for 2 minutes.

Applied linguistics, in one sentence

Linguistics studies language.

Applied linguistics studies how language is used, learned, taught, and analyzed in real life.

Corpus linguistics is one of its methods.

A corpus is MORE than a collection

A corpus is:

  • large enough to be useful
  • systematic and organized
  • machine-readable
  • designed for analysis
  • representative of a language variety or purpose

What makes a GOOD corpus?

  1. Representativeness — it represents the language variety it studies
  2. Finite size — it has a design and a collection goal
  3. Machine-readable — searchable and analyzable by computer
  4. Standard reference — other researchers can reuse it and repeat the study

Concordancing: Key Word In Context (KWIC)

Why use a corpus?

Even expert speakers:

  • do not know everything
  • may be biased by what they notice
  • cannot remember every pattern
  • cannot easily count language use

A corpus helps us test patterns in real data.

What a corpus cannot do

  • There is no negative evidence in a corpus.
  • Absence of evidence is not evidence of absence.
  • Corpora show patterns, but they do not explain everything.
  • Findings depend on the corpus and the research question.

Types of corpora

  • Sample vs. monitor — fixed in time vs. keeps growing
  • General vs. specialized — broad picture vs. one genre/group/domain
  • Written vs. spoken
  • Raw vs. annotated — plain text vs. tagged/labeled text

Quantitative vs. qualitative

Qualitative Quantitative (corpus)
Aim depth and context patterns and frequency
Sample small, purposeful large, systematic
Strength rich detail objective, generalizable
Weakness limited generalization less context

Quick check

You want to know: “Do Americans say ‘delicious’ or ‘tasty’ more often?”

  1. Is this a corpus question or an intuition question?
  2. What kind of corpus would you need?

Part 2: COCA step by step

What is COCA?

COCA = Corpus of Contemporary American English

  • one of the largest corpora of American English
  • over one billion words (1990–2020)
  • texts from TV/movies, spoken, fiction, magazines, news, academic writing, blogs, web pages
  • free to use at english-corpora.org/coca

Because it covers many registers, we can ask: where is a word used, how often, and with what other words.

Step 1: Register an account

  1. Go to english-corpora.org/coca
  2. Click the yellow ID icon (top right) → LOGIN page
  3. Click REGISTER under the “Log in” button
  4. Fill in name, email, password, country, category
  5. Accept the Terms and Conditions, type the colored letters
  6. Click SUBMIT, then confirm via the email from admin@english-corpora.org
  7. Close the browser, reopen it, and log in

Registration page

Do it now

Register your COCA account.

Check your spam folder if the confirmation email does not arrive within 2 minutes.

Raise your hand when you are logged in.

Step 2: Know the interface

  • Profile — your account and usage
  • Saved Words — words you saved in earlier searches
  • Virtual Corpora — build your own sub-corpus from COCA
  • History — your recent searches
  • Help — what each icon does

The four main tabs

Directly under the COCA name:

Tab What it shows
SEARCH where you type words and phrases
FREQUENCY how often your search appears
CONTEXT concordance lines (examples in sentences)
OVERVIEW what the corpus can do

Step 3: List search (the default)

  1. Open the SEARCH tab
  2. Type a word or phrase in the text box
  3. Click Find matching strings

Reading the results

The FREQUENCY tab lists all matching forms and their frequency.

  • Click the ★ to save a word to your Saved Words list
  • Click a form to see its CONTEXT (concordance lines)

Concordance lines

Each line shows your word in a real sentence, with the date, text type (magazine, spoken, fiction…), and source name.

Try it now

Search for linguistics with the List function.

  1. How many times does it appear in COCA?
  2. Look at 2–3 concordance lines. What does it mean?

Click Reset under the search box before your next search.

Trick 1: Wildcard *

Attach * to part of a word to find all words that start or end with it:

  • be* → words starting with “be”

With a space, be * → phrases starting with “be”.

A teacher looking for words with the prefix im- can search im*.

Trick 2: Compare words with |

Put | between words to compare them:

  • talk|speak

Trick 3: Lemmas with [ ]

Words have many forms: talk, talks, talking, talked.

Search [talk] to get all forms of the word, ranked by frequency.

Trick 4: Part of speech

Click [POS] next to the search box to limit the search to one part of speech.

Example: search run as a verb only → select verb.ALL.

Practice 1

Using the List function, compare the frequency of delicious and tasty.

  1. How would you type this into the search bar?
  2. Which word is used more often?

Step 4: Chart — frequency by section

The Chart tab shows where a word appears:

Enter a word and click See frequency by section.

Reading the chart

  • Click a section name (e.g. BLOG) to see its subcategories
  • Click a bar in the bottom row to see concordance lines from that section

Wildcards and | work here too.

Practice 2

Using the Chart function, search for the word linguistics.

Which section (register) uses this word most often?

Why do you think that is?

Step 5: Word — a corpus dictionary

The Word tab gives dictionary-style information plus corpus data:

  • definition and part of speech
  • frequency across sections
  • COLLOCATES — words that most often appear with it
  • CLUSTERS — common 2–4 word phrases containing it
  • CONCORDANCE — example lines with nearby words highlighted

Step 5: Word — a corpus dictionary

Practice 3

Using the Word function, search for amuse.

List the verb collocates of this word.

(Example format: entertain, smile, …)

Step 6: Collocates

Find the most common words next to your word.

Click the + button on the SEARCH tab → COLLOCATES.

Results are grouped by part of speech: NOUN, ADJ, VERB, ADV.

Practice 4

Using Collocates, search for fantastic.

Set the window to 1 word on the left and 3 words on the right.

  1. What is the most common collocate?
  2. What is its frequency in the corpus?

Step 7: Compare two words

COMPARE shows the collocates of two words side by side.

Enter one word in each box and click Compare words.

  • Default sorting: by ratio (most distinctive collocates first)
  • Click FREQUENCY to sort by raw frequency instead

Practice 5

Using Compare, compare the words eat and drink.

  1. What is the most common (distinctive) collocate for eat?
  2. What is the most common collocate for drink?

Were you surprised?

Step 8: KWIC

KWIC = Key Word In Context — the last option under the + menu.

Your word lines up in one column, with its context on both sides. Parts of speech are color-coded.

Sorting KWIC lines

Use the buttons at the top right:

  • L — sort by the words on the left
  • R — sort by the words on the right (default: 3 words right)
  • * — reset the sorting

Sorting this way makes patterns easy to see.

Practice 6

Using KWIC, search for take.

Sort the results. What word most frequently appears immediately after “take”?

Step 9: Narrow your search

Under the search button of each menu:

  • Sections — search only certain registers or time periods
    • List 1 = your target section(s); List 2 = sections to compare with
    • Hold ALT / Command to select more than one
  • Sort/Limit — sort results by frequency, relevance, or alphabetically; set a minimum frequency

Practice 7

Using Sections + List, search for book only in TV/Movies and Fiction.

Which section does it appear in more frequently (per million words)?

Step 10: How a real study works (Wechsler 2018)

  1. Research question — what do you want to know?
  2. Prediction — what do you expect?
  3. Search string — how can COCA find the pattern?
  4. Raw count — how many examples come up?
  5. Error check — are any examples wrong?
  6. Valid count — the clean total
  7. Discussion — do the numbers support your prediction?

Case study: the verb whine

Wechsler (2018) asked: How does the verb whine combine with other words?

Specifically:

  • Does whine use to for a person? (whined to me)
  • Does it use at for a non-person? (whined at the noise)
  • Can it be followed by that? (She whined that…)
  • Can it introduce a question? (whined what…)

Secret codes: COCA part-of-speech tags

Code Meaning Example
_v any verb form WHINE_v → whine, whines, whined, whining
_i preposition to_i → only the preposition to
_cst complementizer that that_cst → only that as a clause starter
VVI infinitive verb TO VVIto + base verb
xxx ... xxx pattern in order xxxWHINE_v that_cstxxx → “whine that”

Results for whine: summary table

Pattern Valid tokens found
WHINE (all verb forms) 2,915
WHINE + that-clause 67
WHINE + to-PP (to a person) 36
WHINE + at-PP 31
WHINE + to-PP + that-clause 4
WHINE + at-PP + that-clause 0
WHINE + infinitival phrase 6
WHINE + indirect question 0

What the numbers mean

The prediction was supported:

  • whine + to-PP + that-clause = 4
  • whine + at-PP + that-clause = 0

So whine can introduce a message with to, but not with at.

Also, whine reports assertions (whined that…) and commands (whined to be held), but it does not report questions — every question search returned 0.

Take-home message from Wechsler (2018)

  • Corpora can test very specific linguistic predictions.
  • Raw numbers can be messy. Always check examples.
  • A zero result is informative when your theory predicts rarity or impossibility.

Practice 8

Which of these are real manner-of-speaking uses of whine? Which are miscounts?

  1. She whined that the test was too hard.
  2. The drill whined to a halt.
  3. The puppy made little whines that turned into yips.
  4. He whined to me about the noise.

COCA search cheat sheet

You want… Type…
one word or phrase beautiful
words starting with “im-” im*
phrases starting with “be” be *
compare two words delicious|tasty
all forms of a word [work]

COCA search cheat sheet

You want… Type…
a specific part of speech word + [POS] menu
words near your word COLLOCATES menu
two words’ collocates COMPARE menu
your word in context KWIC menu
a specific grammar tag (advanced) word_v, to_i, that_cst
a pattern in exact order (advanced) xxxWHINE_v that_cstxxx

Break

10-minute break.

When you come back: make sure you can log in to COCA and run a List search.

Part 3: Your own mini research

The task

Work alone or in pairs.

  1. Choose one topic from the list (or propose your own)
  2. Turn it into a question you can answer with COCA
  3. Run your searches — use at least two different COCA functions
  4. Take notes and screenshots of your key results
  5. Prepare a 2–3 minute report for the class

What to report

Your mini report should answer:

  1. Question — what did you want to find out?
  2. Method — what did you search? Which functions did you use?
  3. Findings — what did the corpus show? (numbers + 1–2 example lines)
  4. Reflection — was the result surprising? What is the limitation of your search?

Topic list: word choice

  1. delicious vs. tasty — which is more frequent? In which registers? What foods do they describe?
  2. big vs. large — do they mean the same? Compare the nouns that follow each.
  3. say / tell / speak / talk — how do their uses differ? Check collocates and KWIC lines.
  4. make vs. do — which nouns go with make, which with do?

Topic list: grammar and form

  1. gonna / wanna — where do they appear: spoken, TV/movies, or academic writing?
  2. a lot of vs. lots of — which is more formal? Compare across registers.
  3. totally / completely / absolutely — what adjectives or verbs follow each one?
  4. -ful vs. -less words (use wildcards: *ful, *less) — which ending is more common?

Topic list: how people really speak

  1. very vs. really — which one is used more in speech vs. academic writing? What adjectives follow each?
  2. literally — read 10 concordance lines. Does it always mean “literally”?
  3. thank / thanks / thank you — which form is most frequent? Where is each one used?

Topic list: advanced (optional)

  1. manner-of-speaking verbs (whisper, shout, whine, mumble) — what kinds of sentences follow them? Try the _v tag.
  2. whine or another emotion verb — does it take that, to, or at? Use the secret codes if you dare.
  3. Your own question — any word or phrase you are curious about. Check with the instructor first.

Tips before you start

  • Write down your exact search strings — your study should be repeatable
  • Compare PER MIL, not just raw FREQ — sections have different sizes
  • Read some actual concordance lines — numbers alone can mislead you

Tips before you start

  • If you use POS tags, remember that COCA’s automatic tags are not perfect — read your examples
  • A surprising result is a good result — report what you found, not what you expected
  • A zero result can also be a finding: ask why

Research time

You have about 30 minutes to search and prepare your report.

Ask for help any time.

Share your findings

Each person/pair: 2–3 minutes.

  1. Your question
  2. Your searches
  3. Your findings (show a screenshot if you can)
  4. One thing that surprised you

Wrap-up

Today you:

  • reviewed what a corpus is and why we use it
  • learned to search COCA: List, Chart, Word, Collocates, Compare, KWIC
  • saw how a real study goes from question to results table
  • answered a real question with real language data

Corpus skills will be useful for your group presentation and final essay.

Learn more (optional)

Video tutorials:

Credits

Part 2 is adapted from:

  • Veloso, I. (2023). Corpus of Contemporary American English (COCA) Tutorial. In L. Goulart & I. Veloso (Eds.), Corpora in English Language Teaching: Classroom Activities for Teachers New to Corpus Linguistics. Montclair State University. Licensed under CC BY-NC 4.0.
  • Wechsler, S. (2018). How to use COCA. Unpublished tutorial, University of Texas at Austin. PDF in the practice folder.