labvanced logoLabVanced
  • Research
    • Publications
    • Researcher Interviews
    • Use Cases
      • Developmental Psychology
      • Linguistics
      • Clinical & Digital Health
      • Educational Psychology
      • Cognitive & Neuro
      • Social & Personality
      • Arts Research
      • Sports & Movement
      • Marketing & Consumer Behavior
      • Economics
      • HCI / UX
      • Commercial / Industry Use
    • Labvanced Blog
    • Services
  • Technology
    • Feature Overview
    • Code-Free Study Building
    • Eye Tracking
    • Mouse Tracking
    • Generative AI Integration
    • Multi User Studies
    • More ...
      • Reaction Time/Precise Timing
      • Text Transcription
      • Heart Rate Detection (rPPG)
      • Emotion Detection
      • Questionnaires/Surveys
      • Experimental Control
      • Data Privacy & Security
      • Desktop App
      • Mobile App
  • Learn
    • Guide
    • Videos
    • Walkthroughs
    • FAQ
    • Release Notes
    • Documents
    • Classroom
  • Experiments
    • Cognitive Tests
    • Sample Studies
    • Public Experiment Library
  • Pricing
    • Licenses
    • Top-Up Recordings
    • Subject Recruitment
    • Study Building
    • Dedicated Support
    • Checkout
  • About
    • About Us
    • Contact
    • Downloads
    • Careers
    • Impressum
    • Disclaimer
    • Privacy & Security
    • Terms & Conditions
    • Third-Party Licenses
  • Appgo to app icon
  • Logingo to app icon
Learn
Guide
Videos
Walkthroughs
FAQ
Newsletter Archive
Documents
Classroom
  • 中國人
  • Deutsch
  • Français
  • Español
  • English
  • 日本語
Guide
Videos
Walkthroughs
FAQ
Newsletter Archive
Documents
Classroom
  • 中國人
  • Deutsch
  • Français
  • Español
  • English
  • 日本語
  • Guide
    • GETTING STARTED

      • Task Editor
      • Frames
        • Canvas Frame
        • Page Frame
        • Website Frame
      • Trial Timeline
      • Stimulus Presentation
      • Correctness of Response
      • Objects
      • Events
      • Variables
      • Task Wizard
      • Trial System
      • Experiment Design
        • Creating a Study
        • Data Frame Patterns
        • Multi-user Experiment Design
        • Longitudinal Experiment Design
    • FEATURED TOPICS

      • Randomization & Balance
      • Eye Tracking
        • Presenting Stimuli in an Eye Tracking Task
        • How to Choose Eye Tracking Calibration Settings
        • Setting Up Eye Tracking in a Task
        • Recording and Exporting Gaze Data
        • Using Shapes as Areas of Interest (AOI) for Eye Tracking
        • The Participant Calibration Experience
      • Web Bridge
        • Selecting Website AOIs | Anchor, Adjust, and Verify AOI Tracking
        • Tracking Mouse Behavior on a Website AOI | Web Bridge Guide
      • Questionnaires
      • Desktop App
      • Sample Studies
      • Participant Recruitment
      • API Access
        • REST API
        • Webhook API
        • WebSocket API
      • Other Topics

        • Precise Stimulus Timings
        • Multi User Studies
        • Head Tracking in Labvanced | Guide
    • MAIN APP TABS

      • Overview: Main Tabs
      • Dashboard
      • My Studies
      • Shared Studies
      • My Files
      • Experiment Library
      • My Account
      • License & Services
    • STUDY TABS

      • Overview: Study-Specific Tabs
      • Study Design
        • Tasks
        • Blocks
        • Sessions
        • Groups
      • Task Editor
        • Task Controls
        • The Trial System
        • Frame Types
        • Objects
        • Object Property Tables
        • Variables
        • System Variables Table
        • The Event System
        • Text Editor Functions
        • Eye Tracking Options in the Task Editor
        • Head Tracking in a Task
        • Multi-User Studies
      • Settings
      • Variables
      • Media
      • Texts & Translate
      • Launch & Participate
      • Subject Management
      • Dataview and Export
        • Dataview and Variable & Task Selection (OLD Version)
        • Accessing Recordings (OLD Version)
  • Videos
    • Video Overview
    • Getting Started in Labvanced
    • Creating Tasks
    • Element Videos
    • Events & Variables
    • Advanced Topics
  • Walkthroughs
    • Introduction
    • Stroop Task
    • Lexical Decision Task
    • Posner Gaze Cueing Task
    • Change Blindness Flicker Paradigm
    • Eye-tracking Sample Study
    • Infant Eye-tracking Study
    • Attentional Capture Study with Mouse Tracking
    • Rapid Serial Visual Presentation
    • ChatGPT Study
    • Eye Tracking Demo: SVGs as AOIs
    • Multi-User Demo: Show Subjects' Cursors
    • Gamepad / Joystick Controller- Basic Set Up
    • Desktop App Study with EEG Integration
    • Between-subjects Group Balancing and Variable Setup
    • Building a Multi-user Study
    • Website Frame Eye Tracking Walkthrough
    • Conditional Webcam Eye Tracking
    • Text Segmentation Study Building
  • FAQ
    • Features
    • Support Policy & Guidelines
    • Security & Data Privacy
    • Licensing
    • Precision of Labvanced
    • Programmatic Use & API
    • Using Labvanced Offline
    • Troubleshooting
    • Study Creation Questions
  • Newsletter Archive
  • Documents
  • Classroom

Word Frequency Effect

Read the words "chair" and "urn." Both are short, both are nouns, both name an object you could sit near. Yet your brain flags "chair" as a word faster than it flags "urn." Psychologists call this gap the word frequency effect: the more often a word turns up in the language around us, the faster and more accurately we recognize it.


Table of Contents

  • What Is the Word Frequency Effect
  • How the Word Frequency Effect Works
    • The Logogen Model
    • The Interactive Activation Model
    • The Discriminative-Learning Account
  • How Word Frequency Is Measured
  • Examples of the Word Frequency Effect in Research
    • Marketing and Persuasion
    • Psycholinguistics and Lexical Decision
    • Reading Development
    • Bilingual and Second-Language Word Recognition
  • Frequently Asked Questions

What Is the Word Frequency Effect

Try it yourself: "table" and "trellis." Both nouns, both two syllables, both concrete objects. One of them you'll tag as a real word almost instantly. The other takes a beat longer. The only real difference between them is how often you've run into each one, and that difference alone is enough to change your reaction time.

Frequency: a word's raw count of occurrences in a real corpus of text or speech, the measurable stand-in for how "common" a word feels.

Baayen and colleagues call it "a strong predictor in most lexical processing tasks" (Heitmeier, Chuang, Axen & Baayen, 2023). That predictive power is well established. What's less settled is why it happens.

How the Word Frequency Effect Works

When you encounter a word, your brain has to retrieve everything stored about it: its spelling, its sound, its meaning. Word frequency changes how long that retrieval takes, high-frequency words come back fast, low-frequency words come back slower. Three accounts explain why, in three different vocabularies.

The Logogen Model

The classic version is Morton's Logogen Model. Every word you know has its own logogen, a recognition unit that fires once enough evidence for that word builds up. Frequency sets how much evidence is needed: a high-frequency word's logogen has a lower threshold, so it fires sooner. A low-frequency word's logogen needs more evidence, so it fires later (Lim, 2016).

The Interactive Activation Model

McClelland and Rumelhart's Interactive Activation Model builds this into a layered system: visual features feed into letters, letters feed into whole-word units. Instead of a fixed threshold, frequency sets each word unit's resting activation level. A high-frequency word's unit starts closer to firing. A low-frequency word's unit starts further away (Lim, 2016).

The Discriminative-Learning Account

A more recent, computational account frames it differently again: as a prediction problem, not an activation race. A system trained on the language you've actually encountered makes a smaller prediction error for words it has seen more often. High-frequency words: smaller error, faster recognition. Low-frequency words: larger error, slower recognition (Heitmeier, Chuang, Axen & Baayen, 2023).

All of the above theories have one thing in common: exposure determines how fast a word clears whatever bar the model is measuring. What all three agree on, and what matters for study design, is that the raw frequency count itself is a poor predictor. The log- or rank-transformed value tracks reaction time far more closely (Heitmeier, Chuang, Axen & Baayen, 2023). You'll see exactly what that transformed value looks like against a real word in the next section.

How Word Frequency Is Measured

Predictors need a number behind them. So where does that number actually come from?

A main way researchers measured word frequency for decades was counting occurrences across printed material: readers, textbooks, the Bible, the English classics, popular magazines, even a set of 120 children's books (Thorndike & Lorge, 1944).

But people don't talk like books. So why use books as the main basis to represent language frequency? Marc Brysbaert and Boris New rebuilt the standard American English word count from scratch, 51 million words pulled from film and television subtitles instead (Brysbaert & New, 2009).

The bet paid off. Tested against 31,201 real words, their subtitle-based measure, SUBTLEX-US, explained 62.3% of the variance in how fast people recognized them. The old norms, built from edited English prose printed in 1961 (Francis & Kučera, 1964), managed just 57.7% (Brysbaert & New, 2009).

A handful of real entries from that SUBTLEX-US dataset show the spread (Brysbaert & New, 2009), from words that show up once in 51 million words to "the," which shows up in every single one of the 8,388 films and shows in the corpus:

WordFREQcountSUBTLWFSUBTLCDLg10WF
enchantments10.020.010.3010
firming10.020.010.3010
septuplets10.020.010.3010
referent10.020.010.3010
punctiliously10.020.010.3010
pragmatism110.220.111.0792
humbled460.900.391.6721
exhibition2074.061.632.3181
faithful4659.124.212.6684
tool54810.754.402.7396
bottles68213.375.472.8344
personality81115.906.722.9096
lovely4,85395.1629.483.6861
dog9,835192.8436.353.9928
eyes11,299221.5556.824.0531
married12,163238.4947.114.0851
together19,553383.3976.904.2912
mind24,715484.6181.904.3930
Thanks31,789623.3185.234.5023
God46,061903.1682.744.6633
Where93,3411,830.2298.474.9701
Yes101,8351,996.7697.045.0079
know291,7805,721.1899.435.4651
the1,501,90829,449.18100.006.1766

New frequency measures based in the SUBTLEXUS database, filename Zipped Excel 2007 file with all 74,286 words in the corpus

  • FREQcount: the raw number of times the word occurs across the entire 51-million-word SUBTLEX-US corpus.
  • SUBTLWF: word frequency per million words (FREQcount divided by 51), the standardized figure that makes words from different-sized corpora comparable.
  • SUBTLCD: contextual diversity, the percentage of the corpus's 8,388 films and TV episodes that contain the word at least once. A word can be frequent but concentrated in a few contexts, or moderately frequent but spread everywhere, this is what captures that difference.
  • Lg10WF: the log10-transformed frequency. This is the value that actually predicts reaction time in lexical decision experiments, not the raw count, which is why researchers use it instead of FREQcount for modeling.

Note

There isn't one universal corpus for measuring word frequency. Which one a researcher uses depends on the language, the dictionary or frequency list in question, and the application. SUBTLEX-US is a widely used example in English-language linguistics research, and it's used here because it demonstrates the core concepts behind the word frequency effect clearly, not because it's the only option. Other corpora exist for other languages, genres, and research needs.

That's the variable, defined and measured. Where it actually shows up, in psycholinguistics labs, in how kids learn to read, in bilingual speakers, in marketing copy, is where things get interesting.

Examples of the Word Frequency Effect in Research

Marketing and Persuasion

Word frequency doesn't only stay in the lab, it shapes how people respond to language generally. Easy-to-process words just feel better, a phenomenon researchers call processing fluency, and word frequency is one of its most reliable levers (Bullock, Shulman & Huskey, 2021).

Brady Hodges, Zachary Estes, and Caleb Warren tested that directly on brand slogans: pulling apart over 800 real examples, using correlational analysis, lab experiments, eye-tracking, and a field study, they found slogans built from common words get liked more (Hodges, Estes & Warren, 2024). The implications of this is that slogans built from rare words are more likely to be remembered more!

The implications can be applied to real business decisions. A newer brand, or one fighting for market share, or breaking into unfamiliar territory, actually benefits from harder-to-process language, it's what makes people remember the name later. For an established brand, the math flips: being forgettable barely costs them anything, people already know who they are, but being disliked does real damage. That's why the safer bet for them is the easy, likable word, even if it's less memorable (Hodges, Estes & Warren, 2024).

Psycholinguistics and Lexical Decision

It sounds almost too simple to matter: look at a string of letters, decide if it's a real word, hit a button, as fast as you can. That trivial-feeling judgment is where word frequency's effect is largest, bigger than word length, neighborhood density, or anything else psycholinguists usually blame for slow reaction times.

Heitmeier and colleagues measured it directly: frequency explained "by far" the most variance in how fast people responded (Heitmeier, Chuang, Axen & Baayen, 2023).

In the video above, an example of the Dual Lexical Decision Task can be seen. This task presents the participant with two words and the participant must make a decision as to whether both presented words are in English or whether one is not.

Reading Development

Sarah Gerth and Julia Festman made use of eye tracking in a study with 66 German-speaking children in fifth and sixth grade, ages 10 to 12, as they read ordinary sentences, not isolated words on a screen, real sentences, read the way a child actually reads in school. The result: low-frequency words held their gaze a full 62 milliseconds longer than high-frequency ones, in readers who were still actively learning to read fluently (Gerth & Festman, 2021).

But frequency turned out to be only half the story. Some of the children were simply faster readers than others, and word frequency alone couldn't explain why.

Word length was the missing piece. The faster child readers had already picked up a habit the slower ones hadn't yet: perceiving a common word whole in one glance instead of working through it letter by letter, a small but real step toward how adults read.

A design modeled on this study, high- and low-frequency critical words embedded in ordinary sentences, matched on word length, with per-word gaze metrics recorded automatically, is demonstrated on Labvanced's Text Segmentation page, built with the Text Segmentation Object.

Bilingual and Second-Language Word Recognition

Here's a pattern most people never notice about their own bilingualism: the word frequency effect hits harder in your weaker language. Jens Schmidtke measured it, bilingual speakers show a bigger frequency effect than monolingual speakers do (Schmidtke, 2014).

It's tempting to read that as bilingual brains processing language differently. For this specific effect, they don't have to be: split your exposure between two languages and each one simply gets less practice than a monolingual speaker's one language gets.

Get more fluent, and the gap shrinks. Higher proficiency was tied directly to a smaller frequency effect.

Frequently Asked Questions

What is the word frequency effect in simple terms?
Common words, the ones you hear or read constantly, get recognized faster than rare ones, even when both words are the same length.
Why are common words recognized faster than rare ones?
Repetition. You've simply encountered common words far more often, and several models of word recognition describe how that repeated exposure sharpens and speeds up recognition specifically for high-frequency words.
How do researchers measure word frequency?
By counting occurrences across huge text or speech corpora. Modern English research leans on subtitle-derived norms like SUBTLEX-US, and analyzes reaction times against the log- or rank-transformed frequency value rather than the raw count.
Does the word frequency effect only apply to reading?
No. It shows up in visual word recognition, spoken word recognition, and picture-naming tasks, and the size of the effect shifts by task and, for bilingual speakers, by how much exposure they've had to the language in question.

Key Takeaways

  • Common words get recognized faster and more accurately than rare ones, a finding known as the word frequency effect.
  • It beats every other variable psycholinguists test, word length, neighborhood density, as a predictor of reaction time.
  • Use log- or rank-transformed frequency, not raw counts, raw frequency barely predicts behavior.
  • Bilingual speakers feel the effect harder than monolingual speakers, because their exposure to each language is split.
  • In brand slogans, common words win on likability and lose on memorability.
  • Labvanced's Dual Lexical Decision Task is a ready-made paradigm for studying this effect, recording trial-level reaction time and accuracy data.
  • For continuous-sentence reading instead of isolated words, the Text Segmentation Object records per-word gaze metrics automatically, demonstrated with a word-frequency reading study modeled on Gerth and Festman (2021).

More psychology concepts and effects, explained.

References

  1. Brysbaert, M., & New, B. (2009). Moving beyond Kučera and Francis: A critical evaluation of current word frequency norms and the introduction of a new and improved word frequency measure for American English. Behavior Research Methods, 41(4), 977–990.
  2. Bullock, O. M., Shulman, H. C., & Huskey, R. (2021). Narratives are persuasive because they are easier to understand: Examining processing fluency as a mechanism of narrative persuasion. Frontiers in Communication, 6, 719615.
  3. Francis, W. N., & Kučera, H. (1964). A Standard Corpus of Present-Day Edited American English (the Brown Corpus). Brown University.
  4. Gerth, S., & Festman, J. (2021). Reading development, word length and frequency effects: An eye-tracking study with slow and fast readers. Frontiers in Communication, 6, 743113.
  5. Heitmeier, M., Chuang, Y.-Y., Axen, S. D., & Baayen, R. H. (2023). Frequency effects in linear discriminative learning. Frontiers in Human Neuroscience, 17, 1242720.
  6. Hodges, B. T., Estes, Z., & Warren, C. (2024). Intel inside: The linguistic properties of effective slogans. Journal of Consumer Research, 50(5), 865–886.
  7. Lim, S. W. H. (2016). The influence of orthographic neighborhood density and word frequency on visual word recognition: Insights from RT distributional analyses. Frontiers in Psychology, 7, 401.
  8. Schmidtke, J. (2014). Second language experience modulates word retrieval effort in bilinguals: Evidence from pupillometry. Frontiers in Psychology, 5, 137.
  9. Thorndike, E. L., & Lorge, I. (1944). The Teacher's Word Book of 30,000 Words. Teachers College, Columbia University.