Lexical density is the ratio of content words (nouns, main verbs, adjectives, adverbs) to the total number of words in a text. It was first proposed by Ure (1971) and developed by Halliday (1985) as a key measure distinguishing written from spoken language.
The basic formula: LD = (number of content words / total words) x 100
A spoken conversation might have a lexical density of 35-45%, while an academic journal article might reach 55-65%. The difference is not random; it reflects fundamental differences in how meaning is packaged in speech versus writing.
Halliday (1985) identified a complementary relationship:
| Feature | Written language | Spoken language |
|---|---|---|
| Lexical density | High: meaning packed into content words | Low: meaning spread across more function words |
| Grammatical intricacy | Low: fewer clauses per sentence | High: more clause chaining and embedding |
Written language says: The rapid industrialisation of previously agricultural regions... Spoken language says: These regions used to be agricultural but then they industrialised really quickly and...
The written version is lexically dense (many content words per clause) but grammatically simple (one clause). The spoken version is lexically sparse but grammatically intricate (multiple clauses). Both convey similar content through different packaging strategies.
Lexical density alone does not determine quality. A text can be dense but incoherent, or sparse but perfectly effective for its context. It is one dimension of register variation, best interpreted alongside other measures like grammatical intricacy, cohesive density, and clause complexity.