Free unique word counter
Unique word counter and vocabulary range
TextLab's unique word counter counts the distinct words in a piece of writing, sets that against its running total, and turns the pair into the measures editors and teachers actually quote: the type-token ratio, a version of it corrected for length, and the count of words used exactly once. Free to run, no account required, and it recalculates on every keystroke. Five bands show how the vocabulary is spread, so you can see whether variety is real or a side effect of a short sample.
- 100% free
- No signup
- Type-token ratio
- Corrected for length
- Once-only words listed
How often each word gets reused
Your vocabulary will be sorted into five bands here — the words used once, twice, and so on — which is a fuller picture of variety than one ratio can give.
How to measure vocabulary range
Step two is where almost every comparison between two texts goes wrong.
Paste a whole piece rather than an extract
Vocabulary measures are sensitive to length in a way that word counts are not, so an extract will not stand in for the essay it came from. Give it the complete piece — the full assignment, the whole story, every answer a candidate wrote — and note the running-word figure at the top, because it is the number you will need when you compare this text against another one.
Read the two ratios in the right order
The plain type-token ratio is types divided by tokens: how many different words per word written. It falls as a text grows, so use it only within one piece. The figure beside it averages the same ratio over every hundred-word stretch, which cancels the length effect out, and that is the one to quote when you are comparing a 300-word paragraph against a 3,000-word chapter.
Look at the bands and the once-only list
The five bands show how the vocabulary is distributed rather than summarised: a text with half its words used once reads differently from one with the same ratio and a flatter spread. Scan the once-only list for words you did not realise you had used, and the workhorse list beside it for the content words the piece leans on hardest.
Technical specifications
| What is counted | Distinct lowercase word forms; hyphens and apostrophes do not split a word, and no stemming is applied, so cats and cat are two entries |
|---|---|
| Type-token ratio | Distinct words ÷ running words, shown as a percentage — comparable only between texts of the same length |
| Length-corrected ratio | The same ratio averaged over every 100-word window in the text, the moving-average method published by Covington and McFall in 2010 |
| Root ratio | Guiraud's index: distinct words ÷ the square root of running words, an older correction that grows slowly with length rather than falling |
| Reuse bands | Five: used once, twice, 3–5 times, 6–10 times and more than 10, each with its share of the vocabulary |
| Once-only list | The first 60 words that appear exactly once, with the full count shown beside them |
| Function-word switch | 127 articles, pronouns, prepositions and auxiliaries can be set aside, which changes every ratio on the page |
| Price | Free, no signup, and no cap on the length of the text you measure |
Frequently asked questions
What counts as a unique word here?
A distinct spelling, after capitals are folded together. The and the are one word; cat and cats are two, and so are run and running, because collapsing them would mean guessing at word families and guessing wrongly in exactly the cases that matter. Apostrophes and internal hyphens do not split a word, so don't and state-of-the-art are one entry each, which keeps this total consistent with the unique-words figure the word counter reports for the same text.
Is there a target type-token ratio to aim for?
The question has no answer without a length attached, which is the single most important thing to know about the measure. A 100-word paragraph will typically score somewhere around 70 per cent because almost every word in it is still new; the same author writing 5,000 words will score nearer 30, having said the and of a few hundred times. Nothing about the writing changed. That is why the length-corrected figure is on this page at all, and why any target ratio quoted without a word count beside it is meaningless.
What are hapax legomena?
Words that appear exactly once in a text — the term is Greek for “said once” and comes from classical scholarship, where a word occurring a single time in the surviving corpus is genuinely hard to translate. In ordinary writing they are unremarkable and abundant: in most texts of a few thousand words, something close to half the distinct vocabulary appears only once, and the proportion is stubbornly stable as the text grows. A low share is more interesting than a high one, because it means the piece keeps returning to the same words.
Does a higher ratio mean better writing?
No. A high ratio can mean a rich vocabulary or it can mean an author reaching for a thesaurus and calling the same thing four different names, which is worse than repetition because the reader has to work out that all four are the same thing. Technical and legal writing deliberately repeats its key terms and scores low as a result; that is precision, not poverty. Read the ratio as a description of the text's texture, and then decide whether that texture suits what you are writing.
Why does the ratio drop every time I add a paragraph?
Because the grammar of English is finite and the subject matter is not. Every sentence you add needs the, a, of, is and to again, so the denominator keeps growing at full speed while the numerator slows down. This is not a flaw in your writing or in the measure — it is arithmetic, and it is exactly why the sliding-window figure exists. Watch that one instead: it stays roughly flat as a piece grows, and a real drop in it means the writing genuinely has become more repetitive.
How many different words does a person actually use?
Far fewer than they know. Estimates from large vocabulary studies put an average twenty-year-old American's known vocabulary at around 42,000 lemmas, and yet everyday speech and writing draw on a small fraction of that: a few thousand distinct words will carry an entire novel. Shakespeare's complete works run to about 884,000 words and use roughly 31,500 different ones, of which some 14,000 appear only once — which is a useful corrective to the idea that a big number in the box above is the goal.
Should I set the function words aside?
Do it when you are comparing subject matter, and leave them in when you are measuring style. With the, of and their 125 relatives included you are measuring the whole texture of the prose, function words and all, which is what stylometry and authorship studies actually work from. With them set aside you are measuring the range of content vocabulary — closer to what a teacher means by vocabulary range in a piece of student writing. The two give noticeably different ratios for the same text, so pick one and stay with it across everything you compare.
What the type-token ratio really measures
The vocabulary of a text has two counts, and linguistics gives them different names. Tokens are running words — every word as it occurs, so a page of 500 words has 500 tokens. Types are the distinct forms among them; if the appears thirty times it contributes thirty tokens and one type. Divide types by tokens and you have the type-token ratio, a measure that has been in use since the 1930s in language acquisition research and is still reported in clinical linguistics, second-language assessment and authorship studies. It is intuitive and it is easy to compute, and it has one flaw that swallows almost every naive use of it.
The flaw is that it is not a property of the writer at all: it falls automatically as the text gets longer. Function words are a closed set — English has a few hundred of them and every new sentence reuses them — so the token count keeps climbing at full speed while the type count slows to a crawl. Compare a 200-word cover letter with a 2,000-word essay by the same person and the letter wins every time, which tells you nothing about either. Statisticians have been patching this since the 1950s: Guiraud divided by the square root of the tokens, Herdan used logarithms, and in 2010 Michael Covington and Joe McFall proposed the fix used above — slide a fixed window along the text, compute the ratio inside every position of that window, and average the results. Because every window is the same size, length cancels out, and two texts of any lengths become comparable. It is the figure to quote if you are quoting one at all.
Once the arithmetic is honest, the interesting question is what variety is for. Repeated key terms are a feature of good technical writing, not a fault: a contract that calls the same thing the Property throughout is clearer than one that alternates with the premises and the site. Variety earns its place in narrative and argument, where repetition dulls attention. So read the bands rather than the single number — a piece with half its vocabulary used once and a small hard-working core is doing something quite different from one with the same ratio spread evenly. To see which words make up that core, ranked and exportable, the word frequency counter is the next page along; to check whether the words themselves are long as well as varied, the average word length tool measures the other half of vocabulary difficulty; and the running total this page divides by is the same one the word counter reports.
Student work and manuscripts are measured in the tab
Every ratio, band and word list on this page is computed by code already running in your browser, which matters when the text is somebody else's: a pupil's assignment, a clinical language sample or a manuscript under review is analysed without being uploaded, stored or shown to anyone. Nothing is retained after the tab closes, so there is no record to delete and nothing to consent to on behalf of the person who wrote it.