Tools › Text
Frequency Analysis
How often each word, word pair, character, letter or letter pair appears, with counts, percentages and bars, optionally skipping common words, with the index of coincidence, chi-squared and entropy for breaking ciphers.
About this tool What it's for, how to use it and an example
What it's for
Count how often each word, word pair, character, letter or letter pair appears in some text, with percentages and bars.
- When you’re checking a piece of writing for overused words.
- When you’re solving a substitution cipher or a puzzle using letter frequencies.
- When you’re finding the most common terms in feedback, survey answers or a list of tags.
How to use it
Paste text and choose what to Count and how many results to Show. Set the shortest word to count, and tick Ignore case, Skip common words (the, and, of…) or Count spaces as needed. The results update as you type. Chart these sends the results to Chart Maker, and Copy CSV copies them for a spreadsheet.
Above the table are three figures for telling ciphers apart: the index of coincidence (how likely two letters picked at random are the same: about 0.067 for English or English rearranged, 0.038 for random letters), the chi-squared distance of the letters from English, and the entropy in bits per character.
Example
Paste this with Words selected:
the cat sat on the mat and the cat slept
The tool reports 10 words, 7 different ones. the comes first with 3 (30%), then cat with 2 (20%). Tick
Skip common words and the and and drop out.
Good to know
Hyphenated words are counted as separate words, while contractions such as don't stay whole. The common-word list is English only.
Guide: Solving an unknown cipher
Private: this tool runs in your browser. Nothing you type, paste or choose leaves this page.
Saved you a few minutes? Say thanks with a coffee.
Something wrong with this tool, or missing from it? Report a bug or suggest a feature.