Readability Scores: What Flesch Actually Measures and Where It Fails
Reading ease formulas are useful, cheap, and routinely misread. Understanding the arithmetic tells you exactly how far to trust the number.
Paste a paragraph into any content tool and it will hand you a readability score, usually a Flesch Reading Ease number and a US grade level. The scores are genuinely useful as a fast signal. They are also among the most over-interpreted numbers in content work, because almost nobody looks at what the formula is counting. It is counting two things, and neither of them is comprehension.
The Arithmetic Is Two Ratios
Flesch Reading Ease is 206.835 minus 1.015 times the average sentence length minus 84.6 times the average number of syllables per word. Flesch–Kincaid Grade Level rearranges the same two inputs: 0.39 times average sentence length plus 11.8 times average syllables per word minus 15.59. That is the entire model. Sentence length and syllables per word. Nothing else enters the calculation.
Notice the size of that syllable coefficient. At 84.6 it dominates the formula, which means word length moves the score far more than sentence length does. Replacing utilise with use across a document will shift the number more than splitting several long sentences. This is worth knowing before you spend an afternrooon restructuring paragraphs to chase a target.
What the Score Cannot See
Because the model reads only length, it is blind to everything else that makes prose hard. It cannot detect passive voice, tangled clause order, undefined jargon, missing logical connectives, or the absence of concrete examples. A sentence can be short, monosyllabic, and completely opaque. Worse, the formula can be gamed in ways that make writing measurably worse: chopping a well-built sentence into three fragments raises the score while destroying the argument.
A high reading ease score proves your words are short. It does not prove your meaning is clear.
The Syllable Count Is an Estimate
Every browser-based readability tool, including ours, approximates syllables by segmenting vowel clusters — counting groups of consecutive vowels as one syllable each. It is fast and roughly right, and it is wrong in predictable ways. Words such as queue and beautiful get over-counted because their vowel runs do not map to spoken syllables. Words such as rhythm and fire get under-counted for the opposite reason. Silent final e, diphthongs, and compound words all introduce error.
The practical consequence is that the absolute number carries less precision than its decimal place implies. Treat a grade level of 9.4 as approximately ninth grade, not as meaningfully different from 9.1. What the estimate does reliably is track direction: if you rewrite a section and the score moves four points, the change is real even if neither endpoint is exact.
It Only Works in English
This is the limitation that causes the most damage in practice. Both formulas were calibrated on English prose, and both the constants and the syllable model assume English word structure. Running German through them punishes ordinary compound nouns. Running Chinese or Japanese produces a number with no interpretable meaning at all, since the syllable estimator is looking for Latin vowels. If your content is not English, these scores are noise, and language-specific alternatives such as LIX for Scandinavian languages exist for a reason.
Reading the Bands Honestly
Roughly speaking, 90 and above reads as very easy, 60 to 70 as plain English for a general audience, 50 to 60 as fairly demanding, and below 30 as heavy going. The mistake is treating 60 to 70 as a universal target. Technical documentation, academic writing, and long-form criticism legitimately sit in the 40s, because the vocabulary is the point and the reader has opted in. A low score is a prompt to check whether the difficulty is doing work, not an error to eliminate.
Audience matters more than any band. Copy aimed at a broad consumer audience, health information, and government forms all benefit from being pushed toward 60 or higher, because the cost of a reader bouncing is high and the subject matter is often stressful. A developer changelog does not need the same treatment.
Where Typography Enters
Readability formulas measure the text. They cannot measure the setting, and the setting frequently matters more. The same paragraph at 130 characters per line, with 1.2 line-height and low colour contrast, will be abandoned by readers regardless of its Flesch score. Measure, leading, contrast, and type size are the variables that determine whether the copy is read at all.
That is the case for using the tools here in combination rather than in isolation. Run the copy through the Text Readability Checker for the linguistic signal, then use the Reading Length Optimizer to bring the measure into the 45 to 75 character range, the Text Line Spacing Calculator to set leading appropriate to that measure, and the Font Size Contrast Detector to confirm the hierarchy actually separates. The Text Repeat Checker catches the repetition that formulas ignore entirely.
How to Use the Number Well
- Use it as a relative signal across drafts of the same document, not as an absolute grade.
- Investigate outliers rather than averages — one 60-word sentence matters more than the mean.
- Never split a sentence purely to move the score.
- Ignore it entirely for non-English copy.
- Read the flagged passage aloud. If you run out of breath or lose the thread, that is better evidence than any formula.
Used this way the score earns its place: a cheap, instant smoke alarm that tells you where to look. It is not a judge of quality, and any workflow that treats a target number as a gate will eventually produce writing that scores well and communicates badly.