Measured against human readers at last, and the composite does not rank them
The first external validation. prosemeter's grade formulas track human readability judgments about as well as the published baselines do, the best of them is not the one the engine actually uses, and the composite carries no ordering above its bottom fifth.
Validation against the CLEAR corpus — 2026-08-29
Every number in eval/ before this one was generated by this repository and judged by a model.
This run scores prosemeter against 4,724 texts rated by roughly 1,800 human teachers.
Three results, in descending order of how much they should change what gets built:
- The composite does not rank readability above its bottom fifth. Pooled over deciles 3–10, r = −0.008 on 3,780 documents.
grade-bandcombines its five formulas the wrong way. The median it uses scores −0.528. prosemeter’s own SMOG, alone, scores −0.575.- The median grade is not exported at all, and it is the only continuous readability signal the engine produces.
The corpus
The CommonLit Ease of Readability (CLEAR) corpus: 4,724 excerpts, each with a BT_easiness score.
About 1,800 teachers each judged 100 pairs of excerpts for which was easier to read, and those
pairwise judgments were fitted with a Bradley-Terry model to give every text a continuous easiness
coefficient. Higher means easier.
That is the design run 7 amended itself toward — pairwise, many raters, aggregated — built at a scale this repository was never going to reach by hand.
It is not vendored here. The data is CC BY-NC-SA 4.0, and it is a 3.3 MB xlsx. Fetch it from
github.com/scrosseye/CLEAR-Corpus, then see “Reproducing” below.
The method was validated before it was trusted
The corpus ships its own pre-computed formula columns, so the first thing to check is whether this correlation code reproduces the published result.
It does. CLEAR’s own Flesch-Kincaid column scores −0.517 here, matching the published −0.517.
That check touches no prosemeter code, so it validates the export and the correlation math against
an external anchor. clear-correlate.mjs prints it first and warns if it ever drifts.
1. The formulas are competitive, and the engine combines them badly
| predictor | Pearson r | Spearman | R² |
|---|---|---|---|
| CAREC (corpus’s best) | −0.577 | −0.578 | 0.333 |
| prosemeter SMOG | −0.575 | −0.580 | 0.331 |
| SMOG (corpus’s own) | −0.563 | −0.572 | 0.317 |
| Dale-Chall | −0.557 | −0.597 | 0.311 |
| prosemeter Flesch Reading Ease | 0.554 | 0.569 | 0.307 |
| prosemeter Gunning Fog | −0.536 | −0.563 | 0.287 |
| prosemeter median-of-five | −0.528 | −0.561 | 0.279 |
| prosemeter Flesch-Kincaid | −0.528 | −0.561 | 0.279 |
| prosemeter ARI | −0.497 | −0.534 | 0.247 |
| prosemeter Coleman-Liau | −0.479 | −0.492 | 0.229 |
The median of the five is worse than three of the five. SMOG (−0.575), Gunning Fog (−0.536)
and Flesch-Kincaid (−0.528, a tie) all match or beat the median that grade-band actually uses.
Coleman-Liau and ARI are the weak ones, and pooling by median lets them pull the result down.
prosemeter’s own SMOG is the strongest grade predictor measured here — it edges the corpus’s own SMOG implementation and effectively ties CAREC, the best formula the corpus ships, on Spearman (−0.580 against −0.578).
So the median is not a free robustness win. It costs about 0.05 of correlation against the best component, and the component that wins is already being computed on every call.
2. The band answers a different question, and the composite inherits nothing from it
grade-band maps the median through a bidirectional band: 1 inside [lo, hi], falling off on
both sides. Correlating that against a monotone ease score gives 0.070 — but that near-zero is
close to a tautology of the design, not a measurement of failure. The two directions cancel.
Split by side of the band, on plain’s 8–12:
| group | n | mean BT_easiness |
band score vs BT_easiness |
|---|---|---|---|
| below band | 1,267 | −0.19 (easiest) | −0.246 |
| in band | 1,909 | −0.94 | +0.013 |
| above band | 1,548 | −1.61 (hardest) | +0.266 |
The group means are perfectly ordered, and within each side the band score moves in the direction it was designed to move. The band does its stated job. An earlier draft of this report called that 0.070 “the signal destroyed”; that was wrong, and the one-sided split is what shows it.
Two findings survive the correction, and neither depends on the 0.070:
- 40.7% of the corpus — 1,925 of 4,724 excerpts — scores exactly 1.0, spanning
BT_easiness−3.59 to 1.71 against a corpus range of −3.68 to 1.71. Inside the band the score is flat by construction (r = +0.013), so the composite inherits no readability ordering from this dimension for two fifths of documents. - The median grade is not exported. It is the engine’s only continuous readability signal, it
carries −0.528, and reaching it requires a regex over
grade-band’s detail string. That is what this report’s own harness does, and it should not have to.
3. The composite has no ordering above the bottom fifth
Deciles by composite, against mean human easiness, with 95% confidence intervals:
| decile | mean composite | mean BT_easiness |
|---|---|---|
| 1 | 68.5 | −1.798 ± 0.087 |
| 2 | 73.8 | −1.183 ± 0.106 |
| 3 | 78.0 | −0.704 ± 0.095 |
| 4 | 81.3 | −0.821 ± 0.101 |
| 5 | 83.8 | −0.880 ± 0.090 |
| 6 | 85.6 | −0.928 ± 0.088 |
| 7 | 87.2 | −0.759 ± 0.081 |
| 8 | 88.6 | −0.912 ± 0.085 |
| 9 | 90.2 | −0.836 ± 0.079 |
| 10 | 92.8 | −0.759 ± 0.075 |
Pooled across deciles 3–10: r = −0.008 on 3,780 documents. No ordering at all.
Two details worth not smoothing over. The first two deciles climb steeply, so the composite does detect genuinely hard text — that is the floor, and it works. And deciles 3 through 6 are significantly anti-ordered: −0.704 and −0.928 have disjoint confidence intervals, so human ease falls as the composite rises across that stretch. Calling deciles 3–10 “flat” would understate the problem.
This is what README.md already claims — “a floor and a loop terminator, not a quality oracle” —
now measured rather than asserted, with the floor located at roughly the bottom fifth.
4. Which dimensions carry readability
| dimension | r | R² |
|---|---|---|
sentence-simplicity |
0.486 | 0.236 |
clarity |
0.152 | 0.023 |
grade-band |
0.070 | 0.005 |
lexical-diversity |
0.014 | 0.000 |
sentence-simplicity is the dimension that tracks human judgment, and matches ARI as a
standalone predictor. That independently confirms what “What does not need re-testing” already
says from a different direction: it was the metric with a stable variance ratio across task sets,
and it is the metric with the strongest tie to human readers.
clarity at 0.152 is consistent with it being a document check rather than a readability measure.
lexical-diversity at 0.014 has no relationship with human readability on this corpus.
5. A calibration gap, in both inputs rather than one
The corpus computes four of the same formulas independently, which makes this a differential test against a published reference.
| formula | r | mean difference | max abs difference |
|---|---|---|---|
| Flesch-Kincaid | 0.974 | +0.56 grades | 21.2 |
| Flesch Reading Ease | 0.984 | −2.74 | 55.6 |
| ARI | 0.972 | +0.37 | 27.1 |
| SMOG | 0.941 | +0.39 | 9.0 |
High agreement with a systematic offset: prosemeter reads text as about half a grade harder than the reference. 5.7% of excerpts differ by more than 2 grade levels; 0.74% by more than 5.
Flesch-Kincaid and Flesch Reading Ease are both linear in the same two quantities, so each row’s pair of scores solves exactly for the words-per-sentence and syllables-per-word each implementation saw. Decomposing the mean offset:
| source | difference | contribution to the offset |
|---|---|---|
| words per sentence | +0.73 | +0.285 grades |
| syllables per word | +0.024 | +0.279 grades |
The offset is an even split between sentence segmentation and syllable counting, not the segmentation story an earlier draft of this report told from three example texts. Both inputs need looking at.
The tail is a different matter and the segmentation reading holds there: among the 270 excerpts differing by more than 2 grades, the disagreement is overwhelmingly driven by sentence length, and the worst cases are all ambiguous-boundary prose — semicolon-chained Robinson Crusoe, a Wilson peace-negotiation memoir, a Reformation article with a parenthetical Latin gloss.
This matters more than half a grade sounds: the profile bands are tuned in grade units, so a systematic offset shifts what falls inside every band in every profile.
What this changes
- Score
grade-bandfrom a pool of the strong formulas rather than the plain median. Done — 0.5.0 pools SMOG, Gunning Fog and Flesch-Kincaid, reaching −0.560.clear-sweep.mjsswept eight candidates; SMOG alone was strongest at −0.575 and failed the telegraphic-prose fixture, because its constant floors it near grade 3.1. Every profile’s band edges are unchanged: the pooled distribution shifts 0.05 grades corpus-wide. - Export the median grade (or whichever statistic replaces it). It is the only continuous readability signal the engine produces and no public API returns it.
- Stop treating the composite as ordered above the bottom fifth. The floor framing in the README is correct and should be the only framing.
- Investigate both the segmentation and the syllable counter before any further band tuning.
sentence-simplicityhas now earned its weight twice, by two unrelated methods.
Limitations
- The register is wrong for prosemeter’s actual use. CLEAR is literary and informational
excerpts of 129–205 words, average publication year 1937.89 (from the paper; the export drops the
year column). There are no READMEs, no API docs, no chat replies, no technical prose. These
results validate
grade-band’s inputs on general English; they say nothing about thereadme,api-docsorchatprofiles on the documents those profiles target. - The ground truth is teachers rating difficulty for students in grades 3–12. prosemeter’s profiles mostly target adult technical readers. “Easy for a seventh-grader” and “clear to a senior engineer” are related but not the same quantity, and this measures the first.
- The one-sided band correlations in §2 are range-restricted and must not be compared against the full-corpus figures. They establish direction, not effect size.
- One corpus, one construction of the ground truth. Bradley-Terry over pairwise judgments is a defensible aggregation and not the only one.
- Nothing here measures the loop. This is a static-scoring validation. Run 6 remains the only test of whether revision improves a document.
Reproducing
The corpus is not vendored (CC BY-NC-SA 4.0). Two steps:
git clone --depth 1 https://github.com/scrosseye/CLEAR-Corpus.git
python3 -m venv venv && ./venv/bin/pip install openpyxl
./venv/bin/python eval/clear-export.py CLEAR-Corpus/CLEAR_corpus_final.xlsx clear.jsonl
pnpm build
node eval/clear-correlate.mjs clear.jsonl
clear-export.py pulls the excerpt text, BT_easiness, and the corpus’s own formula columns into
JSONL. clear-correlate.mjs scores every excerpt and prints every table in this report. Scoring
4,724 excerpts takes about 47 seconds.
It also prints any excerpt it had to drop, with the word count and the reason. An earlier version captured the median grade with a pattern that could not match a minus sign, silently dropped three 149–168 word excerpts whose median grade was negative, and reported them as falling below the 30-word floor. Nothing is dropped now, and a drop would be visible rather than mislabelled.
Source: Crossley, S. A., Heintz, A., Choi, J., Batchelor, J., Karimi, M., & Malatinszky, A. (2023). A large-scaled corpus for assessing text readability. Behavior Research Methods.
The data behind this
- eval/results/run-clear.json — every dimension score for every answer
- eval/corpus/run-clear/ — the answers themselves, marked as experiment output
- the full report — this page renders it