prosemeter

← Research

Measured against human readers at last, and the composite does not rank them

The first external validation. prosemeter's grade formulas track human readability judgments about as well as the published baselines do, the best of them is not the one the engine actually uses, and the composite carries no ordering above its bottom fifth.

run clear · 2026-08-29

Validation against the CLEAR corpus — 2026-08-29

Every number in eval/ before this one was generated by this repository and judged by a model. This run scores prosemeter against 4,724 texts rated by roughly 1,800 human teachers.

Three results, in descending order of how much they should change what gets built:

  1. The composite does not rank readability above its bottom fifth. Pooled over deciles 3–10, r = −0.008 on 3,780 documents.
  2. grade-band combines its five formulas the wrong way. The median it uses scores −0.528. prosemeter’s own SMOG, alone, scores −0.575.
  3. The median grade is not exported at all, and it is the only continuous readability signal the engine produces.

The corpus

The CommonLit Ease of Readability (CLEAR) corpus: 4,724 excerpts, each with a BT_easiness score. About 1,800 teachers each judged 100 pairs of excerpts for which was easier to read, and those pairwise judgments were fitted with a Bradley-Terry model to give every text a continuous easiness coefficient. Higher means easier.

That is the design run 7 amended itself toward — pairwise, many raters, aggregated — built at a scale this repository was never going to reach by hand.

It is not vendored here. The data is CC BY-NC-SA 4.0, and it is a 3.3 MB xlsx. Fetch it from github.com/scrosseye/CLEAR-Corpus, then see “Reproducing” below.

The method was validated before it was trusted

The corpus ships its own pre-computed formula columns, so the first thing to check is whether this correlation code reproduces the published result.

It does. CLEAR’s own Flesch-Kincaid column scores −0.517 here, matching the published −0.517. That check touches no prosemeter code, so it validates the export and the correlation math against an external anchor. clear-correlate.mjs prints it first and warns if it ever drifts.

1. The formulas are competitive, and the engine combines them badly

predictor Pearson r Spearman R²
CAREC (corpus’s best) −0.577 −0.578 0.333
prosemeter SMOG −0.575 −0.580 0.331
SMOG (corpus’s own) −0.563 −0.572 0.317
Dale-Chall −0.557 −0.597 0.311
prosemeter Flesch Reading Ease 0.554 0.569 0.307
prosemeter Gunning Fog −0.536 −0.563 0.287
prosemeter median-of-five −0.528 −0.561 0.279
prosemeter Flesch-Kincaid −0.528 −0.561 0.279
prosemeter ARI −0.497 −0.534 0.247
prosemeter Coleman-Liau −0.479 −0.492 0.229

The median of the five is worse than three of the five. SMOG (−0.575), Gunning Fog (−0.536) and Flesch-Kincaid (−0.528, a tie) all match or beat the median that grade-band actually uses. Coleman-Liau and ARI are the weak ones, and pooling by median lets them pull the result down.

prosemeter’s own SMOG is the strongest grade predictor measured here — it edges the corpus’s own SMOG implementation and effectively ties CAREC, the best formula the corpus ships, on Spearman (−0.580 against −0.578).

So the median is not a free robustness win. It costs about 0.05 of correlation against the best component, and the component that wins is already being computed on every call.

2. The band answers a different question, and the composite inherits nothing from it

grade-band maps the median through a bidirectional band: 1 inside [lo, hi], falling off on both sides. Correlating that against a monotone ease score gives 0.070 — but that near-zero is close to a tautology of the design, not a measurement of failure. The two directions cancel.

Split by side of the band, on plain’s 8–12:

group n mean BT_easiness band score vs BT_easiness
below band 1,267 −0.19 (easiest) −0.246
in band 1,909 −0.94 +0.013
above band 1,548 −1.61 (hardest) +0.266

The group means are perfectly ordered, and within each side the band score moves in the direction it was designed to move. The band does its stated job. An earlier draft of this report called that 0.070 “the signal destroyed”; that was wrong, and the one-sided split is what shows it.

Two findings survive the correction, and neither depends on the 0.070:

3. The composite has no ordering above the bottom fifth

Deciles by composite, against mean human easiness, with 95% confidence intervals:

decile mean composite mean BT_easiness
1 68.5 −1.798 ± 0.087
2 73.8 −1.183 ± 0.106
3 78.0 −0.704 ± 0.095
4 81.3 −0.821 ± 0.101
5 83.8 −0.880 ± 0.090
6 85.6 −0.928 ± 0.088
7 87.2 −0.759 ± 0.081
8 88.6 −0.912 ± 0.085
9 90.2 −0.836 ± 0.079
10 92.8 −0.759 ± 0.075

Pooled across deciles 3–10: r = −0.008 on 3,780 documents. No ordering at all.

Two details worth not smoothing over. The first two deciles climb steeply, so the composite does detect genuinely hard text — that is the floor, and it works. And deciles 3 through 6 are significantly anti-ordered: −0.704 and −0.928 have disjoint confidence intervals, so human ease falls as the composite rises across that stretch. Calling deciles 3–10 “flat” would understate the problem.

This is what README.md already claims — “a floor and a loop terminator, not a quality oracle” — now measured rather than asserted, with the floor located at roughly the bottom fifth.

4. Which dimensions carry readability

dimension r R²
sentence-simplicity 0.486 0.236
clarity 0.152 0.023
grade-band 0.070 0.005
lexical-diversity 0.014 0.000

sentence-simplicity is the dimension that tracks human judgment, and matches ARI as a standalone predictor. That independently confirms what “What does not need re-testing” already says from a different direction: it was the metric with a stable variance ratio across task sets, and it is the metric with the strongest tie to human readers.

clarity at 0.152 is consistent with it being a document check rather than a readability measure. lexical-diversity at 0.014 has no relationship with human readability on this corpus.

5. A calibration gap, in both inputs rather than one

The corpus computes four of the same formulas independently, which makes this a differential test against a published reference.

formula r mean difference max abs difference
Flesch-Kincaid 0.974 +0.56 grades 21.2
Flesch Reading Ease 0.984 −2.74 55.6
ARI 0.972 +0.37 27.1
SMOG 0.941 +0.39 9.0

High agreement with a systematic offset: prosemeter reads text as about half a grade harder than the reference. 5.7% of excerpts differ by more than 2 grade levels; 0.74% by more than 5.

Flesch-Kincaid and Flesch Reading Ease are both linear in the same two quantities, so each row’s pair of scores solves exactly for the words-per-sentence and syllables-per-word each implementation saw. Decomposing the mean offset:

source difference contribution to the offset
words per sentence +0.73 +0.285 grades
syllables per word +0.024 +0.279 grades

The offset is an even split between sentence segmentation and syllable counting, not the segmentation story an earlier draft of this report told from three example texts. Both inputs need looking at.

The tail is a different matter and the segmentation reading holds there: among the 270 excerpts differing by more than 2 grades, the disagreement is overwhelmingly driven by sentence length, and the worst cases are all ambiguous-boundary prose — semicolon-chained Robinson Crusoe, a Wilson peace-negotiation memoir, a Reformation article with a parenthetical Latin gloss.

This matters more than half a grade sounds: the profile bands are tuned in grade units, so a systematic offset shifts what falls inside every band in every profile.

What this changes

Limitations

Reproducing

The corpus is not vendored (CC BY-NC-SA 4.0). Two steps:

git clone --depth 1 https://github.com/scrosseye/CLEAR-Corpus.git
python3 -m venv venv && ./venv/bin/pip install openpyxl
./venv/bin/python eval/clear-export.py CLEAR-Corpus/CLEAR_corpus_final.xlsx clear.jsonl

pnpm build
node eval/clear-correlate.mjs clear.jsonl

clear-export.py pulls the excerpt text, BT_easiness, and the corpus’s own formula columns into JSONL. clear-correlate.mjs scores every excerpt and prints every table in this report. Scoring 4,724 excerpts takes about 47 seconds.

It also prints any excerpt it had to drop, with the word count and the reason. An earlier version captured the median grade with a pattern that could not match a minus sign, silently dropped three 149–168 word excerpts whose median grade was negative, and reported them as falling below the 30-word floor. Nothing is dropped now, and a drop would be visible rather than mislabelled.

Source: Crossley, S. A., Heintz, A., Choi, J., Batchelor, J., Karimi, M., & Malatinszky, A. (2023). A large-scaled corpus for assessing text readability. Behavior Research Methods.

The data behind this