Writing guide · 7 min read
Cohesion and Flow: How the Engine Reads Connection
Lexical Cohesion counts repeated words between neighboring sentences. Transition Signal Rate counts connectives. Here is what each one can and cannot see.
By Alyssa Glasco, Founder · Published
Two readings describe how your sentences hold together. Neither one reads meaning. Both count words. That limit is the most important thing to know about them, so it comes first rather than last.
Lexical Cohesion
For every pair of neighboring sentences, the engine strips out function words, reduces what remains to word stems, and measures how much the two sets overlap. Your score is the average of that overlap across the passage.
The numbers are small, and they are supposed to be. Published fiction sits between about 0.004 and 0.034. Technical writing runs highest, from 0.013 to 0.077, because specifications repeat their nouns on purpose. These floors and ceilings come from the 25th and 90th percentiles of measured published prose in each genre.
There is history in those numbers worth knowing. An earlier version of this benchmark set the floor at roughly six times the median, which flagged more than 90 percent of published writing. That was a permanent flag rather than a benchmark. If you read old advice telling you to aim for 0.3 overlap, ignore it. Nothing real reaches it.
What it cannot see
Word overlap is not logic. A passage can repeat its nouns faithfully and still argue in circles, and a beautifully reasoned paragraph can score low because the writer used a synonym every time. The reading carries a note saying exactly this, and the note is not boilerplate.
The stemmer also misses irregular past tenses. Ran is not matched to run, brought is not matched to bring. Narrative prose in past tense therefore scores slightly lower than its real cohesion, which is part of why the fiction band sits where it does.
Where it earns its place
Coherence Gaps is the more actionable number that comes with it: the count of neighboring sentence pairs that share nothing at all, where both sentences are at least five words long. A gap is not automatically a problem. A deliberate hard cut between scenes is a gap. But if you have eleven of them in a page of argument, some of them are places a reader loses the thread.
Transition Signal Rate
The engine scans for 93 connective phrases in five families: additive (in addition, moreover), contrastive (however, on the other hand), causal (because, as a result), sequential (first, finally), and illustrative (for example). The score is matches per thousand words. Longer phrases are matched first, so the in inside in addition is not counted twice.
This reading only appears for nonfiction, journalism, technical writing, and blogging. Fiction, poetry, and screenwriting never see it, because narrative moves by scene and image rather than by connective, and counting then in a story would measure nothing.
- Nonfiction: 5 to 20 per thousand. Argument relies on explicit connectives between ideas.
- Technical writing: 6 to 25, the widest band. Logical connectives are load-bearing in a specification.
- Journalism: 4 to 15. More than narrative, less than essay.
- Blogging: 3 to 15. Above that, a post reads as over-structured.
What it cannot see
It is string matching. It cannot tell a real logical move from the word then used to mean next in time. A high score does not prove your argument connects, and a low score in an essay is worth a look rather than a rewrite.
Using them together
The two readings answer different questions, and the interesting case is when they disagree. High connectives with low cohesion often means the prose is signposting moves it is not actually making: every paragraph opens with however while the subject changes underneath. Low connectives with high cohesion is usually fine, and is what good narrative looks like.
Open a piece in the editor and run Grade this passage, then read the Coherence Gaps count before either score. It points at specific sentence pairs, which is easier to act on than an average.