Extractability Scorer A model quotes a sentence, not a page.
Score a page
- 7
- Your browser
- Nothing
It lifts one line out of your copy, strips everything around it, and drops it into a paragraph about something else. Whatever that line was leaning on is gone. Score any page for how many of its sentences survive the journey.
What it checks
Two classes, and the difference decides the count. Blocking faults mean the sentence cannot stand up alone at all. Weakening ones mean it survives, but arrives worth less.
Opens with an orphan reference
The subject lives in a previous sentence. Lifted out, the claim is about nothing.
Points at the page around it
Above, below and here do not exist once the sentence is quoted somewhere else.
Nothing a model can verify
Superlatives carry no checkable content, so there is nothing to corroborate off-site and nothing worth quoting. This one is graded: when the superlative *is* the claim — "we are world-class" — the sentence has nothing left once it is doubted. When it only decorates a claim that survives without it — "integrates seamlessly with Salesforce" — the sentence is weaker, not broken, and it is marked that way.
Hedged
An answer engine is picking a sentence to assert. Hedged claims get passed over for ones that commit.
Unquantified quantity
Many and most are not numbers. They cannot be checked, compared, or repeated with confidence.
Figure with no source
A number without attribution is the first thing a careful model declines to repeat.
Too long to quote whole
A sentence that has to be cut to be used gets cut where the model chooses, not where you would.
Questions
What is extractability?
Whether a claim still means something once it is lifted off your page. A model composing an answer does not quote your page — it quotes a sentence from it, stripped of the paragraph above and dropped into a reply about something else. Anything that sentence was leaning on is gone. Extractability is the property of surviving that.
Why does an orphan subject matter so much?
Because it is the difference between being cited and being paraphrased. "It reduces onboarding time by half" is useless quoted alone — a model either skips it, or attributes it to whatever noun happens to be nearest in its own draft, which is how brands end up credited with a competitor's claim. "Acme reduces onboarding time by half" cannot be misattributed.
Is this the same as readability?
No, and they pull in different directions. Readability rewards short sentences with pronouns doing the connective work, because a human reads top to bottom and carries context with them. Extraction has no context to carry. A page can score well on Flesch-Kincaid and be almost entirely unquotable.
Why flag words like world-class and seamless?
Because there is nothing behind them for a model to corroborate. An answer engine is deciding which source to stand behind, and it can check a founding date, a named customer, or a figure with a source. It cannot check an adjective. Superlatives are not just weak writing here — they are structurally uncitable.
Does the tool store my page?
No. The page is fetched once and passed straight to your browser, and every score you see is computed there from the text. Nothing is written down and nothing is queued for a human to look at.
A page that survives being quoted still has to be found among the other two thousand. The last question is which of them you would want a model to read first. Draft the file that says which pages matter