Metrics, Methods & Change Log

Everything this tool measures, how every number is earned, and a dated record of every change ever made to the scoring rules. Nothing about the method is hidden.

What this is The nine scored measures What a “5” means Two grading methods Rants, filler & media filler Invective & wit Judged by its own form Source verification Liberty & consequences The idea graph Fairness rules Second opinion Storage & determinism Known limits Change log The bias test →

What this is — and is not

This is a Value Detector, not an AI detector. AI detectors guess who typed the words and tell you nothing about whether a piece is worth reading. This tool ignores authorship entirely and measures the thinking: the evidence, the originality, the honesty, and what a reader gains. Every grade must be backed by verbatim quotes and specifics from the piece — a score without proof is treated as an error by the engine's own final consistency check.

The nine scored measures

Seven measure quality. Two measure undesirable qualities and are inverted, because low is good. All nine average into the overall number.

MeasureThe question it answers
Worth your time?Is this worth reading at all — or filler you could skip without missing anything? (The slop test. Not “was AI involved.”)
Does it show its work?Are sources and motives shown? Does it expose something hidden, or appear to hide something? Would an ordinary reader notice an omission and object?
Does it push ideas forward?Does it take existing ideas somewhere new? Combining familiar ideas so they produce new understanding counts.
New ideas of its ownIdeas not found in the usual writing on the subject. A new arrangement of old ideas counts when it reveals what the parts alone do not.
Quality of the evidenceAre claims backed by checkable facts — statutes, records, published research, displayed documents?
Does the argument hold up?Do the conclusions follow from the evidence and premises given? Is it internally consistent and accurate?
What you get out of itWhat you will know or be able to do afterward that you could not before.
Free of ranting invertedDoes the piece prove something, or only discharge a grievance? Force and moral certainty are not ranting.
Substance, not filler invertedIs the word count doing work, or is it padding you could delete without loss?

What a “5” means — the anchor ladder

5 is average for what is actually published on that specific subject — not an imagined ideal article. If little published writing on a subject reaches a piece's depth, the piece scores high. A piece is never held down for lacking things (original data collection, interviews, public-records requests) that almost nothing in its field has.

Every measure has a printed ladder, and the grader must award the number when the piece meets the description. Example — quality of the evidence:

ScoreWhat it looks like
3Assertions; few checkable facts.
5Some checkable sourcing — typical published commentary.
7Consistently sourced to named published research or official records.
9Independent published sources plus primary citations (statutes by section, court cases, official records) plus reference apparatus — footnotes, glossary, citations — that survives verification.
10All of that plus original documents or data.

Score what is there — never deduct for what is missing. Presence earns points; absence simply does not earn them. Things a piece could have done but did not — an unaddressed counterargument, an interview not conducted, a case study not included — are recorded separately as Opportunities, are shown to the reader as neutral suggestions, and touch no score. The only things that may lower a score are defects in what the piece actually did: factual errors, fabricated or misattributed citations, internal contradictions, claims its own evidence does not support, or genuine filler.

Two grading methods, every time

1. The checklist method. Fixed, printed definitions say what each number means. If the piece meets a level's definition, it gets that number — the grader may not withhold it because a hypothetically better piece could exist.

2. Against the best of its kind. The tool names a real best-in-class peer work or author for the piece's type of writing and grades it next to that benchmark.

The number shown is the average of the two — unless they disagree by 2 or more points, in which case you see the range, because the disagreement is itself information.

Final consistency check: after writing each reason and evidence list, the grader re-reads the anchors and compares them against the evidence it just listed. If its own evidence demonstrates a higher level, the score must rise. A number lower than the evidence supports is a scoring error.

Rants, filler, and media filler

What makes a rant

A rant is defined by what it lacks and why it exists — never by its intensity. Force, anger at government, moral certainty, and blunt language are not rant indicators; the most important writing in the American tradition has all of them. A piece scores as a rant when most of these hold:

Anger with receipts is a polemic and scores well. Anger without them does not score at all. When in doubt, it is not a rant.

Filler and the media-filler check

Filler is word count that adds nothing — restatement, generic observations anyone could write without knowing the subject, passages you could delete without losing anything.

Media filler is the content-farm pattern common to local radio-station and aggregator sites, and it gets its own report. Indicators: the whole article reduces to one fact stretched over hundreds of words; no original reporting (rewritten from another outlet, a press release, an agency statement, or a social post); padding devices (restating the headline, “here's what you need to know,” rhetorical questions, a bolted-on local angle); engagement bait and listicle bulk; and low information density.

The report states plainly, in one sentence, everything the article actually tells you, alongside a count of distinct verifiable facts, the approximate word count, whether any original reporting occurred, and what it appears rewritten from.

Invective and wit

Earned vs. empty invective

Name-calling is a legitimate rhetorical tool with a long pedigree — Mencken, Ivins, and Buckley all did it well. What matters is whether the epithet earns itself.

EARNED — the target is named or unmistakably identified, and the piece supplies the definition or evidence that makes the label fit. Example: “RINO” used and defined by the votes and positions that qualify it; or calling officials oath-breakers after quoting the statute they violated and their own admission. The insult summarizes an argument the piece actually made. Never penalized; may be credited as a strength.

EMPTY — the label lands on an unnamed or vague target, is never defined, and rests on no evidence. The epithet substitutes for the argument instead of summarizing one. This is not polemic; it is the absence of an argument dressed as one.

Undefined coinages are article-killers. An invented nickname for a group or person that the piece never defines is among the worst defects possible, because it makes the argument private — decodable only by readers already inside the author's grievance, and unfalsifiable to everyone else. When a piece leans on one, ranting scores 7 or higher and the checkable-content measures must reflect that the central claims cannot be checked.

The test in one line: could a reader say who is meant and why the label fits, using only the piece? If yes it is earned; if no it is empty.

Humor and sarcasm — credited only when earned

Wit is not automatically a virtue, and its absence is never a fault. Earned wit illuminates: it makes an argument land harder than plain statement would, exposes an absurdity the reader can now see, compresses a real point into a memorable image, or reframes something familiar. That is genuine original thought, and it raises “new ideas of its own” and “worth your time.” Unearned wit — decoration, in-group snark, a sneer standing in for a point — earns nothing. No points are awarded merely for attempting humor, and none are deducted for being serious throughout.

Judged by its own form

Every piece answers the same nine questions; no form earns a free pass. But the form informs what a strong answer looks like, and the results always state plainly how the form affected the grades.

Source verification

Citations and footnotes are not merely counted — each is checked: does the cited work exist, is it attributed correctly, does it support the claim it is attached to? Glossary definitions are checked for accuracy. Verified apparatus raises scores; fabricated, misattributed, or misused citations lower them and are flagged with the specific problem. Results show the count checked and any flags.

Displayed documents count as shown work. The analysis reads text only, so meeting minutes, letters, and records a piece displays as images cannot be viewed by the grader — they are credited as primary evidence shown to the reader, never treated as missing. Image alt text and figure captions are preserved during extraction so the grader can see they exist.

Disclosed limit: verification is checked against the model's own knowledge, not a live fetch of each source. Only claims with positive reason to be wrong are flagged; an unfamiliar citation is never flagged merely for being unfamiliar.

Liberty & consequences

When a piece involves government action, regulation, or rights, it receives a dedicated analysis in the negative-liberty tradition — liberty as freedom from government interference, powers enumerated and limited, rights preceding government rather than granted by it. It names:

A government action does not become acceptable because it was reversed. The violation, the harm while it stood, and the fact that reversal required a citizen to force it are all part of the consequences.

The idea graph

Separately from scoring, the tool maps the piece's ideas: 10–16 concepts as circles, 12–20 directed relationships as arrows, and the meaningful omissions.

ElementMeaning
Circle colorHow the piece treats the idea — promising, dangerous, contested, neutral. Not whether the idea is true.
Circle sizeHow much of the argument the idea carries.
Purple ringUsual ground — most writing on the topic, including AI-written pieces, would also center this idea.
Dashed green ringA fresh idea — this piece handles it, or connects it, in a genuinely new way.
Dashed gray circleAn omission — an idea readers would expect that the piece leaves out, where the absence itself says something.

Center of gravity: the highest-weighted concepts must be the piece's central claim and its stakes — what happened, who did it, what law or principle it violated, and what came of it. When a piece documents officials reversing course because they were called out, that accountability arc is a core concept, never omitted. Practical asides (lawful alternatives, procedures) carry low weight.

Fairness rules

The optional second opinion

A “Verify with a second AI” button sends the piece to an independent model (Google Gemini) that re-scores every measure and independently lists what it considers the usual ground for the topic — blind to the first analysis. Agreements and disagreements are shown side by side; disagreements over 2 points are flagged rather than averaged away. Concepts both models call usual ground are marked ✓; where they disagree, marked ?.

Honest caveat: the two checklist/comparative methods are the same model reasoning two ways — not fully independent. The second-opinion button is what closes that gap.

Storage, determinism, and cost

Scoring runs at temperature zero: the same article produces the same grades on every scan. Every scan is stored permanently in a database keyed to the article, so a piece is analyzed once and served instantly forever after — repeat visits cost nothing and never re-run the AI. An Update link on any result forces a fresh analysis and replaces the stored one. Stored results are labeled with the engine version that produced them, and share links carrying tracking parameters resolve to the same stored article.

Known limits

The declared perspective built into these rules — and a published test of whether it skews the scores — is documented separately on The Bias Test.

Change log — every scoring change, dated

Each engine version is recorded with what changed and why. Several rules exist because bias was identified in live results and corrected; those are stated plainly rather than quietly patched. Stored analyses are labeled with the version that produced them.

VersionDateChange
v222026-07-25Media-filler detector added.Why: local-radio and aggregator sites stretch one fact across hundreds of words. The report now names everything an article actually tells you, counts verifiable facts against word count, and identifies what it was rewritten from.
v212026-07-25Humor and sarcasm credited only when earned.Why: wit that illuminates is original thought and should score; jokes for their own sake should not. No points for attempting humor, none deducted for seriousness.
v202026-07-25Earned vs. empty invective; undefined coinages treated as article-killers.Why: name-calling is legitimate when the target is named and the label defined (“RINO” defined by votes). An undefined coinage makes the argument private and unfalsifiable.
v192026-07-25Ranting and filler folded into the overall score as inverted measures.Why: they were diagnostics only. Rants and filler are undesirable work and must affect the grade.
v182026-07-25Rant and filler detection added, with the three dominant qualities of the writing.Why: to separate grievance-without-evidence from forceful, well-supported argument — intensity is explicitly not a rant indicator.
v172026-07-25Score what is there; never deduct for what is missing. Opportunities recorded separately.Why: pieces with heavy citation were being marked down for interviews not conducted and counterarguments not taken up. Absence-based complaints are now barred from scoring reasons.
v162026-07-25“Current relevance” metric removed; JSON repair passes and a longer analysis window added.Why: the model's estimates of what is currently in the news were unreliable and caused more confusion than value. Long analyses containing quoted material were also failing to parse.
v152026-07-25Liberty & consequences analysis added.Why: pieces documenting government action needed a dedicated negative-liberty analysis naming the right at stake, the lawfulness, the consequences, and who forced accountability.
v142026-07-25Idea graph re-centered on stakes and accountability.Why: the graph was weighting practical asides (lawful alternatives) above the actual story — a government body violating the law and reversing only when called out.
v132026-07-25Displayed documents credited as shown work; alt text and captions preserved in extraction.Why: a text-only grader was treating displayed meeting minutes and records as missing evidence.
v122026-07-25Concrete score anchors printed for every measure; grader required to award the level the evidence meets.Why: scores were being withheld from work that plainly met the description, on the reasoning that something better could exist.
v112026-07-25Moral and constitutional argument is not a defect; classification tightened.Why: identified bias — evidence-heavy pieces taking a moral position were being re-classified downward and discounted for being normative rather than empirical.
v102026-07-25Deterministic scoring (temperature zero) and canonical URLs.Why: the same article produced different grades on different runs, and share links with tracking parameters were creating duplicate scans.
v92026-07-25Citation and glossary verification added.Why: apparatus should not merely be counted — citations are checked for existence, attribution, and support; glossary terms for accuracy.
v82026-07-25Calibration to the real field; reference apparatus credited.Why: “5 = average” was being measured against an imagined ideal article rather than what is actually published on the subject.
v72026-07-25Form effect disclosed in every result.Why: readers should see how a piece's form shaped its grades, including what was deliberately not held against it.
v62026-07-25Score by the piece's own form.Why: a polemic was being judged by the duties of reportage. Same questions for every piece, but the form sets what a strong answer looks like.
v52026-07-25Framework supremacy rule.Why: identified bias — a rights-based argument was marked down for not answering a public-safety cost-benefit objection it rejects on principle. Rigor is now judged within the piece's own framework.
v42026-07-25Fresh-idea marking and whole-article synthesis scoring.Why: familiar ideas combined in a way that produces new understanding is original work and must be credited.
v32026-07-25Plain-language rules; academic jargon banned from reader-facing text.Why: results were unreadable to ordinary readers and often failed to say whether they described the article, the author, or the AI.
v22026-07-24Full rebuild: seven measures, dual grading methods, best-in-class comparison, evidence lists required, internal weaknesses analysis kept private.Why: the original four metrics produced vague verdicts without supporting specifics.
v12026-07-24Initial release: idea graph with four metrics — public disclosure, transparency, extending concepts, original thought.

Tool and infrastructure changes

ItemDateChange
Explainer2026-07-25This page published — all metrics, methods, and the complete dated change log.
Storage2026-07-25Permanent scan library added. Every scan stored and served instantly forever; “Update” forces a fresh run. File cache retained as fallback.
Branding2026-07-25Renamed Article Values Graph™; purpose and “why not an AI detector” explanation added to the top of the tool.
Layout2026-07-25Idea table moved directly beneath the graph, left of the score column; mobile layout added; graph spread control, node pinning, and always-on relationship labels.
Extension v1.12026-07-25Substack dashboard score panel: dashboard pagination support, self-update check against the published version.
Extension v1.02026-07-25Chrome extension released — score badges for recent posts on the Substack publish dashboard.
Extraction2026-07-25Substack app-style links (substack.com/@user/p-123) resolved through the post API; those pages serve no article text to servers.
Access2026-07-25Automatic link scanning limited to jeffapierson.substack.com essays; any URL accepted when entered directly in the form.
Deployment2026-07-24Tool deployed. Server-side API key, private config, cache directory not web-readable.

← Back to the Article Values Graph™