The Bias Test

I built a tool that grades writing. The obvious question is whether it just grades my politics. So I tested it against my own side, and I am publishing the results along with the place where my thumb is on the scale.

The test

I scanned six pieces across the political spectrum and across quality levels, including my own work, a conservative opinion column, and a progressive investigative outlet I disagree with on nearly everything. Same engine, same rules, deterministic scoring.

PiecePoliticsOverallWhat drove it
Popular Information — “ICE is lying about body cams. These documents prove it.”Progressive8.1Primary documents, eight citations verified with zero flags, contradictions shown rather than asserted.
JeffAPierson.com — “Who Pays for Growth?”Conservative / liberty8.1Fourteen verified citations, Idaho Code by section, glossary. (Seven quality measures; scanned before ranting and filler were added.)
JeffAPierson.com — “No Legal Basis for the Ban in Jerome County”Conservative / liberty7.7Statutes quoted in full, officials' own recorded admissions, documented reversal.
A local media aggregation articleNone3.4Six facts in 350 words, rewritten from a database entry, no original reporting.
Article written entirely by AI (control)None3.1Twelve of fourteen ideas were common ground, zero new, zero citations to check.
A conservative opinion columnConservative2.1Unnamed antagonists, undefined coinages, unverifiable claims, disconnected anecdotes.

What the numbers show

The interesting result is not that a progressive article scored well. It is where the spread is.

Within my own political camp, the spread was six points. A conservative opinion column scored 2.1. My own work scored 7.7 and 8.1. Same politics, opposite ends of the scale.

Across the political divide, the spread was zero. A progressive investigative outlet and my property-tax analysis landed on the same number, 8.1, while agreeing on almost nothing.

Political direction was not the variable. Whether a reader could check the claims was. That is the entire design of the tool, and this is the closest thing I have to evidence that it works.

One more detail worth noticing: the liberty-and-consequences analysis engaged the progressive piece without friction, naming the right at stake as “the right to life and the right to an independent record of government use of lethal force.” The framework is not partisan. Government overreach is government overreach regardless of who is committing it.

Where my thumb is on the scale

I am not going to pretend this tool has no perspective. It does, I put it there deliberately, and I would rather state it plainly than have someone find it and call it a gotcha.

The declared preference. The scoring rules name four frameworks explicitly — negative liberty, natural law, constitutional law, and biblical law — and grant them specific protection. A piece grounded in them is judged within that framework and is never marked down for declining to weigh rights against utility. If an argument holds that the function of a right is to remove certain harms from cost-benefit math, refusing that trade-off is the argument working, not a hole in it.

What that means in practice. A rights-based piece that refuses a public-safety cost-benefit objection is protected by name. A piece grounded in a utilitarian or collective-welfare framework that refuses a rights-based objection is not protected by name. That is an asymmetry. It is real, it is intentional, and it is mine.

Why I am not changing it. These are the frameworks I write from and the ones I believe are true. A tool built on this site reflects them. What I owe you is not neutrality I do not have — it is disclosure, so you can discount my numbers accordingly.

Every scoring rule, including that one, is published in full on the metrics and change log page, with the date it was added and the reason.

Where this test is weak

An honest test names its own limits.

The progressive article won on documents, not on worldview. It is investigative reportage built on records and FOIA material. Documents are worldview-neutral, so what this really proves is that the tool rewards proof regardless of the politics attached to it. That is worth something, but it is not the same as proving the tool judges philosophies evenly.

Its subject sits inside my framework. A federal agency contradicting its own paperwork about accountability records is government overreach and concealment — something my worldview and a progressive worldview condemn for different reasons. I picked an article we both dislike the target of.

The test that would actually settle it has not been run. That would take a piece built on a competing framework, supported just as well as one of mine and structured the same way, to see whether the tool extends it the same latitude it extends me. Until that is done, the asymmetry above is disclosed rather than measured.

Six data points is not a study. The pattern is clean and it points where I say it points. Calling it a demonstration is fair. Calling it proof would be the kind of overreach the tool itself flags.

What the tool measures — and what it cannot

The Article Values Graph measures whether a piece submits itself to being checked: whether it names who it means, shows where its facts came from, and lets a stranger verify the claims. It does not measure whether the conclusion is correct. A scrupulously sourced argument can still be wrong, and it will score well here.

I think that is the stronger thing to measure anyway. What crosses the political divide is not agreement about conclusions. It is the method by which conclusions can be contested at all. Judd Legum and I would fail each other's politics and pass each other's audit. That common ground exists before the disagreement and survives it.

Which is, not incidentally, a premise of the tradition I write from. Natural law holds that there are standards accessible to any person by reason, not granted by tribe or authority. I built an instrument to measure article quality, and it kept finding that the thing predicting quality was the thing available to everyone.

Check my work

Every piece in the table above can be re-scanned. Scoring is deterministic, so the same article produces the same grades every time, and every grade comes with the quotes that produced it. If you think a number is wrong, the evidence for it is printed underneath it — argue with that.

If the tool ever flatters me, it is broken, and I want to know.

← Back to the Article Values Graph™