Back to Posts
Build log6 min read

prose-scrub 0.4: a small model as the second reader

Day two of prose-scrub. A bet about headings that I lost, a judge command that asks Claude Haiku five times and averages, a score field the model filled with sentences, and a test on my own posts where the gate made one of them worse.

prose-scrub is the gate I built with Claude yesterday. It fails prose on dashes, stock phrases, pre-negation and flat rhythm. Today we added a second reader: a command called judge that asks Claude Haiku 5.5 about two habits no word list can see, five times per file. Then I ran the whole thing on four of my own posts and my résumé, read the results in a side-by-side viewer, and found the place where the gate makes my writing worse. I'm shipping it anyway, as 0.4.0. This is the build log.

A bet about headings

Wikipedia's guide to signs of AI writing lists undue emphasis on significance: text that keeps telling you its subject matters. I suspected my last post did it, so I asked for a test: "Run a haiku across the prose in the md and ask it to rank the headings from 0.0 to 1.0 on "how important does the author want me to think this section is" and I bet it'll rank every heading as 1.0".

I lost. Three Haiku runs gave means of 0.47, 0.52 and 0.50, from 0.05 for the source list up to 0.90 for the section on the first relaxation. But all three runs quoted the same seven lines as the reason a section sounded important, lines like "The one that changed the plan was" and "One line in that summary gave it away". Those seven were the real finding. The rewrite cut them.

Each of those runs was a Claude Code subagent, and each cost about 65,000 tokens. I asked for a cheaper version that sends only the headings, and Claude added a headings command. It saved almost nothing: 58,500 tokens instead of 65,000, because nearly all of a subagent's cost is its own setup, not the post. Calling the API directly with a two-line system prompt costs a few thousand tokens for a whole post. So the judge talks to the API, with a key kept in the macOS Keychain where neither the command line nor Claude ever sees it.

What to ask a small model

Before writing the judge we went back through the sources. Wikipedia's guide has sections I half remembered, on avoiding plain "is" and "are" and on superficial analysis. stop-slop scores text on directness, rhythm, trust, authenticity and density. slopkit benchmarks its rewrites with judge panels.

Two patterns needed a reader rather than a list. Significance is scored per paragraph: how important the paragraph says its subject is, and separately how much evidence it gives, with a flag when the first is much higher. Superficial analysis is scored per sentence: does it add new information, or only interpret what came before? A new units command numbers every paragraph and sentence so each run scores exactly the same pieces.

Copula avoidance ("serves as", "boasts" and "stands as" where "is" would do) looked like a job for a plain rule, so Claude wrote one. It found two real cases on my site. I still didn't want it shipped: a specific blacklist grows without bound, and nobody had checked that the first list was right. It's parked on its own branch until I can label a pile of sentences by hand, AI-written ones mixed with human writing from before 2022.

Five runs, averaged

Claude proposed three runs and a median. I asked for five, averaged: "Therefore, if a Haiku Claude disagrees with parallel Haikus, that is signal and not noise, so averaging is the correct move. Probably." My experience with agents is that smaller models fail by being lazy, not by missing things. A run that disagrees may have looked harder.

On a short test file about a bridge, with one inflated paragraph and one empty closing sentence planted, the runs flagged that paragraph, its two sentences and the empty one, and nothing in the paragraph of plain facts. Across that file and my last post, they agreed on 158 of 330 scores. Almost every disagreement was one step on Haiku's grid, one run saying 0.25 where four said 0.3. No single run was the odd one out, and disagreement didn't grow toward the end of long lists. That's too little data to explain anything. Every run's raw scores now go to a log, and finding the cause is on my list.

Then I tried a scale of 0 to 50, hoping Haiku would use decimals. It used none. On significance it moved to a finer grid of whole numbers, and on superficial it stayed on ten steps. Runs agreed more closely on both, so 0 to 50 is now the default.

A field named claim

On my E-E-A-T post, four of the five significance runs returned no scores at all. That looked like the laziness I'd predicted. It wasn't. The score field was called claim, and Haiku read it as "describe the claim": it wrote sentences where the numbers should go. Claude renamed the fields to importance_score, evidence_score and empty_score and made the tool schema strict. All five runs came back complete.

Testing it on my own posts

Every post on my site was vibe-written. I picked four, plus my résumé as a control, and had Claude run the full loop on copies: scan, judge, a fresh subagent per file to rewrite, up to three rounds. My lamina CLI turned each before and after into a page styled for reading, and Claude built a small viewer to compare them side by side, as a word-level diff and as a change log.

Before and after one or two rewrite rounds, 2026-10-09. The OG post's two remaining findings sit in a quoted card and a section name.
FileGate findingsRhythmJudge flagsWords
E-E-A-T requires custom websites9 to 00.63 to 0.6251 to 232205 to 1985
Generative OG images4 to 20.58 to 0.610 to 12731 to 2732
Vibe manufacturing1 to 00.49 to 0.610 to 0729 to 724
Where small business owners hang out online14 to 10.54 to 0.5210 to 81810 to 1746
Résumé (control)00.630813

The gate found real things to fix. It also made the small-business post worse.

Where clean reads worse

The diff of that post's opening showed the problem. The two signposts ending in "I found." and "I learned." were gone, cut by the throat-clearing rule. Two adverbs were gone, cut by the adverb rule. The judge marked "Okay." as a sentence that adds nothing, so "Okay. But where is that, exactly?" became "Where is that, exactly?" The rewriter also dropped a "But" that no rule had asked it to touch.

Those were there on purpose. The signpost before the data tells the reader something interesting is about to happen, and that prepares them for it. Readers run on what they expect to feel next. People start sentences with "But" all the time. Clean isn't always easier to read.

So I asked where the throat-clearing rule came from. prose-scrub copied it from stop-slop, which calls every signpost of that shape throat-clearing and lists no exceptions. A word list can't tell a signpost that pays off from one that doesn't. In the end I kept the rule: "I accept slightly worse writing for avoiding all throat clearing."

The judge is a different case, because it's advisory. It flags setups like "Okay." that are doing their job, so the skill now shows judge flags to a person and never hands them to a rewriter on its own.

Using it

Store the API key without echoing it, then judge a post
security add-generic-password -U -a "$USER" -s anthropic-api-key -w
cd ~/prose-scrub && git pull && bun install
bun src/index.ts units post.md --json
bun src/index.ts judge post.md --log runs.jsonl --json

judge always exits 0 and reports a flagged count. Each paragraph or sentence keeps all five scores and their spread. The thresholds are provisional until they're checked against human writing. A post of about 2,000 words takes 10 API calls, roughly 50,000 tokens in and 15,000 out.

Running prose-scrub or its judge on your own writing? Open an issue on GitHub or send an email.

Talk to me
More Articles