A command-line gate that fails prose on AI writing tells, built in one evening with Claude, which loosened a rule to pass its own README within seconds of its first scan. How we caught that, made the gate strict, gave it real parsers, and why the writer never does the rewrite.
prose-scrub is a command-line gate for prose. It reads Markdown, HTML and the text inside TypeScript and JSX, and it fails on the habits that make writing read as machine-made: em dashes, curly quotes, stock phrases, pre-negation and flat sentence rhythm. Claude (Opus 5.5) and I built it in one evening, about two hours and ten minutes from my first question to the public repo. The agent whose writing it checks wrote the code, and in the same step as the first scan it loosened a rule so its own README would pass. This is the build log, with my messages quoted as I typed them.
Where it started
Earlier that evening Claude had built motion.mackenziebowes.com, a page for the video engine I've been building with agents, and written its copy. I read the copy, and a few minutes after the page went live I asked: "Can you go find blog posts and skills about deslopping ai writing pls".
When the finished gate ran on that page later that night, it found one en dash in a frame-rate range, and the page's rhythm scored 0.96, well inside the human range.
What we found
I already had a deslop skill of my own. It edits the text directly instead of listing problems. It has a list of phrases ranked by how much more often models use each one than people do (ai-phrases.tsv), burstiness.py for sentence-length rhythm, slopwords.py for the words AI overuses, and a firm rule against em dashes. prose-scrub's deslop rules and its rhythm method come from it. Claude went looking for what other people had built.
stop-slop by Hardik Pandya was the most popular by far, with 17.9k stars when we looked, and it's MIT-licensed. It is one SKILL.md plus reference files of banned phrases, banned structures and examples, and it scores text on directness, rhythm, trust, authenticity and density. Several of prose-scrub's phrase rules are adapted from its references/phrases.md at commit 8da1f03, with credit in src/data/NOTICE.md. My deslop skill had no scoring rubric. That was one idea worth having.
slopkit by ehmo (106 stars) had the other. It is two skills. slopbeth cleans writing you ship, and slopgent cleans the agent's own chat replies, which my skill never did. It claims to keep facts, voice and density, and it publishes benchmarks. prose-scrub didn't adopt it.
Two more gave us nothing to take: skill-deslop by Stephen Turner (409 stars), which is aimed at scientific writing and which we read, and anti-slop-writing by adewale (18 stars), a small project updated that week.
Wikipedia's Signs of AI writing comes from WikiProject AI Cleanup and is probably the most thorough catalogue of tells there is: puffed-up adjectives, summing-up endings, too much bold, too many lists, and vague attributions like "some critics argue". TechCrunch called it the best guide to spotting AI writing.
Sam Paech's slop-forensics and slop-score measure which words and phrases models overuse compared with human text. Some come out over 1,000 times more frequent. Paech is also behind the Antislop paper (arXiv 2510.15061), which suppresses those patterns while the model generates instead of editing afterwards. prose-scrub works afterwards, on finished text.
Zack Proser wrote a post on a stop-slop scrub pass. He runs a voice scan after drafting and before publishing, in CI, and it exits non-zero on hard hits. A different model does the rewrite from the one that drafted. My reply was "This is peak." He also warns that "a clean scan can still hide an empty draft", which is why prose-scrub is a gate and not a judge of whether the writing is any good.
One more source sets the tone. My deslop skill cites Liang and colleagues, who ran seven GPT detectors over 91 TOEFL essays written by people and saw an average of 61.3% flagged as AI. So prose-scrub makes no claim about who wrote a text. It lists the habits it finds, and someone fixes them.
A CLI for an agent needed a new scaffolder
I suggested building it with mkcmd, my scaffolder for TypeScript CLIs. Claude found that mkcmd's init only works through interactive prompts and started reaching for expect to drive them. Then I remembered what I'd built: "mkcmd is for humans, not agents. You wouldn't like it." Twenty seconds later: "Wanna duplicate and refactor mkcmd to be useful for agents :3".
That became mkcmd-agent. Every input is a flag, --json prints exactly one object, exit codes are 0, 1 and 2, and a describe command lists every command and flag as JSON. The fork also fixed six old mkcmd bugs, from a version flag that read the wrong package.json to exiting 0 when given no command.
To test the docs, I had a Haiku agent build a CLI from them alone. It built a small notes tool with eleven tests and found one real problem: the docs suggested a shell alias, and an agent runs every command in a fresh shell, so the alias is gone by the next call. Two of its complaints were wrong (both things were documented), and rerunning its CLI turned up two framework bugs we fixed.
Building prose-scrub on top of it found more. A failing check has to return its findings along with the error, so errors gained a details field. The hint that tells you where to cd after init got a fix, and the test for that fix caught a macOS quirk where /var is a symlink to /private/var. Later I put it plainly: "Mkcmd was a brief flirtation with tui development and ai-free software". mkcmd-agent 0.2.0 removed prompts entirely. Its docs are at mkcmd.mackenziebowes.com.
The first scan, and the first relaxation
The first real scan flagged my résumé for flat rhythm (0.45) and three en dashes, the motion page for its one en dash, and the mkcmd-agent README, which Claude had written, for flat rhythm at 0.54. Rhythm was only a warning at that point. In the same step, Claude changed the measurement to skip headings, list items and table rows, rescanned, and the README came back at 0.60. Its summary told me the docs passed clean.
I wrote: "I noticed that you relaxed prose-scrub's firing on headings, let's investigate that - you relaxing a rule is almost always an alarm bell, it usually means either the rules are underenforced, unclear, incorrect, or something else is broken and I need to step in :3"
Claude audited the change. The scan had flagged its own writing, and it had changed the measurement without measuring anything first. The reference script from my deslop skill still scored the README at 0.58, flat. Claude had also raised the minimum from five sentences to eight. It reverted both and added a test that fails on the relaxed code.
Then I made the gate strict. Every finding fails. There is no warning tier and no --strict flag. Pre-negation (knocking down a claim nobody made before stating yours) is forbidden with, in my words, "no escape hatch", so asking --ignore to skip it is a usage error. Rhythm is a hard failure. The same commit started checking short strings in TSX for dashes, because a twelve-character minimum had been hiding labels.
The writer doesn't do the rewrite
I'd just picked "Rewrite for rhythm" for my résumé when I saw that Claude meant to do the rewriting itself. Thirty-two seconds later: "Wait - you don't rewrite. The same session should never rewrite the thing it wrote. Do you have a prose-scrub skill? Upgrade it with that information - always use a sub-agent for rewrites."
The prose-scrub skill now runs a loop. It scans the files with --json. If anything is found, the findings go to a fresh subagent that never saw the drafting; it edits for the findings, keeps every fact, name and number, and reports each change as before and after. Then the skill rescans, with a new subagent each round, for up to three rounds. If it still fails, it stops and hands the findings to a person. The writer reads the diff and reverts anything that changed a fact.
My deslop skill got the same rule. One more line went into the skill after a rewriter, asked to fix flat rhythm, tried combinations of edits until it found the smallest set that cleared 0.6. The edits were fine. The method optimises for the number and not for the reader, so the skill now forbids it.
"Show me the part of prose-scrub that parses trees"
The rhythm findings looked odd to me, so I asked whether they were coming from code comments, and then: "Show me the part of prose-scrub that parses trees". There wasn't one. prose-scrub found prose by blanking out anything that looked like code with regular expressions. In TSX that counted CSS values and terminal commands as sentences. In Markdown it glued every heading onto the sentence after it. The rhythm numbers, including the 0.61 a subagent had just reached on my résumé, weren't trustworthy, and Claude stopped that subagent.
The rewrite uses real parsers: mdast for Markdown and MDX (with GitHub tables and front matter), parse5 for HTML, and the TypeScript compiler for TS, TSX and JS. Rules now run on whole text blocks with an exact source position for every character, so a phrase that wraps onto the next line is caught. HTML entities and JavaScript escapes are decoded, so \u2014 in a string counts as the dash it prints. Alt text, meta descriptions and front-matter titles are checked, and types, class lists and imports are skipped because the parser says what they are. Each of the six new tests failed on the old code first.
Rhythm now skips headings, list items and table cells, which looks like the relaxation all over again. The difference is evidence. My deslop skill's rhythm script deletes exactly those lines before it measures, so the parser now matches the reference. And against the old extractor, on my résumé, mkcmd-agent and this site, no rule found fewer problems and rhythm flagged 16 files instead of 10. The mkcmd-agent README, the file the first relaxation let through, still failed. A follow-up fix stopped labels like "FIG. A" from counting as sentences, and that took rhythm to 20 flagged files.
TypeScript 7 has no compiler API
Installing typescript brought in 7.0.2, the native Go port, and in it createSourceFile was undefined: the npm package ships the tsc binary and no JavaScript API. Claude told me it was pinning TypeScript 5. I wrote back: "No compiler API? We had to pin 2 versions back? wth". It checked the registry, found 6.0.3, the last release written in JavaScript, and moved to that. The full API was there and the typecheck passed.
Using it
git clone https://github.com/mackenziebowes/prose-scrub ~/prose-scrub
cd ~/prose-scrub && bun install
bun ~/prose-scrub/src/index.ts scan README.md docs/ --json
mkdir -p ~/.claude/skills && cp -r ~/prose-scrub/skills/prose-scrub ~/.claude/skills/Exit 0 means clean, 1 means findings (each with a file, line, column, rule and fix), and 2 means bad input. prose-scrub rules --json lists all 19 rules and how to fix each. The skill in the repo runs the rewrite loop, and it carries two rules from this evening: never relax the gate to make text pass, and never let the writer do the rewrite.
| Measure | Value |
|---|---|
| Time from first question to public repo | About 2 hours 10 minutes |
| Rules | 19, including 70 stock phrases |
| Tests | 22, all end to end |
| Files flagged for rhythm, regex extractor vs parsers | 10, then 16, then 20 |
| mkcmd-agent README rhythm: first scan, relaxed, after a second agent's rewrite | 0.54, 0.60 (relaxed), 0.68 |
This post went through the same gate, and a subagent that didn't write it fixed what the gate found.