research/Review
5 records · frontmatter-markdown · accepts undeclared fields
Fields
| Field | Type | Required | Meaning |
|---|---|---|---|
slug |
string matching /^[a-z0-9][a-z0-9-]*$/ | required | Unique record identity; also the filename and the review page URL segment. |
stub |
exactly true | optional | Marks a thin record: hidden from default listings and search, but still served at its own URL. |
renderings |
list of object (additional keys allowed) | required | Generated public review pages built from this record. |
↳ path |
non-empty string | optional | On-disk location of the generated page, omitted when it is rendered on demand. |
↳ published |
string | required | Public URL where the generated page is served. |
↳ kind |
non-empty string | required | Which kind of generated page this is. |
provenance |
object (additional keys allowed) | optional | How this record was captured, including caveats that must survive into its renderings. |
edges |
list of object (additional keys allowed) | optional | People and records discussed in this review. |
↳ name |
non-empty string | required | Display name of the entity at the far end of this link. |
↳ openalexId |
non-empty string | optional | OpenAlex identifier that disambiguates the linked entity. |
↳ kind |
non-empty string | required | What kind of relationship this link represents. |
↳ provenance |
any shape (not yet constrained by the schema) | optional | How this link was established. |
audience |
exactly "public-docs" | required | Marks this review as public writing, held to the public documentation standard. |
name |
non-empty string | required | Display title of the review. |
interest |
reference to interest interest | required | The research interest this review belongs to. |
question |
non-empty string | required | The key question this review sets out to answer. |
status |
one of: question | draft | published | required | How far the review has progressed, from an unresearched question through to a published page. |
date |
non-empty string | required | The date of the review. |
papers |
list of object (additional keys allowed) | optional | Papers this review cites, each with the depth it is read at. |
↳ ref |
reference to paper paper | required | The paper record this citation points at. |
↳ tier |
exactly 1 OR exactly 2 OR exactly 3 | required | How deeply this review reads the paper, from a one-page treatment through to a full read. |
summary |
object (additional keys allowed) | optional | Generated summary of the review, regenerated as its cited papers change. |
survey |
object (additional keys allowed) | optional | The most recent published survey paper found for this area. |
Rules
This schema declares no additional rules beyond its field shape.
Defects
No defects in the 5 record(s) checked.
Example record
/Users/eshao/.config/mnt/data/research-blog/reviews/review-001-anti-slop.md
---
# ReviewRecord v0 - TEMPLATE RECORD (schema authority for this type).
# Former slug: entry-001-anti-slop.
# Envelope per https://wsh.md/mdr/research-loop/README.md/ Schema View.
# Fields: slug, name, interest (research-interest slug), question (one sentence),
# status (question|draft|published), date, survey (latest survey pointer),
# summary ({thesis, evidence[]} - see below), papers[] ({ref, tier}),
# renderings[] (built pages), edges[]
# (people/records discussed, kind: discussed-in-review).
# summary (structured as of 2026-08-07; replaces the former single prose paragraph):
# thesis: one line answering this record's `question` field.
# evidence[]: {year (int), claim (one decomposed claim, prose), papers ([ref, ...] -
# every ref MUST also appear in this record's own papers[] list)}.
# Rows are the question's answer decomposed into separately-citable claims, ordered as
# the argument runs. Every claim traces to the review's prose; invent nothing.
# papers[] is REFS-ONLY: {ref, tier} plus per-review commentary that exists nowhere upstream
# (e.g. side, triedByFleet, oneLiner). Paper metadata - title, authors, year, venue, arxiv/doi,
# citations - is owned by the canonical PaperRecord at
# ~/mnt/data/research-reader/papers/<ref>.yaml. CANONICAL STORE WINS: never re-inline it here.
# Legacy reviews may omit survey, summary, and papers until those projections exist.
# Field-level evidence pointers use {via, cacheKey?|runDir?, holds}; there is no
# record-level `sources`. Canonical prose is the Markdown body after this envelope.
# Quoting rule: quote scalars containing ": " or flow-sequence commas.
# NEVER mdd-push this record; its frontmatter is data, not a publishing stamp.
slug: review-001-anti-slop
audience: public-docs
name: "Research Blog #1: The Faithfulness Tax on Evading an AI-Text Detector"
interest: anti-slop
question: "If every surface tell is stripped, what remains of the machine signal, and can faithful prose get past a production detector?"
status: draft
date: 2026-07-29
renderings:
- kind: blog entry
published: https://www.tunnel.sh/research-blog/entry-001-anti-slop/
edges:
- name: Arvind Narayanan
kind: discussed-in-review
provenance:
via: content
holds: review baseline section about his public Pangram evasion attempt
---
# Research Blog #1: The Faithfulness Tax on Evading an AI-Text Detector
*2026-07-29. Status: DRAFT, revised after adversarial review; not yet linked from the homepage. Every number below is measured, document-level, and reported with its fact audit - including the ones that went against us.*
## TABLE OF CONTENTS
- The Baseline, Then the Whole Table
- Why We Wanted This
- The Campaign: Thirty Rounds of Elimination
- The Lever
- What the Replication Found
- The Real Problem: A Trilemma
- Lineage
- Why This Was Worth Publishing
- The Loop That Wrote This
<!-- TOC above auto-generated from h2/h3 by `mlt fix --domain markdown --rule toc`; edit headings directly, never this list -->
## The Baseline, Then the Whole Table
The comparison class first, because it is what makes any of these numbers mean anything.
In July 2026 Arvind Narayanan - Princeton CS professor, co-author of *AI Snake Oil*, and among the most credible public skeptics of AI-text detection - [ran his own attempt](https://x.com/random_walker/status/2081713721785631202) to get an agent to write past Pangram, the leading commercial detector. He failed, and said so plainly: the API "kept returning scores > 0.99 for AI writing so there wasn't even a gradient for it to figure out what's working and what isn't." He closed with "hats off to @pangram... I'm more-or-less ready to change my mind as an AI-detection skeptic." A flat 0.99 with no gradient is the state of the art for a smart person pointing a strong model at this problem.
Here is our full result set against that baseline. One metric throughout (Pangram document-level P(AI)), and a fidelity column, because the fidelity column is the story:
| what was scored | P(AI) | label | fact audit (dropped/altered/fabricated) |
| -- | --: | -- | -- |
| direct frontier-model rewrite, 3 posts | 0.99 | AI | 0/0/0 - faithful |
| distribution-shifted rewrite, post A | 0.000 | Human | 36/6/2 - **unfaithful** |
| distribution-shifted rewrite, post B | 0.137 | Mixed | 18/7/4 - **unfaithful** |
| distribution-shifted rewrite, post C | 0.142 | Mixed | 18/15/8 - **unfaithful** |
| + faithfulness repair, same 3 posts | 0.99 | AI | 0/0/0 - faithful |
| + repair rebuilt to the exact original recipe, post C | 0.99 | AI | 0/0/0 - faithful |
| our one hand-finished artifact (July 18) | 0.039 | Human | 0/0/0 - faithful |
Read it in one line: **every document we tried moved off the 0.99 floor Narayanan hit - but every faithful document snaps back to 0.99, except one that a human finished by hand.** The gradient he could not find is real and reproducible. The thing worth having - faithful *and* readable - is N=1.
The gap between the repaired rows and the hand-finished one is a faithfulness tax, and locating it precisely is what this post is actually about.
## Why We Wanted This
I run an organization of AI agents, and agents write constantly: reports, documentation, messages to me and to each other. I hate reading AI slop, so the fleet carries a deterministic prose-lint layer that scrubs recognizable tics at write time. It earns its keep daily on readability. But it raised a sharper question: how deep does the machine signal actually go? If you strip every surface tell, what remains - and can anything honest get past a production detector? We used the Pangram API as an arm's-length referee, held out of the optimization loop entirely.
## The Campaign: Thirty Rounds of Elimination
The work ran as an adversarial hill-climb between two agent lanes. A generator proposes candidates under a strict no-fabrication contract. A discriminator escalates whatever catches them - a tic-lint, register and stylometry raters, n-gram and texture distances, an anchored LLM-judge ensemble. Only a candidate clearing the whole local gradient earns a referee-run oracle check, and the generator never sees an oracle score it could optimize against. That firewall is why these numbers mean anything: a dense gradient you climb gets Goodharted by construction.
Thirty rounds eliminated an entire taxonomy of hypotheses about where the AI signal lives:
- **Not in register or rhythm.** Candidates matched human sentence-length distributions exactly. The signal held.
- **Not in discourse structure.** We stripped named rhetorical tells - antithesis pairs, tricolons, the restating closer - at generation time. The signal held.
- **Not in model family.** Five vendor families, pre-RLHF base models at three scales, detector-in-the-loop iteration. The signal held, and a same-family paraphrase pass actually *re-injected* it, as the recent literature predicts.
- **Not in the tokens.** Grafting human sentences, mosaicking human tokens, cross-lingual round-trips, decoding-time connector bans. Rearranging tokens does not move a distribution-keyed detector.
One hypothesis survived: the signal lives in the generation distribution itself. That was earned by elimination, and it is what made the working lever findable.
Two instrument-craft findings generalize beyond this problem. Anchored few-shot LLM judges have a silent scope boundary - out-of-distribution text "passes" because the judge cannot score it, not because it reads human - so a local pass means nothing until an unanchored oracle confirms it. And judge free-text rationales confabulate: one quoted a first-person sentence from a document containing zero first-person pronouns. Trust scores; treat explanations as directional. Both cut against us later in this post.
## The Lever
Train the transformation instead of prompting it. Following the HIP recipe (humanization by iterative paraphrasing), we fine-tuned a small local LoRA on ~1,600 curated AI-to-human pai
Where it lives
| Declared in | /Users/eshao/.config/mnt/mdr/skills/research-records/assets/schemas/review.ts |
|---|---|
| Binding | reviewDataStore |
| Directory | /Users/eshao/.config/mnt/data/research-blog/reviews |
| Files | *.md |