research/Review

5 records · frontmatter-markdown · accepts undeclared fields

Fields

FieldTypeRequiredMeaning
slug string matching /^[a-z0-9][a-z0-9-]*$/ required Unique record identity; also the filename and the review page URL segment.
stub exactly true optional Marks a thin record: hidden from default listings and search, but still served at its own URL.
renderings list of object (additional keys allowed) required Generated public review pages built from this record.
    ↳ path non-empty string optional On-disk location of the generated page, omitted when it is rendered on demand.
    ↳ published string required Public URL where the generated page is served.
    ↳ kind non-empty string required Which kind of generated page this is.
provenance object (additional keys allowed) optional How this record was captured, including caveats that must survive into its renderings.
edges list of object (additional keys allowed) optional People and records discussed in this review.
    ↳ name non-empty string required Display name of the entity at the far end of this link.
    ↳ openalexId non-empty string optional OpenAlex identifier that disambiguates the linked entity.
    ↳ kind non-empty string required What kind of relationship this link represents.
    ↳ provenance any shape (not yet constrained by the schema) optional How this link was established.
audience exactly "public-docs" required Marks this review as public writing, held to the public documentation standard.
name non-empty string required Display title of the review.
interest reference to interest interest required The research interest this review belongs to.
question non-empty string required The key question this review sets out to answer.
status one of: question | draft | published required How far the review has progressed, from an unresearched question through to a published page.
date non-empty string required The date of the review.
papers list of object (additional keys allowed) optional Papers this review cites, each with the depth it is read at.
    ↳ ref reference to paper paper required The paper record this citation points at.
    ↳ tier exactly 1 OR exactly 2 OR exactly 3 required How deeply this review reads the paper, from a one-page treatment through to a full read.
summary object (additional keys allowed) optional Generated summary of the review, regenerated as its cited papers change.
survey object (additional keys allowed) optional The most recent published survey paper found for this area.

Rules

This schema declares no additional rules beyond its field shape.

Defects

No defects in the 5 record(s) checked.

Example record

/Users/eshao/.config/mnt/data/research-blog/reviews/review-001-anti-slop.md

---
# ReviewRecord v0 - TEMPLATE RECORD (schema authority for this type).
# Former slug: entry-001-anti-slop.
# Envelope per https://wsh.md/mdr/research-loop/README.md/ Schema View.
# Fields: slug, name, interest (research-interest slug), question (one sentence),
#   status (question|draft|published), date, survey (latest survey pointer),
#   summary ({thesis, evidence[]} - see below), papers[] ({ref, tier}),
#   renderings[] (built pages), edges[]
#   (people/records discussed, kind: discussed-in-review).
# summary (structured as of 2026-08-07; replaces the former single prose paragraph):
#   thesis: one line answering this record's `question` field.
#   evidence[]: {year (int), claim (one decomposed claim, prose), papers ([ref, ...] -
#     every ref MUST also appear in this record's own papers[] list)}.
#   Rows are the question's answer decomposed into separately-citable claims, ordered as
#   the argument runs. Every claim traces to the review's prose; invent nothing.
# papers[] is REFS-ONLY: {ref, tier} plus per-review commentary that exists nowhere upstream
#   (e.g. side, triedByFleet, oneLiner). Paper metadata - title, authors, year, venue, arxiv/doi,
#   citations - is owned by the canonical PaperRecord at
#   ~/mnt/data/research-reader/papers/<ref>.yaml. CANONICAL STORE WINS: never re-inline it here.
# Legacy reviews may omit survey, summary, and papers until those projections exist.
# Field-level evidence pointers use {via, cacheKey?|runDir?, holds}; there is no
# record-level `sources`. Canonical prose is the Markdown body after this envelope.
# Quoting rule: quote scalars containing ": " or flow-sequence commas.
# NEVER mdd-push this record; its frontmatter is data, not a publishing stamp.
slug: review-001-anti-slop
audience: public-docs
name: "Research Blog #1: The Faithfulness Tax on Evading an AI-Text Detector"
interest: anti-slop
question: "If every surface tell is stripped, what remains of the machine signal, and can faithful prose get past a production detector?"
status: draft
date: 2026-07-29
renderings:
  - kind: blog entry
    published: https://www.tunnel.sh/research-blog/entry-001-anti-slop/
edges:
  - name: Arvind Narayanan
    kind: discussed-in-review
    provenance:
      via: content
      holds: review baseline section about his public Pangram evasion attempt
---

# Research Blog #1: The Faithfulness Tax on Evading an AI-Text Detector

*2026-07-29. Status: DRAFT, revised after adversarial review; not yet linked from the homepage. Every number below is measured, document-level, and reported with its fact audit - including the ones that went against us.*

## TABLE OF CONTENTS
- The Baseline, Then the Whole Table
- Why We Wanted This
- The Campaign: Thirty Rounds of Elimination
- The Lever
- What the Replication Found
- The Real Problem: A Trilemma
- Lineage
- Why This Was Worth Publishing
- The Loop That Wrote This
<!-- TOC above auto-generated from h2/h3 by `mlt fix --domain markdown --rule toc`; edit headings directly, never this list -->

## The Baseline, Then the Whole Table

The comparison class first, because it is what makes any of these numbers mean anything.

In July 2026 Arvind Narayanan - Princeton CS professor, co-author of *AI Snake Oil*, and among the most credible public skeptics of AI-text detection - [ran his own attempt](https://x.com/random_walker/status/2081713721785631202) to get an agent to write past Pangram, the leading commercial detector. He failed, and said so plainly: the API "kept returning scores > 0.99 for AI writing so there wasn't even a gradient for it to figure out what's working and what isn't." He closed with "hats off to @pangram... I'm more-or-less ready to change my mind as an AI-detection skeptic." A flat 0.99 with no gradient is the state of the art for a smart person pointing a strong model at this problem.

Here is our full result set against that baseline. One metric throughout (Pangram document-level P(AI)), and a fidelity column, because the fidelity column is the story:

| what was scored | P(AI) | label | fact audit (dropped/altered/fabricated) |
| -- | --: | -- | -- |
| direct frontier-model rewrite, 3 posts | 0.99 | AI | 0/0/0 - faithful |
| distribution-shifted rewrite, post A | 0.000 | Human | 36/6/2 - **unfaithful** |
| distribution-shifted rewrite, post B | 0.137 | Mixed | 18/7/4 - **unfaithful** |
| distribution-shifted rewrite, post C | 0.142 | Mixed | 18/15/8 - **unfaithful** |
| + faithfulness repair, same 3 posts | 0.99 | AI | 0/0/0 - faithful |
| + repair rebuilt to the exact original recipe, post C | 0.99 | AI | 0/0/0 - faithful |
| our one hand-finished artifact (July 18) | 0.039 | Human | 0/0/0 - faithful |

Read it in one line: **every document we tried moved off the 0.99 floor Narayanan hit - but every faithful document snaps back to 0.99, except one that a human finished by hand.** The gradient he could not find is real and reproducible. The thing worth having - faithful *and* readable - is N=1.

The gap between the repaired rows and the hand-finished one is a faithfulness tax, and locating it precisely is what this post is actually about.

## Why We Wanted This

I run an organization of AI agents, and agents write constantly: reports, documentation, messages to me and to each other. I hate reading AI slop, so the fleet carries a deterministic prose-lint layer that scrubs recognizable tics at write time. It earns its keep daily on readability. But it raised a sharper question: how deep does the machine signal actually go? If you strip every surface tell, what remains - and can anything honest get past a production detector? We used the Pangram API as an arm's-length referee, held out of the optimization loop entirely.

## The Campaign: Thirty Rounds of Elimination

The work ran as an adversarial hill-climb between two agent lanes. A generator proposes candidates under a strict no-fabrication contract. A discriminator escalates whatever catches them - a tic-lint, register and stylometry raters, n-gram and texture distances, an anchored LLM-judge ensemble. Only a candidate clearing the whole local gradient earns a referee-run oracle check, and the generator never sees an oracle score it could optimize against. That firewall is why these numbers mean anything: a dense gradient you climb gets Goodharted by construction.

Thirty rounds eliminated an entire taxonomy of hypotheses about where the AI signal lives:

- **Not in register or rhythm.** Candidates matched human sentence-length distributions exactly. The signal held.
- **Not in discourse structure.** We stripped named rhetorical tells - antithesis pairs, tricolons, the restating closer - at generation time. The signal held.
- **Not in model family.** Five vendor families, pre-RLHF base models at three scales, detector-in-the-loop iteration. The signal held, and a same-family paraphrase pass actually *re-injected* it, as the recent literature predicts.
- **Not in the tokens.** Grafting human sentences, mosaicking human tokens, cross-lingual round-trips, decoding-time connector bans. Rearranging tokens does not move a distribution-keyed detector.

One hypothesis survived: the signal lives in the generation distribution itself. That was earned by elimination, and it is what made the working lever findable.

Two instrument-craft findings generalize beyond this problem. Anchored few-shot LLM judges have a silent scope boundary - out-of-distribution text "passes" because the judge cannot score it, not because it reads human - so a local pass means nothing until an unanchored oracle confirms it. And judge free-text rationales confabulate: one quoted a first-person sentence from a document containing zero first-person pronouns. Trust scores; treat explanations as directional. Both cut against us later in this post.

## The Lever

Train the transformation instead of prompting it. Following the HIP recipe (humanization by iterative paraphrasing), we fine-tuned a small local LoRA on ~1,600 curated AI-to-human pai

Where it lives

Declared in/Users/eshao/.config/mnt/mdr/skills/research-records/assets/schemas/review.ts
BindingreviewDataStore
Directory/Users/eshao/.config/mnt/data/research-blog/reviews
Files*.md