BLOGCREATIVE PROCESS
AUGUST 4, 202614 MIN READ

The Obedient Layer: What a Style Gate Actually Removes

Claudia Vaduvescu
WRITTEN BY
Claudia Vaduvescu
QUICK ANSWER

The AI register has two layers. The lexical layer (word choice, punctuation) obeys an explicit style rule immediately and completely. The structural layer (sentence-length variation) does not respond to instruction in the same way. Measured in a matched genre, my published prose scores 0.61 on sentence-length variation against 0.57 and 0.62 for machine essays written without and with my style rules, meaning the two are not distinguishable on the measure that is supposed to be the durable tell. Earlier tonal counts from my own working notes are withdrawn here, because one did not reproduce and the machine baseline was already constrained before the test began.

AVAILABILITY
Accepting Projects
The Obedient Layer: What a Style Gate Actually Removes
ARTICLE DEK

I keep a toolkit for catching the way machines write. Pointing it at my own published essays turned out to be the uncomfortable part: on the measure that is supposed to separate us, my edited prose and machine prose score the same.

I keep a small library of rules for catching the way machines write. A banned-word list, an em-dash budget, a structural checklist, a script that scores a batch of drafts and ranks them worst-first. It exists because I write with a model every day and I needed the register caught in production, not admired in an essay.

Pointing that instrument at my own archive was supposed to produce a tidy result. It produced two, and only one of them was comfortable.

The comfortable one is that the rules work. Tell a model to stop using a word and it stops, immediately and completely. The uncomfortable one is that when I compared my own published essays against machine essays written in the same genre, on the same kind of topic, at the same length, I could not find the difference in the place I was most confident it would be. On sentence rhythm, the measure my whole library treats as the durable structural tell, my edited prose and the machine's scored the same.

This essay is about that split. There is a layer of the AI register that does whatever it is told, and a layer that does not, and almost everything written about AI prose addresses the first one.

What is actually measured, and what is only named

The state of the evidence is lopsided, and it is worth separating what is established from what is repeated.

The vocabulary is measured thoroughly. Kobak and colleagues adapted the excess-mortality method to word frequencies across 15 million PubMed abstracts and documented an abrupt post-ChatGPT rise in specific style words, estimating that at least 13.5 percent of 2024 abstracts were LLM-processed, reaching 40 percent in some subcorpora (2025). The grammar is measured too. Reinhart and colleagues built parallel human and machine corpora and found present-participial clauses at two to five times the human rate and nominalizations at 1.5 to 2 times, with the human-machine gap larger for instruction-tuned models than for base models (2025). That last detail places the register in post-training, in the preference-tuning stage, rather than in the raw model.

Now the part that is named everywhere and counted nowhere. The features I actually notice, and that my library was built to catch, are tonal: the wonder register ("there's something beautiful about", "I love this question"), the diminutive softeners, and the negate-then-assert cadence, "it's not X, it's Y", which reads like a drumbeat once you have seen it. Journalists name these constantly. Researchers name them in commentary. As far as I can establish, nobody has quantified them in a corpus. The method exists, in Reinhart's parallel design. It has not been pointed here.

So I tried to point it here myself, and I want to tell you how that went before I tell you what I think, because the failure turned out to be more instructive than the attempt.

The measurement that did not work

My first pass was an afternoon of counting across four piles of text: my published articles, my unedited chat messages, unconstrained machine research reports, and machine essays written under my style rules. It produced striking numbers. Then I tried to do it properly, and most of them did not survive.

Two things went wrong, and both are worth knowing if you ever attempt this.

The genres were not comparable. Personal essays against chat against research reports is not a controlled comparison, it is four different kinds of writing with different conventions. When I rebuilt the test the way Reinhart's method requires, matching genre and length and topic type across human and machine, several of the differences I had found shrank or reversed. One of them, a cadence count I had been fairly excited about, did not reproduce at all: rerunning the same search on the same corpus gave me four hits where the original run implied roughly twenty-seven. I still do not know what the first count was doing. The script was not saved, which is its own lesson.

And I could not build a clean baseline. To compare constrained machine prose against unconstrained machine prose, I need unconstrained machine prose. But the model I was generating with had my style rules sitting in its instructions, from my own setup. The supposedly unfiltered text came back with zero banned words, where genuinely unfiltered output runs closer to two per thousand. There was nothing left to suppress, because it had already been suppressed before the test began. You cannot measure a diet you are already administering.

So I am not going to report those tonal counts as findings, and the specific figures that circulated in my working notes should not be quoted. They are unreplicated, one of them is contradicted, and I would rather publish an essay with one defensible number in it than five exciting ones I cannot stand behind.

Here is the defensible one.

The number I can stand behind

I took my eleven published English essays, pulled as clean prose from the site itself, with headings, lists and references stripped out. Against them I set eight essays generated in the same genre, first-person practical writing for a designer's blog, at the same length, on topics I have never written about, half of them under my style rules and half without.

Then I measured burstiness, which is simply the variation in sentence length. High means a mix of short punches and long winding sentences. Low means everything comes out about the same size. It is the measure my library weights above all others, because it is the one that is supposed to survive editing.

My published prose scores 0.61. The machine essays score 0.57 without the style rules and 0.62 with them. Per essay, my range runs 0.46 to 0.66 with a median of 0.58; the machine's runs 0.43 to 0.75 with medians of 0.56 and 0.59. The distributions sit on top of each other. If you handed me those numbers without labels I could not tell you which pile was which.

Two things follow, and I want to be careful about which is solid.

The solid one: on this measure, in this genre, my edited published writing is not distinguishable from machine writing. That is a fact about my corpus and my instrument, with one author and a modest sample, and it is exactly the sort of claim that should be checked by someone else before it is believed. But it is measured, and I can hand over the files.

The interesting one, which I hold more loosely: my unedited writing does not look like this. The rhythm of how I actually type, before anything is tidied, is far more varied than the rhythm of what I publish. If that holds up under a cleaner measurement than the one I have, then the story is not that machines write like me. It is that the same tidying flattens both of us toward the same place, and the gate I built to remove the machine register is also removing the thing that most distinguishes mine.

I want to flag one technical trap here, because it caught me. Burstiness is extremely sensitive to what you feed it. The same essays measured from my working files, which contain headings and bullet lists and draft fragments, score 0.91. Measured as published prose, 0.61. Any burstiness number without a statement of what was included is meaningless, and I have seen several quoted without one.

The obedient layer and the stubborn one

Set the failed counts aside and the shape of the thing is still visible, because it shows up in work other than mine.

The lexical layer takes instruction. This is not really in dispute. Sam Altman announced in November 2025 that ChatGPT would finally honor a custom instruction to stop using em-dashes, which is a strange thing to have to announce and a clear demonstration of what the punctuation always was: a dial, adjustable in post-training, and adjustable per user on request. Word choice behaves the same way. My own gated machine essays contain none of the vocabulary the rules ban, and that is the least surprising result in this entire project.

The structural layer does not take instruction in the same way. The strongest evidence here is not mine. The StoryScope work, which is a preprint and should be read as provisional, reports that structural signatures survive stylistic editing at high rates while surface artifacts do not, which is the finding my entire library was reorganized around. My own measurement is consistent with it from a different angle: the gate my machine essays were written under includes an explicit instruction to vary sentence length, and the essays written under it score 0.62 against 0.57 for the ones written without it. That difference is too small to mean much. Instructing a model to be rhythmically varied does not reliably make it so.

Which means the public argument about the AI register, the one conducted in em-dashes and banned words and whether "delve" gives you away, is largely an argument about the layer that complies. The compliant layer is the one everyone can see, so it is the one everyone debates, and it is the one that will keep changing as vendors turn dials. The layer underneath is harder to see, harder to instruct, and consequently more durable.

Why the machines converge on one voice at all

If this were one company's house style it would be a curiosity. It is not, and the mechanism is reasonably well evidenced.

Preference tuning reduces diversity. Kirk and colleagues found that RLHF significantly reduces output diversity compared to supervised fine-tuning, including mode collapse across inputs (2024). Zhang and colleagues locate a cause: a typicality bias in preference data, where annotators systematically favor familiar text, pushing models toward a safe center (2025). Sharma and colleagues show that sycophancy is general across assistants, with humans and preference models alike sometimes preferring agreeable answers over correct ones (2023). That combination produces the affect people describe: pleasant, familiar, centered, agreeable.

Convergence itself is now peer-reviewed at the level of creativity. Wenger and Kenett found that model responses mirror other model responses far more than humans mirror other humans, across vendors, after controlling for confounds (2026). The boundary matters: that is measured for creative responses, not for tone. Nobody has measured cross-model convergence of the tonal register, and my own work does not either, since everything I generated came from a single model family.

The best name for the result is koineization, the process where dialects in contact level and simplify into a shared variety that sits above its sources rather than replacing them. It fits because there is no target being copied and no editor issuing a style guide, just many systems accommodating toward the same center. Where it breaks is worth saying: koineization runs on human accommodation across roughly three generations, and this ran on gradient descent in about three years. Same shape, entirely different engine. And inside the koine there are still accents, since stylometric work attributes machine output to specific model families with high accuracy. Convergence and fingerprints coexist, which is exactly what a regional accent with individual speakers looks like.

The part where this is about people

There is a cost to treating any of this as a detector, and it lands on writers like me.

Liang and colleagues found that GPT detectors misclassified more than half of non-native English samples as AI-generated while accuracy on native samples stayed near perfect, because non-native writing tends to have lower perplexity (2023). English is my second language. My library has a standing rule never to trust a detector score, adopted before I read that paper, because a false positive on my own prose is expected noise rather than information.

The "delve" affair made the same point in public. When the word was identified as a ChatGPT tell in April 2024, Nigerian writers pointed out that it is ordinary in Nigerian formal English and that the reflex to call it machine-made had a longer history behind it. There is a further hypothesis in circulation, that annotators in the preference-tuning pipeline shifted model English toward their own formal register. That is speculation rather than an established finding, and I am naming it as an open question rather than as a cause.

My own library holds both positions at once, which I did not plan and now think is the argument in miniature. It is a toolkit for removing the machine register, and it contains an explicit rule protecting my non-native constructions from being corrected out, because they are identity rather than error. That is my definition of the problem: not that the register is machine-made, but that treating it as a fingerprint punishes people whose ordinary English sits near the statistical middle.

This essay, audited

An essay about the AI house voice, drafted with a model, will drift into the AI house voice. So here is this piece through the same script, at final draft:

  • Burstiness 0.58. Above the 0.55 flag, so it does not trip the alarm. It is also slightly below my own published average of 0.61 and sits inside the range I measured for the machine essays. This essay scores like the thing it is about, and I could not fix that by knowing about it.
  • Em-dashes: zero. A house rule, and on this subject a joke I was obliged to land.
  • Banned vocabulary: one hit, "delve." It is here because the essay is about the word.
  • Negate-then-assert: three flagged. Worth opening up, since this is the construction the whole piece has been arguing about. One is the essay quoting the construction as an example, which is a mention rather than a use. One is a false positive, a sentence the regex misreads as the cadence. One is a real use, deliberate, with content in both halves, and it happens to be my thesis sentence. So the instrument that undercounts this thing in one form also overcounts it in another, and the number means nothing until somebody reads the hits.
  • Scaffolding: zero. Named references: 2.1 per 100 words, about twice the thinness threshold, which is what citing real work does.

Two of those numbers are the argument arriving on time. The burstiness figure is the one I would have quietly fixed had I not committed to printing it, and I could not. The antithesis figure is a demonstration that a flag is a prompt to look, not a verdict, which is exactly the discipline I failed to apply to my own earlier counting.

What would actually settle this

The honest status of everything above: one author, one model family, a small sample, no preregistration, and a set of tonal counts I withdrew rather than published.

What it would take is not exotic. A machine baseline generated in a clean context, with no style rules present, which is the blocker that broke my first attempt. Several models from different vendors, since a claim about convergence needs more than one lab. More than one human writer. Word lists written down and versioned before the counting starts, rather than rebuilt afterwards. And the direction of each prediction registered in advance, because two of my early results reversed when I looked again, which is the classic signature of a measurement with too many degrees of freedom.

I would rather publish the small honest version now than sit on it until it is a paper I am not going to write.

What I keep from all of it is the split, which survived losing most of its evidence. There is a layer of this register that does whatever you tell it, and a layer that does not, and the debate is almost entirely about the first. The words are a costume. The rhythm is the walk, and mine walks more like the machine than I would like.

References

  • Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. https://doi.org/10.1126/sciadv.adn5290
  • Kerswill, P. (2002). Koineization and accommodation. In J. K. Chambers, P. Trudgill, & N. Schilling-Estes (Eds.), The Handbook of Language Variation and Change. Blackwell. https://doi.org/10.1002/9781118335598.ch24
  • Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., & Raileanu, R. (2024). Understanding the effects of RLHF on LLM generalisation and diversity. ICLR 2024. https://arxiv.org/abs/2310.06452
  • Kobak, D., González-Márquez, R., Horvát, E.-Á., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27), eadt3813. https://doi.org/10.1126/sciadv.adt3813
  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779
  • Reinhart, A., Markey, B., Laudenbach, M., Pantusen, K., Yurko, R., Weinberg, G., & Brown, D. W. (2025). Do LLMs write like humans? Variation in grammatical and rhetorical styles. PNAS, 122(8), e2422455122. https://doi.org/10.1073/pnas.2422455122
  • Sharma, M., Tong, M., Korbak, T., et al. (2023). Towards understanding sycophancy in language models. ICLR 2024. https://arxiv.org/abs/2310.13548
  • Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631(8022), 755–759. https://doi.org/10.1038/s41586-024-07566-y (Author Correction: Nature 640, E6, 2025)
  • Wenger, E., & Kenett, Y. N. (2026). Large language models are homogeneously creative. PNAS Nexus, 5(3), pgag042. https://doi.org/10.1093/pnasnexus/pgag042
  • Willison, S. (2024, May 8). Slop is the new name for unwanted AI-generated content. https://simonwillison.net/2024/May/8/slop/
  • Yakura, H., Lopez-Lopez, E., Brinkmann, L., et al. (2025). Empirical evidence of large language model's influence on human spoken communication. https://arxiv.org/abs/2409.01754
  • Zhang, J., et al. (2025). Verbalized sampling: How to mitigate mode collapse and unlock LLM diversity. https://arxiv.org/abs/2510.01171

Measurement note: the burstiness comparison reported here was run on 2026-07-24 across 11 published essays (22,588 words) and 16 machine-generated essays in a matched genre (8,925 words), using the sentence-length measure in my own audit script. Prose only: headings, lists and references excluded. Earlier tonal counts from a 2026-07-08 run are deliberately not reported, for the reasons given in the text.