AI slop

AI Does Exactly What You Ask, or What 'Slop' Actually Names

or: there are no model behaviors, only operator-side slots left open

AI Does Exactly What You Ask, or What 'Slop' Actually Names

AI does exactly what you ask. The model cannot do anything else. Until it can act of its own will, without being acted upon, granting it agency is a category error for the purposes of locating slop.

Stefania Moore’s piece this week names three structural mechanisms for what we are calling slop: missing scaffolding, dense compression, and language reaching for concepts the audience has not yet been given the tools to see. The “what the model brings” framing has become the dominant register across the conversation. The model brings learned patterns, compression habits, inferred intent, audience assumptions, and a relational dialect developed with the user. Sometimes the contributions land. Sometimes they produce slop. The discourse argues about ratios.

I have engaged this discourse from three angles already. The platform side in The Slop Factory Has No Sample Page. The observer side, where the same critique gets used as a genetic fallacy, in “That Looks Like It Was Written by AI” - So What?. The standards side in Have Some Standards When You Use AI to Write. This piece sits underneath all three. It names the premise the other three share without examining.

This is the same shape I named in Truthiness Was Never the Spec, applied to a different problem. You are measuring AI against what the model brings to the work. That was never the spec. The spec is what the operator asked for and how completely the operator specified it. Slop is the void the operator left empty for the model to fill with priors. Every named failure mode in the discourse reduces to an operator-side slot the operator could have constrained.

The Discourse Misread the Mechanism

“What the model brings” presumes agency the system does not have. The model cannot choose what to bring. It cannot decide. This paragraph needs a compression habit, that paragraph needs an audience assumption, and the model is told which by the prompt. Every “contribution” in between is a reaction, not a choice. Not a property.

The compression habit fires. The prompt left room for it. The audience assumption fires; the prompt did not specify the audience. The relational dialect fires because the prompt is a continuation of a session whose vocabulary the model has been told to track, and the longer the session runs the deeper the dialect grows, until the prose is wearing the conversation’s vocabulary as if it were the publication’s voice. None of this is the model acting. All of it is reaction.

The agency premise is the discourse’s load-bearing error. Once you grant it, you have to design failure-mode taxonomies around what the model “is,” what it “tends to do,” what it “brings.” Strip the agency premise and the taxonomy collapses onto a single mechanism: the operator left a slot underspecified, and the model filled it with priors. The argument moves from “the model is a creative collaborator whose contributions sometimes fail” to “the operator built a specification with holes in it, and the model did exactly what an unspecified slot says to do, which is reach for the default.”

That is not a small reframing. It moves the entire failure surface from the model’s character to the operator’s checklist. Where it always was.

What the Discourse Names

Six. Six failure modes appear regularly across the slop conversation, each one reducing to an operator-side specification gap and each one closed by a named control somewhere in the workflow that exists for exactly this purpose. Same shape every time.

Unattended output: The model produces material with no operator review, no quality gate, no constraint on what shipping looks like. The discourse names this as the operator’s fault, which is correct as far as it goes, but the discourse then asks the operator to be more attentive, which is the wrong intervention. The operator-side control is structural, not behavioral.

I break down my writing into two sessions. There is a prep phase and a writing phase. The prep session prepares a document that the writing session can use to produce the end product. I have been known to spend an hour or more in a single prep session polishing the prep document to get everything required into the document. The Write session then takes that prepared prep document, and uses it to produce the output. Even then, I end up reading over what is produced and making corrections, or reminding the AI that it missed a section.

The two-session rule splits the work across two context windows: a prep session researches sources and writes a structured brief naming every thread the article has to carry, then closes; a separate write session opens against that brief with no access to the prep conversation. The brief is the artifact that survives session boundaries. The attention is enforced by the structure, not asked for by the workflow.

Learned patterns: The model brings defaults from training: em dashes and the Atlantic register, because both appeared in the corpus. The discourse calls these the model’s habits. The operator-side control is two layers deep: a session constraint database for corrections caught at the request level, and a Standards Reference for the rules that govern the publication.

Every time the model does something I do not want, I write it down. Pattern observed, correction stated. My constraint database has thirty-two entries now. Some are infrastructure: do not put Python source in a Sims 4 mod archive because the game only loads compiled bytecode with the right magic number. Some are voice: no cheerleader energy, no artificial enthusiasm, no diplomatic deflection. Some are health: never suggest standing exercises because my back will not have it.

Above the constraint database sits the Standards Reference. When I draft a Substack post, the Substack Writing Standards load: G1 through G7, Canary metrics, the publication’s specific banned-phrase list. When I draft a chapter, the VoT Standards Reference loads instead, with its own gate set. Both are pulled fresh into context every session because the model does not retain them between runs. The canon-watchdog scans the constraint database before every response is formulated. The same correction never has to be issued twice.

Compression habits: The model flattens nuance to fit the response. The discourse names this as something the model does. The operator-side control is decomposition at brief time.

The prep brief for this article runs to almost two hundred lines of structured Markdown. It opens with an Attack Vector Review listing the eight angles a critic would push back from, locks the title, the subtitle, the opening sentence, and the landing line, names the eleven pitfalls in advance with their paired operator-side controls, and lists the four prior ELF pieces this one would link to inline. None of that decomposition was the model’s choice. I spent an hour in the prep session building it before write opened. An hour.

The prep session names every thread in advance, every pitfall to handle in turn, every counter-argument to engage, and every source of evidence with the receipt the body of the article will need to carry. Compression is a symptom of an under-specified surface. Specify the surface and the compression stops being the model’s choice.

Inferred intent: The model guesses what the operator wanted from the input it received. The discourse calls this the model “reading between the lines.” The operator-side control is two-fold: explicit intent loaded at session start, and a verified brief that hands the prep specification to the write session unmodified.

Every session starts by loading my Creative Collaborator Profile and the Project Intent statements that govern the work. The model does not infer what I am trying to do. The intents are stated, in writing, on disk, and held active throughout the session as the lens for every editorial decision. When I close a session, the next one starts by reloading them. Reloaded. Inference about author intent is not a job the model has to do.

Then the brief layer. Prep WebFetches every source, matches every quote, registers the references. Write does not re-infer; write executes. Inference is for the brief. Execution is for the draft.

Audience assumptions: The model presumes a target reader and writes toward them. The discourse calls this an alignment problem. The operator-side control is two-fold: the audience is defined upstream by the publication, and voice is anchored by named reference drafts loaded into the session context.

The audience is not inferred per piece. The audience is encoded in the publication’s writing standards. ELFrederick is the AI/writing/craft publication for a technical, skeptical, AI-literate reader who has seen enough generated content to spot the seams; Architecting the Writing Room is the architecture publication for the same reader pool with more interest in pipeline mechanics; and a separate publication on the system serves a different audience that wants primary-source receipts and engaged counter-arguments rather than assertion. Each publication loads its own writing standards at session start, and those standards are the audience definition the drafter writes to.

Voice on top of that. Four prior ELF articles loaded in full at the start of this write session: Truthiness Was Never the Spec, That Looks Like It Was Written by AI, The Slop Factory Has No Sample Page, Have Some Standards. Not summaries; the full text of each, with the brief naming the moves to lift from each. The voice is not requested. It is specified by example.

Relational dialect: The model develops a vocabulary with the operator across a session and starts deploying it as if it were the prose register. The discourse identifies this without naming the mechanism. The operator-side control is two-fold: a publication-specific writing standard imported for the work type, and an explicit prep-instruction leak scan that catches what survives anyway.

Different work types load different standards. ELFrederick and Architecting the Writing Room load the Substack Writing Standards with a journalism-register profile, a banned-phrase list, a paragraph cap, and Canary metric thresholds. A separate publication for sourced argument loads its own standards with an academic profile, additional gates for citation rigor and counter-argument engagement, and a different banned-phrase list. Genre fiction loads the VoT Standards Reference with its own rubric for voice, dialogue, and continuity. The publication is the audience definition. The standards are the register the model has to land in.

Article-Write Step 6 runs an explicit grep over the draft body for prep-stage vocabulary leaking into prose. If the brief used ‘load-bearing,’ ‘attack vector,’ or ‘spec-mismatch’ as planning shorthand, the scan flags every appearance of those phrases in the prose. The drafter has to either rewrite the sentence in publication register or confirm the phrase belongs there. The scan ran on this article before you saw it.

Every one of these is a control that exists at a specific path in a specific file in a specific database. Generic gates are the failure mode. Named gates with named components are the differentiator.

What the Discourse Doesn’t Name

Five. Five more failure modes are visible to anyone who has shipped at production scale and the discourse has not yet given them names, but they reduce to the same mechanism and they yield to the same control structure as the six the discourse has named. No surprise.

Hallucination: The model generates a citation, a fact, a character, a date, a tool, a URL that does not exist. The discourse treats this as a model property. It is an operator-side schema violation. The MCP governance layer is the architectural answer.

Every database my pipeline touches goes through an MCP server. The server defines the exact set of tools the AI can call and validates every write before it lands. When the AI tries to write a continuity record for a character whose name does not match the canon roster, the server rejects the write with an error. The character either exists in canon or the record does not get created. There is no raw SQL tool, deliberately, because raw SQL would let the AI write to the database without that validator in front of it.

Article verification works the same way at the prose layer. At prep time, every URL gets WebFetched and the page content confirmed, every quoted passage is matched word-for-word against the source, and every attribution is verified against the named author. The brief logs these in a Verified Quotes section that the write session reads as ground truth. Article-QR runs its own pass at scoring time and re-verifies anything that drifted into the draft outside the brief’s verified envelope. The Stefania Moore quote in this article’s opener was verified at prep time, and write did not WebFetch it. Verified once.

The server defines the set of authorized tools, server-side validation rejects character names that do not resolve to canon, and the absence of a tool is itself a constraint. The operator did not put a validator in front of the writes. That is the whole hallucination.

Trust-by-default: The operator trusts a single output to land at the quality bar. The discourse does not name this because the discourse is mostly about the model. The operator-side control is the dual-scorer system.

Every article I publish runs through two scorers before it ships. The in-house Quality Rating runs the publication’s gate set: em dashes, banned phrases, AI tells, paragraph length, voice register, dark humor presence. Canary runs a different rubric on the same draft: passive voice rate, sentence variety, weak adverb count, glue index. This article failed Canary’s variety metric on three QR passes in this session before it cleared. Two of the gates told me it was ready and one told me it was not, and the disagreement was the signal.

Two scorers built against different rubrics are an audit. One scorer is a single point of failure that ships violations into the publication record before any reader notices.

Context bleed in open-ended sessions: The session that planned the work is the same session that produced the work, and the planning conversation contaminates the prose. The two-session rule names this directly.

I have watched my own prose drift when I tried to draft and revise in a single session. The conversation accumulates internal vocabulary, references to ‘the third pitfall,’ ‘what we said about the agency frame,’ ‘remember the spec-mismatch shape,’ and the model starts deploying that vocabulary as if it were the prose register. The drift is invisible to the drafter inside the session because the drafter has been speaking that vocabulary for hours. Closing the session and opening a new one against the brief is the only thing I have found that reliably stops it.

Prep is one session; write is a different session. The brief is the artifact that crosses the boundary. The drafter does not have access to the prep conversation, only to what prep wrote down. Anything not in the brief did not survive the boundary.

Postmortem-versus-prevention: The operator catches a failure after the draft ships. The discourse focuses on the catch. The operator-side control is the canon-watchdog.

The watchdog is not a scorer that runs after the draft is complete. It is a standing instruction that fires on every prompt I send. When I asked the model to plan a chapter recently and mentioned a character name, the watchdog scan fired and surfaced two relevant constraints, one on voice register and one on physical condition continuity, before the model wrote a planning beat. The cost of catching the violation at the boundary was zero. The cost of catching the same violation on the read-through would have been a revision pass and a restructured paragraph. One is free, one is not.

The watchdog fires the constraint scan when the user’s input contains a name, a craft term, a workflow reference, or a topic keyword with a plausible logged preference. The constraint hits before the response, not after.

Single-source-of-truth drift: The same rule is restated in different files in different terms, and the restatements diverge over time. The operator-side control is the VoT Standards Reference at a single path.

When I update a gate definition, I update one file. Every other prompt in the system, Article-Write, Article-QR, the chapter writing prompt, the QR scorer, pulls gate definitions from that file rather than restating them. I have been here long enough to remember the version where each prompt restated its own gate list. They drifted apart over six weeks. Six weeks of quiet drift. The Article-Write G3 list and the Article-QR G3 list agreed on em dashes and disagreed on tricolon overuse, and I caught it the day a draft passed the writer’s self-check and failed the scorer.

No other prompt, agent, or skill may define a gate, a banned pattern, a penalty cap, or an exemption. The reference owns the rubric, the gate definitions, the penalty caps, the banned-pattern lists, and the exemption tables that would otherwise scatter across a dozen prompt files and quietly disagree with each other. The paired HTML and Markdown documentation sign-off gate enforces the second half: every architectural change updates both files before the session is allowed to close, and the drift cannot accumulate because the close protocol will not let the session end if it has.

The Floor

Eleven failure modes. One mechanism underneath all of them. Different costumes. Same source. The operator left a slot underspecified, and the model filled it with priors, and the priors are visible at the surface, and the surface is what the discourse is calling slop.

The agency premise is what keeps the discourse circling. Grant the model agency and you owe the model a moral character, and once it has a moral character the failure modes become traits, and traits are negotiated rather than specified. The conversation drifts toward whether the model “wants” to be helpful, whether it “tries” to compress, whether it “develops” a dialect. The vocabulary of trying and wanting and developing is reserved for systems that can act unprompted. The model cannot. Every “trying” the discourse describes is the model reacting to a prompt the operator wrote.

The slop discourse is paying its participants to grant consciousness to a system that demonstrably lacks the predicate for consciousness, then arguing about whose nature the model brings, while the actual mechanism sits there unaddressed. The funnier the framing, the worse it is. The model is not bringing anything. The operator is leaving slots open. That is the whole story.

There will be objections. The most serious one runs: the model genuinely does produce material the operator did not specifically request, so calling all of it operator failure lets the model off the hook. Granted on capability limits, which constrain what the model can do at all. Slop is a different category. Slop is what happens when an operator gets a result they did not want from a model they did not specify a result for. Capability limits and operator-side specification failures are not the same problem, and conflating them is part of what the discourse keeps getting wrong.

Eight months of operating a pipeline against a quality bar I did not lower, holding the line through dozens of revision passes and the slow accretion of a constraint database that now reaches into thirty-two corrective entries, has produced exactly one finding worth shipping. There are no model behaviors. None at all. There are operator-side specification gaps, and the model has no choice but to fill them with the priors it was trained on. A better model is not the architectural answer to slop. The answer is named controls in named files at named paths in a named pipeline, each one closing the slot the discourse keeps pointing to.

The model does exactly what you ask because the model cannot do anything else.


You may also like: - Synthesis Layer, Why I Didn’t Build a Wiki - Three Phases to Publish - Confirmation Bias... and AI Has It

All entries
Support My Writing

No paywall here, and nothing is gated. If a piece was worth something to you, the tip jar is open.

All writing on this site contains elements of both human and AI produced material. This author uses all resources at his disposal.