In January 2026, Dario Amodei published The Adolescence of Technology, a roughly 20,000-word essay on the risks of powerful AI. When NBC News asked whether Claude had written it, he said he used Claude for ideas and research, and that the writing itself was his.
His reason took one sentence:
“I don’t think Claude is quite good enough yet to write the whole thing.”
In the same interview he described some of Anthropic’s own engineers, who told him they barely write code anymore.
Claude will writes all the code and engineers only need to review and fix it.
Here’s the paradox: AI fundamentally cannot tell the difference between elite prose and derivative garbage, just like it can’t tell clean, architectural code from a chaotic mess of glue code holding broken logic together.
While Anthropic runs victory laps bragging about model capabilities, AI is clearly not “smart enough” to cross this qualitative chasm. You can still spot the low-grade, plastic AI slope a mile away by the sheer predictability of its sentence structure.
For all the hype about impending AGI, the machines still can’t write a genuinely compelling article—and deep down, Anthropic’s CEO Dario knows it.
Dario passes this off as a temporary product quirk, claiming his own writing style is simply “unique” and that Claude just isn’t quite there yet—as if the next model weights drop will finally crack it.
In fact, he is inadvertently explaining the fundamental limitations of AI itself:
It’s not about whether the AI can generate a particular style of article; it’s about whether the AI model truly understands prose, tone, and style.
And this is the strongest sign of so-called "AI intelligence."
AI has no taste of code quality nor article prose and style, because the AI model is not able to “see” the result while in the “generating process”.
The MODEL can only SEE/VERIFY the OUTCOME AFTER RESULTS ARE GENERATED.
Yet I am still seeing so many writers arguing whether AI has awareness or consciousness.
One way to judge whether an AI model truly can be aware or feel anything is to ask the AI model to write an article with style and articulacy,
yet unfortunately, during the writing process, AI keeps producing more AI slops than beautiful, descriptive, and touching language.
You will find most AI writings are bad because of:
Repetitive patterns of short sentences, mostly beginning with “A.”
Very limited selection of words; you will see how articles look exactly the same because AI always chooses particular words.
Missing “I/they/we/our,” and mostly adjectives and descriptive words. The AI model tends to express this in a pure coding expression rather than an emotion-based expression.
The AI model slams all conclusions and theses without any introduction.
Too much gibberish; the longer the article, the longer the gibberish.
Almost all articles start strong but fade very quickly with super weak endings.
The charts lose meaning to the point where you will question the AI model’s taste and judgment as to whether the AI really understands what you are writing about.
To fix these failures, I built an entire stack—workflows, Markdown guidelines, execution contracts, and harness structures—to train an AI agent to write in my exact style.
Until recently, I believed every bit of that effort had failed miserably.
I started by stuffing Markdown files with rules and negative constraints, eventually creating far more files than anyone should. More rules simply didn’t work: the LLM can only scan a fraction of them and fails to faithfully execute the whole corpus.
Next, I moved to contracts and a harness layer to regulate the Markdown, attempting to precisely dictate every step of what to do and what to avoid. The result was grim. While the probability dropped slightly, the core afflictions remained: the relentless repetition of ‘A’ openings, staccato short sentences, and unmistakable AI slope.
When an article gets long, the architecture completely collapses. The intended trajectory—a deliberate, linear drift from point A to point B—decays midway, a failure clearly exposed by the charts embedded in the piece. Despite an entire custom-built stack, I still end up having to rewrite the charts and the prose by hand.
Ultimately, the harness layer can regulate execution steps and workflows, but it cannot solve the prose. It cannot fix word selection, sentence cadence, vocabulary, phrasing, or any of the nuances that actually matter.
In the end, after tearing through every LLM paper I could find, I realized the core issue goes down to how models are trained, their underlying operating mechanics, and how they rely entirely on a ‘scoreboard’ to interpret reality.
You cannot understand the world through grading alone. Context matters in reality, and it operates entirely differently than a pure point-scoring system.
Take clothing: wearing a full three-piece suit is ideal for a high-end financial conference, but it’s completely absurd for a casual date. Even though a suit scores ‘higher’ in terms of formality and price compared to a simple shirt, the shirt remains the objectively superior choice for the date.
Language works precisely the same way—and so do articles.
Now let us breakdown how LLM truly learns of everything.
I. TRAINING = SCORING?
What DeepSeek-R1 was rewarded for
In January 2025, DeepSeek published the training recipe behind R1. On January 27, the release helped wipe $589 billion off Nvidia’s market value in a single session—the largest single-day loss for any stock in market history. The catalyst was simple: the world finally got to see what was actually inside the AI blackbox.
The reward mechanism behind the model was almost embarrassingly simple.
For math, the model had to place its final answer in a designated format so a deterministic rule could verify it against the correct number.
For coding, a compiler executed the output against predefined test cases. Right answers gained points; wrong answers lost them.
This is what we call ‘machine learning’—the attempt to teach a machine binary definitions of ‘right’ and ‘wrong.’ Yet that entire process is locked onto quantifiable constraints, completely blind to complex human context.
Coding is almost an ideal environment for machine learning because reality eventually gives you an answer. The program compiles or it does not. Tests pass or fail. Run the program and something observable happens.
Mathematics offers a similar advantage. Even when the reasoning path is complicated, many problems eventually collapse into an answer that can be checked. Games are better still. They come with rules, states and scores built in.
While these training methods are direct and easy to measure, they are miles away from making AI anywhere near AGI. Human intelligence relies heavily on hidden, highly contextual, cultural, and emotionally grounded knowing—dimensions that binary metrics cannot touch.
Writing a genuinely good article with cultural fluency across different languages is practically the ultimate proof of this structural gap.
Suppose I give you two versions of a function: one passes the unit tests, the other fails.
At least we have a baseline to work with.
Now suppose I give you two paragraphs.
Both are grammatically correct.
Neither contains a factual error.
They make the exact same argument, use appropriate vocabulary, and any automated writing evaluator would happily give both high scores.
Yet one belongs in the essay, and the other ruins it, sounds completely sloppy.
Explaining why can take longer than reading either paragraph.
Perhaps the second version introduces an idea too early. Perhaps the previous paragraph already did the heavy lifting, making a redundant explanation drag the reader down. Maybe the punchy sentence everyone loves is precisely the one that needs to be slaughtered because it steals momentum from the core argument.
There is no compiler waiting at the end with a green check mark.
And that is precisely why prose is the ultimate litmus test for what LLMs are actually learning—and where they fundamentally fail.
As Ilya Sutskever has pointed out, we are reaching the hard limits of the pure scaling era.
Training an AI to master human language cannot rely on a glorified grading ledger or a flat optimization scoreboard.
We need entirely new architectural methods—foundational breakthroughs that go far beyond assigning points to outputs, shifting from static scoring to true experiential understanding.
[CHART 1 — Why more rules never fixed AI prose]
II. WHAT A SCOREBOARD CANNOT SEE
The model writes without a sense of direction
In November 2025, Ilya Sutskever told Dwarkesh Patel that the field is leaving the age of scaling and returning to an age of research. Part of his case rests on something people have and models lack.
Humans carry an internal sense of whether a task is going well long before the final outcome arrives, and Sutskever suspects that emotion is part of what tunes it. Reinforcement learning, as it is practiced today, mostly waits for the outcome and then hands out the grade.
That is exactly the gap described above. While drafting, writers feel a paragraph start to drag, notice that a phrase already appeared two pages earlier, and sense when the argument has said enough. The model receives its signal only after the text exists, and for prose that signal often never arrives.
Checking the work afterward does not close the gap.
In 2024, researchers from Google DeepMind and the University of Illinois tested what happens when a model reviews its own answers with no outside signal telling it whether the first answer was right. Across four models and two benchmarks, the reviewed answer never beat the first one. GPT-4 fell from 95.5% to 89.0% on grade-school math. GPT-3.5 fell from 75.8% to 41.8% on a commonsense quiz, largely by changing answers that had been correct the first time. Llama-2 lost more than 25 points on both tests.
[CHART 2 — Accuracy before and after the model reviews itself]
Those were math and logic questions, where a correct answer exists and the model still could not find its way back to it. Prose has no answer key at all, even after the fact.
The scoreboard always picks the suit
When the reward has to come from people instead of a compiler, it inherits their habits.
In October 2025, researchers from Northeastern University and Stanford, including Christopher Manning, took 6,874 pairs of answers from the HelpSteer dataset that had received identical correctness ratings.
Raters still preferred whichever answer read as more typical, and the effect was statistically overwhelming. Train a model to win those ratings and it learns to reach for the familiar choice every time.
The paper calls the result mode collapse.
This is the three-piece suit at the casual date.
The scoreboard ranks the familiar option higher in every room, so the model keeps wearing it, whatever the occasion calls for.
The result shows up in public data.
In July 2025, Dmitry Kobak's team at the University of Tübingen published a Science Advances study of more than 15 million PubMed abstracts from 2010 to 2024, and estimated that at least 13.5% of 2024 abstracts had passed through an LLM. Their raw yearly counts show what that looks like on the page.
"Delves" took twelve years, from 2010 to 2022, to climb from about 1 to about 7 abstracts per 100,000.
By 2024 it appeared in 357, roughly 48 times its 2022 level. "Underscores" and "showcasing" rose about 14-fold over the same two years, and "meticulously" about 10-fold.
III. LLMS LOVE PULLING SSR CARDS (The Probability Illusion)
This gaming terminology perfectly describes all the AI problems above. In most mobile games, ‘SSR’ stands for the rarest and best cards, obtainable only through extremely low probabilities.
When it comes to coding, writing, or problem-solving without a verification loop, using an LLM is basically playing gacha. You pray for a high-quality outcome, but mostly end up with a mess. Most of the time, you have to inspect and fix the output with heavy manual effort—you end up cleaning all of the AI’s messes yourself.
The funny part is that you expect AI to work for you, but it doesn’t.
Hoping that a high-quality outcome will save you three hours of work is like rolling for an SSR, and I am essentially doing it every single day. For every bad piece of writing, I spend extra hours fixing it just like debugging code. Sometimes the quality is so abysmal that I feel exhausted doing this day in and day out.
The LLM’s word generation follows this exact same pattern: words emerge out of naked probability rather than true context. So even in articles and essays, the entire piece can fall apart and hallucinate. We call this the **AI Tax**—a concept I will dive into further in the future.
The words themselves are not technically wrong. That is exactly why the problem survives. Use one at the right moment and nobody notices. Use enough of them and the article develops that unmistakable synthetic smell.
Sentence structure behaves the same way. Once the model finds a rhythm that sounds forceful, it tends to return to it. One sentence establishes a pattern. The next sentence repeats the pattern with a different noun. Soon the paragraph is marching. I call this the A/A/A problem.
The bizarre part is that every individual line can look good on its own.
That is what makes the SSR analogy so useful. Twenty rare, overpowered cards in a single deck do not make a good deck. Writing needs connective tissue, changes in pace, and moments where the language simply gets out of the way.
LLMs just keep pulling SSRs and we can do nothing about this.
[CHART 2 — Every word is one pull from a weighted machine]
IV. WHERE THE SAME WORDS KEEP COMING FROM
What 15 million abstracts show
In July 2025, researchers led by Dmitry Kobak at the University of Tübingen published a study in Science Advances that tracked vocabulary across more than 15 million biomedical abstracts indexed by PubMed from 2010 to 2024. Their question was simple: did the arrival of ChatGPT leave fingerprints in how scientists write? From the excess vocabulary alone, the authors estimated that at least 13.5% of 2024 abstracts had been processed with an LLM, and the shift in word use was larger than the one Covid caused.
The team published its raw yearly counts, so the fingerprints can be checked directly. In 2022, “delves” appeared in about 7 of every 100,000 abstracts. In 2024 it appeared in 357, roughly 48 times as often.
“Underscores” and “showcasing” rose about 14-fold over the same two years, and “meticulously” about 10-fold. Before ChatGPT, the same words moved slowly: “delves” took twelve years, from 2010 to 2022, to go from about 1 to about 7 per 100,000 abstracts. It then added another 350 in two.
[CHART 4 — Four style words, indexed to their 2022 level]
At 13.5% of the 1.44 million abstracts in the study’s 2024 sample, that is at least 190,000 papers in a single year. Their authors studied different diseases in different labs, and they reached for the same handful of words in the same two years. What they shared was the tool.
In order to solve this problem, can we solve by simply banning words?
This bothered me enough that I started explicitly banning words.
It works for a while until it’s not.
Take away one favorite term and the model finds another. Tell it to stop inventing unnecessary concepts and the language becomes cleaner, until a different family of abstractions begins creeping in. Ask for more vocabulary and the result can become even worse because the model interprets variety as a request for rarer words.
That is when I realized I had been treating the symptom as a vocabulary problem.
Why the word output eventually converges
What happens is closer to convergence. Once the model recognizes the genre it is churning out, certain words become irresistible hooks. Macro essays start sounding like every other macro essay. Technical writing drifts straight into the familiar vocabulary of systems, layers, and frameworks. The subject shifts, but the prose stays frozen in place.
Part of the reason sits in how models are tuned after pretraining. In October 2025, researchers from Northeastern University and Stanford, including Christopher Manning, traced the sameness of aligned models right back to the human preference data used to train them.
They took 6,874 pairs of answers from the HelpSteer dataset where both answers had the exact same correctness rating, meaning neither was more accurate than the other. Raters still preferred whichever answer read as more typical, and the effect was statistically overwhelming. Train a model to win those ratings, and it quickly learns to stay dead in the middle of the road. The paper calls the result mode collapse.
The variety is still inside the model.
When that same team simply asked models to list several possible answers with a probability assigned to each, creative-writing output became 1.6 to 2.1 times more diverse. The model still holds the whole range.
Training just taught it which end of the range gets the treats.
Look, human writers fall into this exact same trap. We pick up weird habits, steal rhythms from people we binge-read, and overuse words we happen to like.
The difference is that a real editor eventually gets sick of AI Slope and more bullshit.
Eventually all the AI power users will stop using AI for all the slopes& garbages AI model has generated.
V. I TRIED SOLVING THIS WITH MORE RULES
What happens at 500 AI instructions
In July 2025, researchers at Distyl AI tested how well models follow many instructions at once.
Their benchmark, IFScale, gives a model a business-report writing task and up to 500 simple instructions, each requiring a specific keyword to appear. Twenty frontier models from seven providers were tested.
At 500 instructions, even the best one, a preview version of Gemini 2.5 Pro, followed only 68.9% of them. GPT-4.1 managed 48.9% and Claude Opus 4 managed 44.6%.
The strongest models held near-perfect scores up to around 100 to 150 instructions and then slid, while weaker ones collapsed early: GPT-4o was already down to 49% at 100 instructions.
Across models, rules placed earlier in the prompt were followed more often than rules placed later.
[CHART 5— Instructions followed as the rulebook grows]
Keyword inclusion is the easiest rule to add in the md files and rule books.
But speaking of Rules about article prose, rhythm, restraint, and pacing?
There’s no way you can let LLM handle them through google search and other tools.
I already build software with an unusually heavy control system around AI agents. Important behavior is locked down in contracts. Changes have to pass gates. Invariants are explicit. State changes require hard evidence. This approach actually cut down engineering accidents because you can eventually turn messy software bugs into something deterministic.
Writing dragged me straight into that exact same trap.
Do not use repetitive parallel constructions.
Stop reaching for the same abstract vocabulary.
Avoid unnecessary headings.
Do not manufacture a concept when ordinary language works. Watch paragraph rhythm.
Stop summarizing points that have already landed.
The size rulebook just kept growing.
The early gains were real. Five or ten carefully chosen instructions can tweak an output pretty well. But once you pile on too many rules, every extra page gives you diminishing returns until it does jack shit.
When the rules go over 500 rules, the entire rule book falls apart entirely.
VI. THE INDUSTRY’S MEASUREMENT PROBLEM
Why DeepSeek refused to use a judge
Go back to the R1 paper for one more detail. DeepSeek deliberately chose not to use neural reward models, the learned judges that score answers the way a human rater might, for its reasoning tasks. Their stated reason was that models learn to game those judges during large-scale training. The industry calls this reward hacking.
That decision tells you where the frontier labs trust their own signals. Where a rule can check the answer, they scale. Where only a learned imitation of human taste can check it, they get nervous, because the model learns to please the judge instead of doing the task.
Prose has only the second kind of signal.
[CHART 6 — Where a rule can grade the answer, and where it cannot]
What the charts measure
Seen from this angle, the current AI boom becomes easier to understand.
We are making extraordinary progress where training can receive a clean signal. Naturally those capabilities improve fastest. Researchers pursuing them are being perfectly rational, exploiting the feedback loops available to them.
The trouble begins when progress in those environments becomes evidence for a much broader claim about intelligence.
Every benchmark gives us a number.
Numbers produce curves. Curves can be compared across model generations, and a rising curve is wonderfully easy to put on a slide, most importantly, CAPITAL REWARDS NUMBER.
NUMBERS ≠ JUDGEMENT
Nobody can cheaply generate ten million examples of great editorial taste and attach an unquestionably correct reward to every decision. The interesting part of an edit may be the explanation of why it works, and producing that explanation can itself require the skill we are trying to train.
This creates a nasty asymmetry. The parts of intelligence that are easiest to measure are also the easiest to optimize at scale. Once they improve quickly enough, it becomes tempting to treat them as proxies for the whole thing.
That is where much of the AI industry’s strange, occasionally fraudulent atmosphere comes from.
The models really are improving.
The leap from “the measured capabilities are improving” to “we are rapidly solving intelligence” is doing far more work than the charts admit.
Prose happens to expose the gap because the failure is difficult to disguise. Models can know every fact required for an essay, understand the argument, follow a detailed style guide and still produce something a good writer would never publish.
No benchmark number can make the paragraph feel alive.
VII. Now let’s talk about DARIO
So I don’t find it strange at all that Dario Amodei uses Claude heavily to grind through research and still writes his own essays.
Research is precisely where I want a hyper-powered machine sitting in the passenger seat. I want it reading, comparing, digging up buried facts, and taking a sledgehammer to the weak spots in my arguments. I use AI for that every single day.
Then comes the blank page.
Suddenly you’re staring at thousands of pristine, beautifully structured sentences. The real challenge is figuring out whether the essay actually needs a single one of them.
That is a completely different kind of intelligence test.
Let’s stack the whole mess together. Models sprint ahead wherever a hard compiler can verify the answer—which is how R1-Zero went from 15.6% to 77.9% on AIME. Prose doesn’t have a compiler. The closest substitute we’ve got is human preference, and humans naturally gravitate toward the familiar even when two answers are equally correct. That’s why AI writing keeps pulling the same recycled SSR cards, and why buzzwords like ‘delves’ magically exploded across medical abstracts.
Piling on more rules hits a brick wall. Hand a top model hundreds of formatting rules and it will still slip up. And even when the model can recite the rule back to you word for word, it can’t reliably catch itself violating it. Take away the external guardrails and every model scores worse trying to fix its own mess than it did on the first try.
Claude can help Dario dig up data. Dario still decides what sounds like Dario, what earns its keep on the page, and what gets thrown straight into the trash.
For all the marketing hype about how much modern models know, I’d bet everything the next real breakthrough won’t come from another bloated benchmark.








