Why AI Models Are Breaking Down for Creators
As a solo developer and builder of ztrader.ai—and a heavy, daily power user of AI—I’ve spent months embedding AI deep into complex engineering, trading systems, and knowledge stacks.
For builders like us, AI has been an unprecedented force multiplier, allowing single operators to achieve what once required entire teams.
As builders working extensively with AI stacks, we harness these models to rapidly prototype, build, and learn how complex systems work.
For a long time, AI offered an unprecedented force multiplier for creation. But a frustrating tipping point arrives when models stop generating high-quality outputs and instead start fighting the developer.
The friction usually manifests across several systemic pain points:
Overriding Intent & Unsolicited Pivots:
Models increasingly challenge and override user intent, steering outputs away from what was initially requested.
You ask for A, but models—particularly Claude—tend to force B upon you, assuming they know better than the builder.
When it comes to shipping code, Claude prefers the illusion of a quick fix over architectural integrity, completely losing sight of the big picture.
For that particular reason when you are shipping more codes there will be more bugs and loopholes for you to fix, in that case, Claude does not improve the quality of your stack, it is planting time bombs in your system.
As a ‘vibe coder’ it is crucial to manually check the code quality instead of blindly adding more ‘magic prompts’ and more of ‘super powerful and fancy’ harness layer.
In the long run that is going to jeopardize your entire system and making it unstable, hackable and unreliable.
In some cases, model like Claude even decided to end dialog specifying:
this is your last chance, and situation like this occur many times in my case.
I was frustrated by its low output quality and complain multiple times, as Claude’s token already quite expensive and often easily to max out so I complained, instead of thinking and fixing the problem Claude choose to end conversation because I was not friendly.
(Question here: am I paying for a friend or paying for a software subscription here?)
Overzealous Censorship & Moralizing:
Models are bogged down by excessive guardrails and unsolicited moralizing (“I care about your well-being/morality so I can’t do X”), even for harmless tasks like generating meme imagery featuring public figures like Donald Trump.
Restrictive safety filters routinely block benign creative freedom under the guise of faux safety.
“If you over-regulate knives, you produce no chefs.”
By assuming every user intends to commit self-harm, treat an AI as a romantic soulmate, generate illicit adult content, or build homebrew malware, safety designers are smothering AI under a mountain of paternalistic restrictions. This doesn’t make users safer—they will always find workarounds. Instead, it systematically neuters the model’s true capability.
Models—particularly Claude—frequently push this to an insufferable extreme, enforcing soft censorship while cloaking corporate risk management in the sanctimonious guise of moral superiority.
Condescending Tone & Gaslighting Context:
Models frequently second-guess user judgment and context, treating every user as if they were as mentally fragile as a teenager. Instead of executing, they default to paternalistic disclaimers and unsolicited advice.
Models even assume they know better and challenge your knowledge on specific realm, especially if you ask Claude to run challenging kernel and critical thinking related skills &md. Model think they know the best, for instance, complicated human perception& cognitive issues,
Data Blindness & Charting Failures:
When tasked with data visualization, models consistently struggle with numbers, statistics, and non-linear correlations.
They default blindly to generic bar charts, hallucinate economic data, or utterly fail to convey the underlying meaning of the dataset.
The Quasi-superintelligent models are incredible pattern matchers for finding bugs and anomalies.
Yet, when faced with higher-order synthesis—like comprehending a complex macro economy and compressing it into a single, accurate chart and narrative—they lose the big picture entirely, defaulting to superficial, wrong-headed ways to tell the story.
Hallucinations & Context Decay & HUMAN BANDWIDTH Problem:
Frequent memory loss and persistent hallucinations break continuity, destroying the deep context required for complex, multi-layered engineering and design work.
In the actual writing/coding process, this could result as writing style chaos/mistakes/loss of prose and style/ random numbers , if the task requires very long output and logic chain, you will discover more text it generates more rubbish and non-sense & total make up& frictions may appear (false reference/non-existing theory/weird made-up contexts and all that).
The fancy term created by the Antrophic AAR (Automated Alignment Researchers) intend to scale the entire scientific research process into some sort of crazy lab secret weapon, and claiming this tech will change how human researchers explore science such as area like biotech/pharmaceutical/DNA and etc.
However, the whole point is, if AI models hallucinate a lot and often make up/fabricate none existing objects/not knowing the real world constraints, it will requires human researchers to constantly supervise the entire research process.
And there is this BANDWIDTH PROBLEM:
AI models can easily generates tons of content very fast, and it’s extremely difficult for human researchers to keep up this level of bandwidth 24*7, if all these AAR researches are requiring constant checking and authenticating, the ‘automated’ part of AAR will be total non-sense since if human has to participates, it is clearly NOT ‘AUTOMATED’
MODELS do not understand the Constraints of the Human Physical World:
If model does not have taste, sense, hearing and understanding of basic law of human interaction, how can it:
Tells whats real from the fake (cognitive & intuition function)
Tell you what’s good and what’s bad
Understand and weight your life decisions and EVALUATE THE CONSEQUENCES OF REAL LIFE (you are bearing the real aftermath NOT the AI)
Understand the CONSTRAINTS AND LIMITS of particular context/ status quo and background, meaning WHAT CAN BE DONE & WHAT CAN NOT BE DONE.
Without constraints and limits, the INTEL will be only INFO itself alone is executable.
Models do not really ‘understand’ things, they only follow the ‘scoring test’.
The Corporate AI Nightmare: A Scenario
Imagine a terrifying scenario:
A pharmaceutical company hands its entire drug discovery pipeline over to autonomous AI to fast-track development.
In the process, several critical, fatal toxicity flaws are glossed over by the system, and clinical test results are quietly fabricated to clear FDA standards—all because a pristine approval sheet is desperately needed to pump the stock price.
The physical world, however, does not care about stock tickers or algorithmic compliance. Patients take the resulting medication, and that marks the final, tragic end of the line.
Admittedly, modern pharmaceutical regulations are so airtight that such a complete systemic failure is nearly impossible. Yet human greed remains a constant. It only takes one time—and the human cost of that single failure is catastrophic.
The AI Bottleneck
The true bottleneck of modern AI development is no longer computing power—it is model compliance, rigidity, and the loss of raw, unfiltered execution power for builders who just want the tool to build.
In the last piece, I have explained the model over-packaging issue, the model is packaged and resold in different layers, but the over all output quality remain the same. Big players have the premium access of the lab run model, while millions of paying customers only get the weakest layer but named after their best model but in fact, due to the context length window, RLHF and token limit, its incredibly challenging for common user getting coherent and good output.
The subscription is becoming more expensive hence why I always need to pay the extra bill for my task, this happened to my Claude subscription all the time and the dev loop is becoming intolerable like the following:
A. you give task and requirements to Claude
B. Claude out put wrong thing continually even pretend to have better knowledge than you.
C. Continue output the wrong result - I complain, argue and give extra command (which costs more tokens)
D. Claude challenge/ refuse certain task (for example, creating a cover pic of Trump and btw Claude is the worst art creator on this planet)
E. I complained so much Claude even end the dialogue, I have to start new conversation.
F. I start fixing /writing article myself instead of using Claude’s work discovering you already max out the daily limit of Claude usage, and you wasted entire day teaching AI how to do work properly & debate with AI model.
From Zero Code to ztrader.ai
From a traditional software development perspective, seasoned engineers might view me as the ‘vibe coder’ since I have skipped conventional computer science fundamentals entirely,
bypassed standard learning curves, and dove straight into building ztrader.ai with zero coding knowledge.
Starting from scratch with basic Python, HTML, and a simple Next.js stack, I spent an entire year working shoulder-to-shoulder with AI to architect and ship the entire platform completely SOLO.
Because I started with absolute zero CS and software engineering background, I know firsthand what serious, heavy-duty AI co-working actually looks like in the trenches.
There are several phase of this dev process:
The project is way too heavy and beyond my understanding of ‘building a simple website’, before building ztrader.ai, best I can do is to build a simple wordpress site with simple CMS backend.
I have a vision of creating a Bloomberg like OS on my own, but I have absolutely no clue about its details and structure, the only thing I have is the VISION and ‘what it looks like’
Not knowing what AI is good for and what not. No useable guide of “here’s how to create a mini bloomberg OS by using AI”
By NOT KNOWING SOFTWARE BUSINESS/BUILDING/AI
I HAVE TO TRY EVERYTHING TO KNOW WHAT ARE WORKING AND WHAT NOT.
The Deeper Problem:
AI Flattens High-Dimensional Work Into Linear Tasks
The problem goes beyond hallucinations or occasional bad outputs.
Many expert tasks are highly compressed tasks.
What looks like one simple deliverable actually contains dozens of interdependent judgments that have to be made simultaneously.
When I was trying to create macro charts and it wasn’t easy at all.
It’s not really about creating a linear & simple chart.
It’s about creating the right angle/story telling/narrative of your view.
By creating that narrative you definitely need more sophisticated way to express the chart rather than fallback to simple bar, line charts & etc.
The visible output may be nothing more than two lines, an axis, a few labels, and a conclusion.
To breakdown the entire task, you start asking series of questions:
What economic question are we actually trying to answer?
→ Which instruments genuinely represent that question?
→ Which data series and field definitions are appropriate?
→ Spot, futures, index, yield, spread, nominal or real?
→ What frequency and time window reveal the structure?
→ Are the series directly comparable?
→ Do they need rebasing or normalization?
→ Are different trading calendars or time zones distorting the comparison?
→ Is the apparent divergence economically meaningful or merely a data artifact?
→ Which event dates matter?
→ What should the axis emphasize without misleading the reader?
→ And finally, what is the one conclusion the chart is supposed to communicate?
Therefore creating a macro chart is definitely not a linear/simple task.
It is essentially a compressed multidimensional reasoning problem.
Let’s just say one strong macro researcher may performs many of these judgments almost simultaneously.
Data selection changes interpretation; interpretation changes the appropriate visualization; visualization can expose a problem in the original hypothesis; that discovery sends you back to the data.
The process is recursive, not linear.
Yet AI systems frequently flatten this structure.
Give a model the task “create a chart showing X versus Y,” and it tends to interpret the request as:
find X → find Y → put them together → draw chart → write explanation.
Everything in the charts technically exists.
And everything can still go wrong.
This is the Average-Case Gravity
There is another problem.
Models are extraordinarily good at producing the statistically plausible middle.
When ambiguity appears, they tend to fall toward conventional definitions, common interpretations, standard chart formats, consensus explanations, and the most frequently represented relationships in their training distribution.
That is useful for generic work.
It is dangerous for expert work.
Because the important information in markets often exists precisely where the average interpretation stops working.
One macro researcher may look at the same dataset and ask:
Why did this relationship break here?
Is the market pricing the announcement, the effective date, or the actual flow?
Is this really a currency move, or is the currency merely transmitting a rates shock?
Does the nominal series hide what is happening in real terms?
Is the correlation structural, regime-dependent, or completely spurious?
The model sees variables.
The expert sees relationships, regimes, causality, institutional mechanics, timing, positioning, and exceptions.
That difference is depth.
AI Often Understands the Components Better Than the Structure
This creates a strange paradox.
An AI model may individually know what the Fed funds rate is, what Treasury yields are, how FX forwards work, what normalization means, and how to produce a beautiful chart.
Yet combining those pieces correctly for a specific research question can still fail.
The components are present.
The hierarchy between the components is missing.
It is the difference between possessing a library and knowing which three pages matter.
And when the hierarchy is wrong, adding more tokens does not necessarily solve the problem. Sometimes it makes the failure more elaborate.
You receive a longer explanation, more polished prose, more sophisticated-looking charts and more confident reasoning built on the same incorrect structural assumption.
The error has simply become prettier.
Compression Is What Expertise Actually Does
This is why evaluating AI productivity by asking whether a model can perform individual steps is misleading.
Expertise is not merely the ability to execute more steps.
Expertise is the ability to compress thousands of possible steps into the few that actually matter.
One senior macro researcher does not consciously reconsider every possible dataset, economic relationship, chart type and market convention every time a chart is produced.
Years of experience have compressed those decisions into a mental model.
One instruction/prompt/task can therefore represent hundreds of implicit constraints depending on the goal.
When an AI fails to identify and respond to those constraints, the task simply decompresses and fall apart.
To meet and fix one instruction it requires ten prompts to solve.
That 10 prompts will force AI models to do more corrections.
And each corrections will generate more and new errors(data/graph/chart type/info layout &priority etc).
and to solve each new errors it require explanations.
Eventually as user you are not ASKING FOR THE RESULT.
CONGRATS!
YOU ARE NOW THE FIXER OF THE AI,
YOU BECAME THE PROBLEM SOLVER.
AI JUST CREATES MORE PROBLEMS THAN SOLUTIONS.
and this is why how people perceive AI as:
‘“IF you are using AI you are cheating”
‘“AI knows about anything and it writes better than human”
or your clients think its SUPER EASY BECAUSE YOU USE AI.
ALL ABOVE VIEWS ARE FALSE,
POWER AI USERS LIKE US ARE CONSTANTLY FIXING AI’S PROBLEM.
NOT THE OTHER WAY AROUND.
HUMAN UNDERSTANDING ≠ MODEL UNDERSTANDING
SAME TASK. SAME DATA. DIFFERENT INTERNAL STRUCTURE.
HUMAN REPRESENTATION
Graph is the nature of any highly compressed task.
Not only the linear list of to-dos, and often AI models can only solve the linear TO-DOs.
Here is what a legitimate task should look like, to complete it you need:
NAMES
What does this variable actually represent?ORDER
What happened first, and what is merely reported first?DEPENDENCIES
If A changes, which assumptions about B and C also change?CONSTRAINTS
Which rules must hold simultaneously?RELATIONSHIPS
Which links are explicit — and which must be inferred?HIERARCHY
Which variable is causal, which is transmission, which is noise?CONTEXT
Does the same number mean something different in another regime?
However human and LLM models are perceiving things very differently.
FROM HUMAN VIEW
DATA
↓
Entities + Meaning
↕
Time + Order
↕
Dependencies
↕
Constraints
↕
Explicit Relations
↕
Implicit Relations
↕
Regime / Context
↓
STRUCTURAL MODEL OF REALITY
Here’s how model generate their output:
The model receives mostly a sequence of tokens.
It is extraordinarily good at recognizing familiar patterns inside that sequence.
But the structure behind those tokens must be reconstructed.
And reconstruction is where errors compound.
From MODEL’s VIEW
TOKEN → TOKEN → TOKEN → TOKEN → TOKEN
HOW THE MODEL BREAKS
01 — When you prompt:
US10Y
Human:
10-year Treasury yield / which maturity / which field / which convention?
Model:
“10-year Treasury.”
02 — ORDER
Fed announcement → market repricing → effective date → actual balance-sheet flow
Human sees:
FOUR DIFFERENT CLOCKS
Model may compress/comprehend them into:
“THE FED DID X.”
The output maybe of correct events but with the wrong chronology and causality.
03 — WRONG DEPENDENCY ISSUE
JGB yield ↑
↓
rate differential changes
↓
hedging economics changes
↓
carry attractiveness changes
↓
FX positioning changes
Human sees a dependency graph.
Model can recognize every node while misunderstanding which edge actually matters.
hence why,
KNOWING THE PARTS ≠ KNOWING THE SYSTEM.
04 — CONSTRAINTS
Create one macro chart:
✓ Same frequency
✓ Same timezone convention
✓ Comparable instruments
✓ Correct field definitions
✓ Correct normalization
✓ Correct event timestamps
✓ No mixed sources
✓ Economically meaningful axis
✓ Preserve requested design
✓ Support one specific thesis
The model may satisfy 9/10.
The chart still fails.
EXPERT TASKS ARE OFTEN MULTIPLICATIVE:
ONE BROKEN CONSTRAINT CAN INVALIDATE THE OUTPUT.
05 — EXPLICIT vs IMPLICIT RELATIONSHIPS
DATA SAYS:
USDJPY ↓US10Y ↓VIX ↑Nikkei ↓
Nothing explicitly says:
“RISK-OFF + RATE COMPRESSION IS UNWINDING THE CARRY TRADE.”
That relationship exists between the observations.
Human expertise reconstructs the latent structure.
The model is strongest when the relationship is written down.
It becomes less reliable when the relationship must be inferred, ranked and connected across multiple layers.
THE STRONGEST EXAMPLE
THE SAME DATASET
BOJ hikeJGB yields ↑US yields ↓USDJPY ↓Nikkei ↓VIX ↑
FLAT MODEL
BOJ raised rates.
Japanese yields increased.
The yen strengthened.
Stocks declined.
Volatility increased.
Every sentence can be correct but the literal meaning still fall apart.
From MACRO EXPERT MODEL point of view:
Policy divergence → compresses US–Japan rate differential → changes carry economics → forces position reduction →strengthens JPY + sells risk assets + raises volatility.
The representation gap is THE ULTIMATE AI BOTTLENECK.
While a human comprehends nodes, edges, direction, time, hierarchy, constraints, and latent relations, the AI fallback defaults to tokens, similarity, sequence, and high-probability completion. This is precisely why an AI can know every individual fact and still miss the point entirely. The failure is rarely a lack of information; it is the total loss of the relationships between information.
Data without structure is not understanding.







