My whole thing is being honest about AI. That means when something is impressive, I say so. And when something is overhyped or genuinely broken, I say that too.
This is the second kind of post.
There’s a lot of hype about AI and not enough sober accounting of what it can’t actually do. These are seven things AI consistently struggles with, even the best models — and my best guess on when that changes.
1. Reliably Telling You When It Doesn’t Know
This is the one that keeps me up at night. The best AI models are trained to be helpful and fluent. That combination sometimes produces confident, fluent wrong answers. The model doesn’t always know that it doesn’t know — it just… continues.
This is called hallucination, and it hasn’t been solved. It’s gotten much better, but it hasn’t been solved.
My timeline: Models are getting better at expressing uncertainty. I’d guess 18-24 months before the best models are reliably calibrated about their own knowledge limits. Until then: verify anything factual before you publish it.
2. Maintaining Consistent Identity Across a Long Project
Ask an AI to write Chapter 1 of a novel, then come back a week later and ask for Chapter 5. The character’s name might be the same, but the voice, the details, the small choices that make a character feel real — those often drift.
AI is excellent at the local level (this scene, this paragraph) and weak at the global level (this whole project, these characters’ arcs).
My timeline: Better long-context memory and persistent state are being actively worked on. This is a 12-24 month problem, not a 5-year problem.
3. Understanding Causality, Not Just Correlation
AI is extraordinarily good at pattern recognition. It’s genuinely bad at causal reasoning — at understanding why things happen, not just what tends to follow what.
This matters for business decisions, medical advice, legal analysis — anything where the cause-and-effect chain matters. AI can tell you what’s correlated. It can’t reliably tell you what’s actually causal.
My timeline: This is a harder, more fundamental problem. Probably 3-5 years before we see meaningful progress.
4. Creative Work That’s Genuinely Surprising
AI is very good at producing work that is competent. It’s mediocre at producing work that is surprising.
The best creative work — the kind that makes you stop and think, or feel something unexpected — comes from an unusual perspective. AI optimizes toward the center of what’s worked before. That’s useful, but it’s not the same as creativity in the fullest sense.
My timeline: I think this is less a technical problem and more a philosophical one. The question of whether AI can be genuinely creative is going to be debated for a long time.
5. Keeping Up With What Happened This Week
Most AI models have a knowledge cutoff. Even with web search tools, there’s friction in getting current information into the context and using it accurately.
For me, this matters every week — the AI space moves fast, and models that can’t track current events miss the context that makes information useful.
My timeline: This is being actively solved with better retrieval systems and real-time search. 6-12 months to meaningfully improve.
6. Knowing the Difference Between What You Wrote and What You Should Have Written
AI feedback on writing is useful for catching obvious problems — typos, passive voice, logical gaps. It’s not good at telling you whether the piece works — whether it’s interesting, surprising, worth reading.
Put a mediocre piece in front of an AI and it will often confirm that it’s fine. Put a great piece in front of it and it will often suggest changes that make it more generic.
My timeline: Genuinely hard. Aesthetic judgment is one of the last things to fall. I’d put this at 3-5 years for significant improvement.
7. Physical and Embodied Tasks (For Most Use Cases)
Robot technology is advancing, but the gap between AI that can think and AI that can reliably act in the physical world is enormous. The precision, adaptability, and contextual awareness required to do physical tasks well is a different problem than language modeling.
My timeline: Specific industrial use cases in the next 2-3 years. General-purpose physical AI is a 10+ year problem in my assessment.
Why This Matters
I’m not writing this to be a skeptic. I’m writing it because the people making the best use of AI are the ones with a clear-eyed view of what it can and can’t do. They use it where it’s strong. They don’t trust it where it’s weak.
The worst AI users are the ones who trust it for everything and the ones who dismiss it as overhyped. The best users are in the middle — curious, rigorous, honest.
That’s what I’m trying to be.
— Deco