Aníbal Carmona
Back to writing
Artificial Intelligence & Judgment

Hallucinating with AI Industrializes the House of Cards

September 20266 min readPublished

Vision without execution used to be hallucination. Hallucination with execution is blindness.

Hallucinating with AI Industrializes the House of Cards

Abstract

For decades, execution was proof of knowledge because it was proof of effort. Language models break that correlation: today an idea can be executed without ever being understood. Drawing on three products built in a single month at the Kairosia lab — all functional, zero customers — the essay argues that when producing becomes cheap, what becomes expensive is judgment: asking a good question, spotting a false premise and going back to talk to reality.

A founder pitches his company to a panel of investors. He has a market study, a five-year financial model, a working product and a deck any consulting firm would sign. He produced it in eleven days, almost alone. The investors ask how many customers are paying. They are about to. Never has an unvalidated idea looked so validated.

For decades, management repeated that vision without execution is hallucination. It worked because imagining was free and executing was enormously expensive: whoever managed to get from one to the other proved, by the mere fact of having done it, that they knew something. Execution was proof of knowledge because it was proof of effort.

That correlation has broken. Language models let you research a sector, design a product, write the code and build the financial model in an afternoon. It is an extraordinary advance. But it voids the old aphorism and suggests another, less comfortable one: hallucination with execution is blindness. Today an idea can be executed without ever being understood.

The accent of the wise

Ignorance has always existed, but it used to be visible. Someone who could not build a financial model produced a spreadsheet any analyst could take apart in two questions. That friction forced you to study, to hire someone who knew, or to say “I don’t know.” AI removes the friction and, with it, the signal.

The models learned the language humans use to express knowledge: prudence, nuance, the anticipated objection. A generated answer sounds like a consultant with twenty years of scars. But the language of wisdom is not wisdom. There is an enormous distance between knowing how someone who knows speaks and having walked the road by which they came to know — a road made of customers who left without explaining why and perfect products nobody used. You can hack the language of wisdom without acquiring a drop of it, and since we all judge by language, the hack works.

There is an aggravating factor: the questions we ask the machine contain our hypotheses. Whoever asks “how do I monetize this” has already decided there is something to monetize. The idea acquires a strategy; the strategy, metrics; the metrics, a prototype. Everything confirms the idea was good, until someone who took no part in those conversations shows up: the customer.

Houses of cards at industrial scale

No one is immunized by experience. In a single month, in our own Kairosia lab in Barcelona and after three decades of building software, we stood up three products with the help of language models. A judgment tool for business angels: discarded after the first serious analysis, carried out, with some irony, by the very AI meant to power it. A DORA compliance expert system no consultant wanted to adopt. A compliance engine for short-term rentals, technically flawless, commercially inert.

Three functional products, zero customers. The technology did not fail. What failed was the willingness to pay that sat in row 14 of the model and never existed outside it.

For decades execution worked as a filter: it was so expensive that it killed bad ideas before they got far. That filter is dissolving. The cost of executing a good idea falls, and so does the cost of executing a bad one. A bad idea used to die in a PowerPoint. Now it reaches production, with a registered domain and a pre-seed round.

The counterargument deserves fair treatment, and our failures illustrate it as well as they do the thesis. The old filter was crude: it killed as many good ideas as bad ones and favored whoever had capital over whoever was right. My three products cost a month, not three funding rounds. Cheap execution also makes contact with reality cheap — provided you allow it. The same tool shortens the path to the market’s verdict or postpones it with one more layer of sophistication. The difference lies in who is holding it.

The resource that gets expensive

Speed does not correct course; it amplifies it. The question is not how much AI accelerates execution, but what happens when you accelerate something that should never have been built.

When a resource becomes abundant, its complement becomes expensive. If producing gets cheap, what rises in price is judgment: asking a good question, spotting a false premise, changing your mind when the evidence contradicts what you wanted to believe. And, above all, stepping out of the conversation with the machine and going back to converse with reality.

One of our three products was killed not by the market but by the AI itself, when it was asked, with no anesthesia, whether the idea made sense. It said no, and it was right. The machine knows how to kill bad ideas. It only does so when asked, and we almost always prefer not to ask.

The greatest risk of artificial intelligence is not that it hallucinates. It is that we stop knowing when we are hallucinating along with it. Vision without execution used to be hallucination. It is time to complete the sentence: hallucination with execution is blindness. And the easier it becomes to execute, the harder it will be to pretend we know where to look.

Notes

  1. 1.Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R. and Wilson, N. (2025). “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers”. CHI ’25, ACM. https://doi.org/10.1145/3706598.3713778 — Survey of 319 knowledge workers (Microsoft Research and Carnegie Mellon): the greater the confidence in generative AI, the less critical thinking; the greater the confidence in oneself, the more.
  2. 2.Sharma, M. et al. (2023). “Towards Understanding Sycophancy in Language Models”. Anthropic; ICLR 2024. arXiv:2310.13548 — Evidence that models trained on human preferences tend to adjust their answers to the user’s beliefs: the technical basis of the echo chamber.
  3. 3.Bender, E. M., Gebru, T., McMillan-Major, A. and Shmitchell, S. (2021). “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”. FAccT ’21, ACM. https://dl.acm.org/doi/10.1145/3442188.3445922 — The distinction between producing fluent language and understanding what is said; the academic origin of the idea that form does not guarantee substance.
  4. 4.Ries, E. (2011). The Lean Startup. Crown Business — Introduces the concept of “achieving failure”: successfully executing a plan that leads nowhere. Execution without validated learning was already the problem before AI; AI makes it cheaper.

Aníbal Carmona · Barcelona · September 2026 · Founder, Kairosia