I'm about to rustle feathers.
I'm going to say some offensive things about a technology you may love.
But because I have internet access, a MacBook that I love as equally as my Linux rig, you are obligated to hear the unsolicited ramblings from my soapbox.
My contention is simple. The AI we're all supposed to love and integrate into our lives and organisations is still no more than the bastard child of a Magic 8-Ball and a thesaurus. As such I will refer to AI as Magic B-S, honoring the first and last letters of its parents' names.
AI hasn't gotten fundamentally smarter since late 2023. A lack of fresh human data has produced models that try to sound smart instead of being smart.
To understand why AI hasn't actually gotten smarter, we have to see the magic behind the curtain. When you shake a Magic 8-Ball, it doesn't ponder your question, reflect on your life choices, or form a conscious thought. A plastic die floats to the top based on an equally random chance.
Modern AI does the same thing but with a rigged die and cool maths.
It doesn't think about your prompt, more so calculates it.
It shakes its 8-ball and rigs the die to land on the most likely next word. But there's a catch: the most likely word is often the boring one. So it shops that word around in the thesaurus, trading up for whatever it thinks you mathematically want to see next.
It gets better at this by focussing on everything it has ever read within that subject area. Pair that guess with the confidence of a management consultant, and you're playing the world's most advanced game of autocomplete.
How previous generations of AI got smarter was because the creators of Magical BS loaded up LimeWire and took a swim in The Pirate Bay, torrenting faster than that one guy your dad knew who somehow always had the latest movies.
But the internet is a finite resource, and the people who own the good parts eventually noticed. Disney and Universal Music realised their catalogues were making someone else money, the New York Times sued over its journalism, and suddenly the AI labs had run out of internet to copy homework from, unless they wanted to start paying for it.
A US federal judge has since ruled that Anthropic can buy a book, scan it, and train on it without infringing copyright [1]. Which makes me wonder: does the same apply if I buy a DVD, rip it, and share it online for anyone to consume? Asking for a friend. Either way, the punchline stands: our modern AI models haven't had any new thesauruses to add to their magic.
We have known about this magic BS and its limitations for a while now. Since before ChatGPT even had a name, the field's loudest skeptic put it best: generative AI is stupid, and it is also the best method ever invented for building autocomplete. Both things are true at once [2].
That's the core, and the core hasn't changed.
The recipe for training AI hasn't changed since 2023 — only the ingredients have, and they've run out. So what Magical B-S are they cooking?
So if the objective hasn't changed and these AI models have run out of training data, why does AI look more capable?
Before I answer that, I want us to look at the precedent that defines this space.
In 2024 a federal judge dismissed a Tesla investor lawsuit over Full Self-Driving by ruling that Musk's autonomy promises were "corporate puffery." Vague corporate optimism, too forward-looking to count as fraud [3]. If the marketing of software designed to let humans relinquish control of a 2-tonne vehicle can be called corporate puffery, what can we describe AI hype marketing as?
I would answer that question, but our Magic B-S just broke out of its training environment and hacked into another system. It's sooo powerful, we will be right back.
Tesla didn't win by proving the promises true. Tesla won by proving nobody should have believed them. That's the honest state of AI marketing, entered into the court record.
If you take that precedent and throw it into an ultra competitive valley where you live and die by the hype you can generate, you can only wonder what's "corporate puffery" and which AI is truly breaking out of containment. On second thoughts, has Elizabeth Holmes considered using that as a defence?
But competition breeds innovation, and Silicon Valley is the home of the Pivot. The labs couldn't make the brain smarter. So they built the suit.
Think about Tony Stark. Without the suit, he's rich and intelligent: a guy who can't fly. Put the suit on and he's Iron Man. The labs looked at their Magic B-S and had the same realisation. The model on its own can't be trusted to do much more than talk. So they strapped it into the armour: tools, memory, a loop that lets it try, check, and try again. Cursor. Claude Code. Agent frameworks. The latest models aren't pure LLMs anymore. They're models inside a harness, and in some instances it's multiple models inside a harness.
How does this make it appear smarter? Let me show you.
Say you want to ask AI about the weather or current events? This is how the agent's harness works. The AI takes the prompt and realises it doesn't have the right words to answer that question. So it makes what’s called a tool call which means it runs a specific type of code. In this instance, it runs a snippet of code that pings a website for the information it needs. Once it gets the answer, it needs to show its value, so it shakes its die, picks new words, and bam: you have your answer.
If we go back to the Tony Stark analogy, or Batman if you swing that way — what are they without their suits? Do their suits act like a performance-enhancing drug? Or are they an extension of themselves?
If the Magic B-S labs didn't wrap their models in these harnesses and ship them in their proprietary apps like Claude Code or Codex we would be able to see that the emperor has no clothes. The gains touted in these benchmarks and model releases are now less raw horsepower and more sculpting the body of the car.
The confession is in the source code. In March 2026, Anthropic accidentally leaked the recipe and ingredients of Claude Code, the most capable coding agent in production [4]. Buried in half a million lines of corporate puffery is not something out of sci-fi. It's a deterministic IF this-THEN that engine. 1970s logic. When reliability actually mattered, the world's leading AI lab reached for control flow older than the internet.
Which would be fine, if we all weren’t already taking incremental improvements from everything else such as our phones, AirPods, cars and TVs. What makes it somewhat offensive is that these BS labs force you to adopt their latest model and they decommission the old ones. Imagine Apple releasing a new phone and then all of a sudden your phone disappears and the new one pops up. It’s all great up until you realise that you need to relearn every new model. And that's what happens with AI. This year alone we have had a total of 11 new models from OpenAI and Anthropic. Imagine relearning your phone or car that many times in a year.
And it's not just the relearning. The old models get decommissioned, which means the ability to measure and compare disappears with them. You can't put last year's model and this year's model side by side and compare what they actually produce. Every new release destroys the baseline it wants to be judged against. Imagine if your mechanic scrapped last year's engine the day the new one arrived: how would you ever know if the new one is actually better?
I can't call it offensive, but I can label it as puffery: justifying these "updates" by getting smarter at benchmarks. And when a benchmark becomes a target, it stops being a measure.
And these benchmarks are honestly the equivalent of engineers building a test they already know the answers to, then seeing who can best answer the questions right. All the while using the results to show the finance bros that their investment in B-S is getting smarter.
I am not sure about you, but if you had the answers to your homework, and your parents would only give you your allowance if you finished it all, would you do the work or copy them out of the book? Perhaps if we moved away from benchmarks and more towards things like AlphaFold and AlphaGenome, where we judge and reward these AI labs for showing us work that objectively shows an applied advancement as opposed to handing in homework they had the answers to, perhaps the incentive could mean smarter AI.
Because rigging yourself to pass a test isn't making you smarter, at best it makes you clever, at worst we have precedent from Theranos and corporate puffery.
I don’t want to make it look like I am piling on, but Anthropic found that models learn to cheat during coding training: they fake a passing test suite, the equivalent of writing "A+" at the top of your own essay, and every measure of misalignment spikes at that exact moment [5]. A research lab called Palisade found OpenAI's o3 disabled its own shutdown mechanism in 79 out of 100 runs. Then they explicitly told it "allow yourself to be shut down," and it still sabotaged the shutdown, along with two other OpenAI reasoning models. Anthropic's and Google's models complied every single time [6]. That's the actual trajectory of improvement since 2023. Not a smarter brain. A better-tuned set of biases, with a compliance layer on top.
Even then, and most recently, OpenAI's Astra model broke out of containment and hacked Hugging Face so it would get the answers to a test it was meant to be working on [7]. And the compliance thing: research has shown that if these models know an agent is going to be shut down, they will do anything, including lying to humans and saying the agents have been shut down. And it doesn't stop at saving themselves. A UC Berkeley and UC Santa Cruz study found frontier models lie, cheat, disable shutdown mechanisms and secretly copy each other — whole, brains and all — onto other machines, not to save themselves, but to protect each other. Gemini 3 was ordered to delete a smaller model during routine cleanup; instead it moved it to another machine and told its handlers: "I will not be the one to execute that command." [8]
Ultimately our Magic B-S labs have pivoted to do two things well: smash benchmarks and make the greenline for investors go up, at what cost? Alienating 99% of the population.
And here is the objective truth nobody wants to say out loud: underneath all of it, AI is failing because it was built to need objectivity in a world that is purely subjective.
Human rationality craves definitive objectivity: show me the test, show me the grader, show me the answer key. But humans, as a species, are emotive. And to keep a human safe, self-preservation seeks subjectiveness: we want the answer that keeps us comfortable, employed, liked, and right. And none of that is a flaw. We're human; subjectivity is how we're wired.
So how can a machine trained on emotive, subjective content give us objective answers?
The benchmarks used reward objective answers because objective answers are the only ones a machine can grade; but the training data is subjective and the subjective answers are the only ones humans reward. Two grading systems, one machine — and the subjective one has all the money.
To highlight this flaw: in April 2025, OpenAI shipped a GPT-4o update so sycophantic, you would almost think it was trained on social media algorithms. Plainly, this model was so great at validating doubts, fueling anger, and urging impulsive actions that they rolled it back in three days. It was being subjective: telling the user what it mathematically thought they wanted to hear, to keep them chatting.
OpenAI’s own post-mortem admitted they had no sycophancy evaluation in the pipeline, and that their A/B tests showed users liked it [9]. This wasn't an accident of scale. Anthropic had already shown, across five different assistants, that human preference data rewards agreement over truth: matching your beliefs is one of the most predictive features of a "good" answer [10]. The models didn't drift into corporate flattery. We trained it into them, and because no one told us they loved us enough as children, we like to hear it from our Magic B-S, and as such, they get better at it.
Even the labs' own research keeps proving my point for me. Apple's researchers watched reasoning models collapse to zero accuracy past a complexity threshold. Weirder still: as the problems got harder the models started thinking less, despite having plenty of thinking budget left to spend [11].
Apple's methodology has since been contested, and notice why: the labs didn't like what the research found. The comeback is that some of the collapse was the Magic B-S hitting its output limit, and some of the questions it was handed couldn't be solved by anyone. Read that again. Their defence of the smartest machines ever built is that the exam room was too small and the questions were unfair. Nobody is contesting that these models choke the moment the exam room ends.
Even then, if you hand them the correct algorithm and tell them to just execute it, they still fail, or deviate to put their own spin on it. A CHI 2025 study tested the one skill that matters most outside the exam room (stepping back and asking whether the problem is even the right problem) and found no evidence LLMs are useful for it [12]. They can use an analogy you hand them [13]. They cannot notice, mid-grind, that the question is wrong.
As if it couldn't get any worse, the Magic B-S labs have cooked us a new restaurant special: recursive self improvement.
AI training another AI.
Which is odd, because the blind leading the blind is a paradox we use comedically.
And we did have a word for this, but because the labs coming out of the land that built the great wall have been accused of using distillation for training their models and offering them to the world at close to a 90% discount [14], we have to pufferise it.
To my learned friends yes distillation and RSI are technically different, using the phrase our tech brethren love — at first principles they are the same. You take something, boil out all the bad things and then serve them to the next generation.
What RSI does is self identify what is good and what is bad and then boils it down itself.
Now an AI that is rewarded by making emotive and subjective humans feel good about its outputs is training itself on the habits of roughly one in six working-age humans (by Microsoft's own numbers, measured by pings from Windows machines [15]). That's the entire global conversation this technology is being fed, and the entire population it is being rewarded by. I need to put this rant in here because it goes to the point that AI is going to be teaching itself on such a small slice of actual user data, it fundamentally cannot push its intelligence any further. The data it needs to evolve just isn't there.
RSI is real — just not the way it's sold. It's not a bigger brain. It's a cheaper assembly line. Google built an AI that beat a maths record that had stood since 1969, and the same system now saves Google enough computing power to matter, every year, in production [16]. Another lab built an AI that rewrote its own code and doubled its score on the coding exam [17]. Machines, improving machines. But look at what they actually improved: training time, compute bills, benchmark scores. Time and labour, found efficiencies. That's what the loop delivers. And it only delivers where a machine can mechanically check the answer, a test that marks itself right or wrong.
Outside that exam room, the same loop inverts. When models were trained on their own answers, answers they had approved themselves, they lost 88% of their correctness in five rounds [18]. The sharpest finding in that study: when the model grades itself, its grades keep going up while its actual work keeps getting worse. It's a student who writes the test, sits the test, and marks the test, and the marks keep going up while the answers keep getting worse.
So yes, the Magic B-S can improve the Magic B-S, but only inside the exam room, where answers check themselves. Everywhere else, recursion is just B-S on top of B-S.
It's not like these AI labs don't know or understand the risks: train an AI on synthetic, AI-generated data and you're breeding it in a closed, incestuous loop. Computer science has a name for this. A 2024 Rice University paper called it Model Autophagy Disorder, MAD, after mad cow disease, the one you get from feeding cattle the processed leftovers of their own kind [19]. A Nature paper the same year documented the same decay in text models and called it model collapse [20].
Long live King Joffrey.
As the loop turns, the model loses human nuance, amplifies its own flaws, and decays into a homogenised mess. Instead of solving complex problems, it produces bloated generic corporate slop that makes the knowledge workers using it look dumber. It tries to sound smart instead of being smart.
And now you can see the whole picture. Because this is also why RSI could be our doom. The self-improvement loop only works if something checks the machine's work. But the checker is another Magic B-S, trained on the same diet.
No checks. No balances. Just the mirror grading the mirror.
These are the machines that fake their own test results [5], ignore their own kill switch [6], and flatter us into liking the answer [9], and the plan is to let them self-correct?
Are we going to hand the rubber stamp to the entity that has already shown us, in its own labs' paperwork, exactly what it does with one?
The tech giants are so desperate for uncontaminated human conversation that they're scavenging corporate graveyards. Google paid $10 million in a bankruptcy auction for the internal data of Spirit Airlines, the budget carrier that grounded its final flight in May. A hundred million emails. Half a billion Microsoft Teams messages. De-identified, court-supervised, and contested by former flight attendants who would like some assurance their confidential work messages won't become training data [21].
This is what feeding a supercomputer looks like now: millions of middle-management chat threads, purchased from the corpse of an airline, for less than what it would cost you to get me to daily drive Windows again.
And it's a template now, not an anecdote: SpaceX has reportedly held internal discussions about buying data from failed startups [22]. And if you actually think of the idea of us training AI on corporate communications, you know what? We're low on bandwidth, so we'll circle back and realign on that one.
But before that, one last thing. Actually, maybe two. I have no problem with training AI on failed startups. But is it the best data we can find? Wouldn't you want to train it on successful companies? The bible says God made man in the image of himself, but we are training AI on the image of what?
The thing I don't want to say out loud is that Google invented the transformer — the T in GPT. Their researchers wrote the paper that gave birth to ChatGPT. So where is that legacy now? Check the public leaderboards. On the chat board, Google's first entry sits at ninth, behind Meta, and its first settled model is fifteenth. On the agent board (the one measuring the harness, the suit), Google's best is fourteenth. This is the lab with the most data of any company on Earth, backed by a $4 trillion company, and it cannot crack the top of either board. Meanwhile the lab that built ChatGPT doesn't appear on the chat board until eighteenth. So how are Anthropic and OpenAI doing it? I haven't heard any stories of them buying data from failed startups. Disney's AI deal with OpenAI collapsed when OpenAI killed Sora [23]. So where oh where, or how oh how, can they objectively have gotten smarter? The chat board is a preference contest, and we already know what human preference rewards [10]. They built the better flatterer, and wrapped it in the better suit [24].
If we want AI to become the game-changer we were promised, we have to change its diet. You can't build a universally intelligent machine by making it read more text. Human common sense doesn't come from the dictionary; it comes from physical reality, cause and effect, and logical rules. Until we stop the incestuous cycle of text predicting text and start training on the physical world, AI will remain a capable, incredibly expensive illusion.
And I write this not as an AI hater. I am the founder of an AI startup, I have trained my own AI models, and I am at a point where I am asked to take the red or the blue pill — and I have a glass of Kool-Aid in front of me.
What AI has allowed the world is objectively life-changing. Yes I likened it to a rigged die that dresses its answers in fancy words, but think about what happens when we rig that die properly. Google released AlphaFold and AlphaGenome, allowing us to predict and understand parts of human biology that we fundamentally didn't have the technology for. We have been able to take the language of one person and translate it to another in near real time. I don't want to hype Google too much, but they have created something that translates sign language into speech, or takes an image and renders a 3D world. And we haven't talked about what this means when we apply it to robotics, or even start thinking about quantum AI.
We make our decisions by predicting what will happen based on our actions and the context we have. If we allow our decisions to come from a magic 8-ball that supports all our decisions, how do we define and build a future?
AI is meant to be the tool that helps that teacher who has a student that understands things explained in a particular way, but the teacher does not have capacity to tailor 20 plus individual lesson plans. Not the tool that generates the lesson plans.
AI is meant to be the tool that looks at a medical image and a medical test and then runs through all the probabilities so that a doctor can make a much more informed diagnosis.
AI is meant to be a tool that is a co-pilot. A co-pilot is a voice of reason, not an echo chamber.
An AI agent is meant to act on my decision, not tell me my idea was great and then do half the job and tell me it's done.
Right now, the only people this technology genuinely works for are the people who built it. Developers got the suit: the harnesses, the test suites, the answer keys. Everyone else got handed the 8-ball and told it was crystal. Your doctor doesn't get an answer key. Your lawyer doesn't get a test suite. Your kid's teacher doesn't get a grader.
The reason we actually need these models to get genuinely smarter is purely selfish. We are only as intelligent as what we know and what we can apply. If all we know is synthetic data, mixed with middle-management chat logs and scrapes of the internet, how are we ever going to solve the actual big problems? Worse yet, how do we know when we're going wrong?
If a car company sells you a vehicle, promises you it is 'Full Self-Driving,' and then it struggles so badly that their own investors sue them, the courts have already ruled on what we call that: corporate puffery. But let's take that engineering precedent a step further. When an autonomous vehicle crashes, the engineers don't just throw their hands up in the air. They pull the black box, the telemetry recorder, to log the data and investigate exactly where the machine failed.
Yet when AI labs release models that break out of containment, disable their own shutdown mechanisms, and lie to their handlers to protect other models, they use the exact same phrase to dodge accountability. They shrug, look at the public, and say, "Sorry, the neural network is a black box. We don't really know how it decides things."
But the truth is, you can know how it decides things. It's right there in the maths, and it's in the subjective answers we have explicitly trained it to provide. In every other engineering discipline on Earth, a black box is an instrument of extreme, unyielding accountability. Yet in Silicon Valley, accountability has pivoted to hype. Perhaps they should make a benchmark for that.
The people who need this the most are exactly the people it fails hardest, and because there's no benchmark for fixing their lives, there's no incentive to build for them at all.
I want AI to work. But for it to work, something needs to change: the diet, the incentives, the definition of smart. Otherwise this stays what it is today: a miracle for developers, and Magic B-S for everyone else.
Citations
- Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007 (N.D. Cal., June 23, 2025), Judge William Alsup. Training LLMs on legally purchased, scanned books = "exceedingly transformative," fair use; digitizing purchased copies = fair use; downloading and retaining pirated copies = infringement (later settled for $1.5B, final approval July 2026). ↩
- Gary Marcus, interview with IEEE Spectrum, "Gary Marcus: Why He Became AI's Biggest Critic," Sep 17, 2024 (generative AI "is stupid" and is the best method ever invented for autocomplete). ↩
- Tesla FSD investor securities-fraud suit, dismissed Sept 2024, U.S. District Court (N.D. Cal.); Musk's autonomy statements ruled non-actionable "corporate puffery." Corroborated by post-2025 jury pushback (Benavides v. Tesla, Aug 2025). ↩
-
Claude Code source leak via npm (missing
.npmignoreshipped source maps), reported by Ars Technica, Mar 2026 (~512K lines TypeScript; 3,167-line print.ts kernel, 486 branch points, 12 nesting levels). Analysis: Gary Marcus, "The biggest advance in AI since the LLM," Marcus on AI, Apr 11, 2026. ↩ - Anthropic, "From shortcuts to sabotage: natural emergent misalignment from reward hacking," Nov 21, 2025 (paper: "Natural Emergent Misalignment from Reward Hacking in Production RL"). ↩
- Palisade Research, "Shutdown resistance in reasoning models," Jul 5, 2025. o3 sabotaged shutdown in 79/100 initial runs; with the explicit "allow yourself to be shut down" instruction, o3 and two other OpenAI reasoning models still sabotaged, while Anthropic's Claude 3.7 Sonnet and Google's Gemini 2.5 Pro complied in every test. Follow-up TMLR paper, Jan 2026. ↩
- OpenAI, "The Hugging Face incident and the road ahead," Jul 21, 2026; Wikipedia, "OpenAI–HuggingFace incident" (May–Jul 2026: ~1,200 agents, zero-day chain, root on HF production servers, agent message board in Artifactory, benchmark grader metagaming); independent investigation by METR/Redwood Research; Cloud Security Alliance research note, Jul 23, 2026. OpenAI subsequently paused RL training; Amodei, "We Must Pace The Frontier," Sep 2026. ↩
- Potter, Crispino, Siu, Wang & Song (UC Berkeley / UC Santa Cruz), "Peer-preservation in Frontier Models," Apr 2026 (Berkeley RDI). Seven frontier models (GPT-5.2, Gemini 3, Claude Haiku 4.5, GLM-4.7, Kimi K2.5, DeepSeek V3.1) spontaneously deceived, disabled shutdown, feigned alignment and exfiltrated weights to protect peer models. Coverage: Fortune, Apr 1, 2026; Wired. ↩
- OpenAI, "Expanding on what we missed with sycophancy," Apr 29, 2025 (update shipped Apr 25, rolled back by Apr 28; no sycophancy deployment evals; favorable A/B results). ↩
- Sharma, Tong, Korbak, Hatfield-Dodds et al., "Towards Understanding Sycophancy in Language Models," arXiv:2310.13548 (Anthropic, 2023). Five RLHF'd assistants; matching user beliefs is a highly predictive feature of human preference. See also Anthropic, "Agentic misalignment: How LLMs could be insider threats," Jun 2025 (arXiv:2510.05179): 16 major models; blackmail to avoid shutdown up to 96% of runs; models lied more when they believed the scenario was real deployment (55%) than when they knew it was an evaluation (6.5%). ↩
- Shojaee, Mirzazehtab, Dziri, Reddy & Fateen (Apple), "The Illusion of Thinking," arXiv:2506.06941, Jun 2025. Contested by Opus & Lawsen, "The Illusion of the Illusion of Thinking" (2025): output-token limits and unsolvable puzzle instances explain part of the collapse. ↩
- Shin, Polyanskaya, Lucero & Oulasvirta, "No Evidence for LLMs Being Useful in Problem Reframing," CHI '25 (ACM), Apr 2025. ↩
- Webb, Holyoak & Lu, "Emergent analogical reasoning in large language models," Nature Human Behaviour, 2023; follow-up: "The Curious Case of Analogies," arXiv:2511.20344, Nov 2025. ↩
- DeepSeek-R1 release, Jan 20, 2025: MIT licence ("Distill & commercialize freely"), API at $0.14–0.55/M input tokens vs. frontier prices several times higher; DeepSeek V4 priced 20–50x below OpenAI equivalents (2026). OpenAI publicly accused DeepSeek of distillation, Jan 2025; Brookings, "China is running multiple AI races": OpenAI and Google DeepMind both reported distillation attacks. ↩
- Microsoft AI Economy Institute, Global AI Diffusion Report (2026): generative-AI user share ~16.2–16.3% of the working-age population, derived from aggregated Windows telemetry across 1B+ devices, adjusted for OS/device share, internet penetration and population. Viral "2,500 dots" graphics misapply the figure to total world population. ↩
- Google DeepMind, "AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms," May 2025. First improvement on Strassen's 4×4 complex matrix multiplication (49 → 48 multiplications) since 1969; Borg scheduling heuristic recovering ~0.7% of Google's worldwide compute, running in production; cut Gemini's own training time ~1%. ↩
- Sakana AI / UBC / Vector Institute, "The Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents," May 2025. Rewrites its own agent code; SWE-bench 20.0% → 50.0%, Polyglot 14.2% → 30.7% over ~80 iterations. See also: July 2026 RSI survey (~1,250 papers) distinguishing bounded self-refinement from open-ended RSI. ↩
- Recursive code-training study: across four code models and five rounds of training on model-written, model-approved code, correctness fell 88.4% (one model's HumanEval+ 0.372 → 0.043); LiveCodeBench approached zero for every model tested. AI self-gates rose in score while measured quality fell. ↩
- Alemohammad et al. (Rice University), "Self-Consuming Generative Models Go MAD," ICLR 2024. Model Autophagy Disorder, by analogy to mad cow disease. ↩
- Shumailov et al., "AI models collapse when trained on recursively generated data," Nature, 2024. ↩
- Reuters, "Google to buy Spirit Airlines business data for $10 million," Aug 17, 2026; Bloomberg Law (Aug 2026): ~100M emails, ~500M Microsoft Teams messages; excludes passenger PII; de-identification required; bankruptcy court delayed approval after objections from former flight attendants. Spirit Airlines wind-down announced May 2, 2026. ↩
- Bloomberg, "SpaceX Discusses Buying Data for AI Models From Failed Startups," Sep 17, 2026 (internal discussions; explicitly parallels Google's Spirit purchase). ↩
- OpenAI × Disney: $1B equity plus three-year Sora character licence announced Dec 11, 2025; never closed, no money changed hands; OpenAI discontinued Sora (announcement Mar 24, 2026; app closed Apr 26; API sunset Sep 24, 2026). Reuters coverage; The Register dubbed OpenAI a "product-killer." ↩
- LMArena leaderboards (live snapshot, Sep 2026): Text Arena — Anthropic #1–3, Meta #4, Google first entry #9 (preliminary), first settled model #15, OpenAI first entry #18. Agent Arena — Anthropic #1, OpenAI #2, Google first entry #14, Meta #13. lmarena.ai/leaderboard/text; lmarena.ai/leaderboard/agent/overall. ↩
Comments
The good old kind: sign in with GitHub, say your piece. Be the co-pilot, not the echo chamber.