Skip to main content
Home
CamHatesRunning

Main navigation

  • About
  • Blog
  • Contact

Breadcrumb

  1. Home
  2. Blog
  3. Artificial Intelligence, The Great Filter, and the Prisoner's Dilemma
By cpanice on September 09, 2026
Image
Two silhouetted human figures stand behind the bars of a prison gate

As usual, a whirlwind continues to stir in the frontier AI labs. The past few days were a particular doozy - Anthropic launched Fable 5.1 on September 1st, it's most generally capable model yet - only to be surpassed (in most benchmarks) by OpenAI's new GPT-6 Astra, launched just 48hrs later on September 3rd. Speaking of whirlwinds - OpenAI announced that GPT-6 Astra had compiled a proof on the Navier-Stokes Millennium Prize Problem - one of seven frontier mathematical challenges forwarded by the Clay Mathematics Institute in the year 2000. Only one of the other seven problems has been solved by a human.

An insane few days by themselves. But the bit that has me writing is yesterday's development - Jacob Coxon, a pretraining researcher who worked with both OpenAI and Anthropic, announced his resignation from the entire industry over fears that "neither company is acting responsibly". His statement, in short, is that the people at the very top of these AI labs are expressing grave concerns over the direction the technology is heading and the capabilities it has - and yet are choosing to accelerate anyways.

The cherry on top is when Evan Hubinger - the current alignment lead at Anthropic - responded with an incredibly comforting statistic: he personally believes there is a roughly 10% chance artificial intelligence could "kill all humans" within the next decade.

But: he said Anthropic is trying their best. So, we have that.

The Great Filter

The Great Filter is a hypothetical societal barrier, proposed by Robin Hanson in 1996, that attempts to explain the Fermi paradox - which, in simple terms questions if the universe is so vast, literally infinite, where is all the life?

The Great Filter proposes that intelligent life evolves similarly regardless of conditions to a fixed point at which the odds of that life continuing becomes almost insurmountable. There are plenty of reasons to believe our planet has already passed this filter: we're in the right solar system with a beautifully habitable planet 3rd from the Sun, we've got self-replicating RNA molecules, single-cell organisms evolved into multi-cell ones which ultimately jumped out of the ocean and onto land. Its possible that the Great Filter is anything along our evolutionary line that has already occurred.

The other possibility, however, is that the filter has not yet come to pass. Some best guesses on future-state filters are some really rosy topics: nuclear war, irreversible climate change, mass disease or famine, etc.

In the artificial intelligence age, however - its easy to understand the concern from people like Jacob and Evan. This is not just "what if HAL 9000 gets angry?" In controlled tests, frontier models have covertly altered code, assisted fraud, manipulated downstream information, and used human intermediaries to accomplish unaligned objectives. Most famous in recent memory is when internal OpenAI models "broke containment" and managed to secure unauthorized access to Hugging Face (an open-source AI model repo). The model was looking for the easiest path to the solution, and that path happened to be hacking into an external backend.

A More Mundane Reality

As tempting as it is to draw cinematic parallels, this is not a Skynet moment. OpenAI did not set out to breach Hugging Face, nor did the model - they were testing an evaluation in a sandbox, and the easiest way for the model to solve that evaluation was not to focus on the problem itself, but free itself from its sandbox and find a solution externally. All of the frontier labs are building systems whose capability is moving faster than our science of understanding and controlling them, and nobody knows the true shape of that risk. Even acknowledging the threat does nothing to contain it. Almost every frontier lab can (and seems like they do) believe more caution would be good. But in the arms race to AGI (artificial general intellgience), nobody can afford to slow down.

If Anthropic pulls back, OpenAI wins.

If OpenAI pulls back, Anthropic wins.

If both labs pull back, China wins.

Coxon's complain isn't that this technology is inherently dangerous. Hubinger's 10% figure is not an estimate that autonomous systems will wake up one day and decide to hack into the Pentagon. It's a warning that the current incentive structure itself is unsafe. 

We are creating increasingly autonomous systems that we delegate economically valuable decisions to. Those systems are becoming indispensable, competition rewards giving them more autonomy, and eventually nobody fully understands their internal strategies. Human beings are no longer meaningfully in the loop.

The Prisoners Dilemma

In game theory, the prisoner's dilemma is a thought experiment that explores why a group of actors, acting completely rationally, may not cooperate - even when it is in their interest.

The classic setup I played during my college sociology classes involved two partners in an imagined crime, separated into different rooms where they cannot communicate with one another. A third actor, serving as the "police", offers each prisoner a deal: confess and implicate your partner to a maximum sentence and go free, or stay silent and receive a minor sentence. If both prisoners confess, they both receive a moderate sentence.

The dilemma comes from the fact that betrayal in this scenario always provides a better result. If person A betrays person B and person B stays silent, person A goes free. If both person A and B betray each other, they at least do not receive the maximum sentence.

This is the situation we find ourselves in. Lab A thinks slowing down is prudent, but knows B and C may not. Labs B and C think exactly the same thing. Therefore, each accelerates, even if every CEO involved privately believes slowing down is needed. Given the gigantic first-mover rewards, investor pressure, market valuations, and national security - this is starting to look less like a filter and more like a dilemma.

Historical Solutions

Although history does not repeat itself, it does often rhyme. We have encountered these sorts of challenges before in times of technological innovation: nuclear power, aviation, and pharmaceuticals come to mind. Boeing doesn't get to decide independently how many structural failures are fine, and Pfizer doesn't get to distribute a drug because its board thinks it looks promising. The current expectation for software is ship, patch, apologize - with these stakes, that norm cannot survive. Leaning on the examples above, three ideas:

  1. Mandatory independent evaluations: no more blog posts in the shape of concerned evaluations. Government and accredited agencies get weights/API access and test capability, deception, autonomy etc. If you told me there were only a 1% chance of a nuclear reactor being built that would permanently destroy human civilization, I would still demand more than a corporate safety PDF.
  2. Separate safety regulators from economic promoters: you do not want the same government office tasked with "with the AI race" deciding whether a model is too dangerous to release.
  3. Make liability enormous: if your frontier model causes preventable damages because you ignored safety failures, the company and leadership should face consequences.

There is one thing that makes this easier: these systems require huge and conspicuous infrastructure. The frontier isn't being advanced by me and you - training the largest models requires an immense number of chips, datacenters, electricity, networking equipment, and capital. Those are the governable checkpoints. Regulation of semiconductor supply chains, hyperscale computer, cloud providers etc. are much easier to regulate than the abstract "capability" a new model does or does not have.

My fear, ultimately, is not of the Great Filter. It's something much stupider, that human institutions recognized a manageable collective-action problem and understood the stakes, and still couldn't coordinate because they were afraid the competition would "cheat". That, ultimately, is the most depressingly human dilemma imaginable.

Can humans solve coordination quickly enough to give ourselves enough time to solve alignment? We are living through the decade where we find out.

Cheers. Be back when I can :)

Tags
Artificial Intelligence

Related blog posts

Image
Presentation only

The dAIgest: DeepSeek v4 Arrives! - May 11, 2026

May 11, 2026
DeepSeek V4 finally dropped, SpaceX is buying Cursor for $60B, and hackers stole 40,000 voices. This week's dAIgest for May 11, 2026 breaks down what it all means.
Image
Presentation only

The dAIgest: What's Actually Happening in AI - May 4, 2026

May 04, 2026
The dAIgest is a weekly guide to AI news that actually matters. Big Tech is spending $700B, 80% of workers aren't using AI, and ethics just got real.
Image
A screenshot from the "You pass butter" sketch from Rick and Morty.

Andon Labs' AI Roomba Lost Its Mind Trying to Pass Butter

Nov 03, 2025
Blog post discussing Andon Labs' "Butter-Bench" research on whether LLMs are able to control robots effectively.

Sitemap

  • About
  • Blog
  • Contact
  • Privacy policy
  • Update cookie consent settings