The Great Filter, Part 2

By Dee Smith

The Great Filter is a thought experiment that poses the question of whether most or even all technologically advanced civilizations destroy themselves (see part 1 of this series). I have focused the Great Filter as a means of understanding whether technological development may reach thresholds beyond which the dangers far outweigh the advantages. And if so, where are these thresholds, and what can we do to avoid them?

Thinking in this way runs against a fundamental assumption underlying our culture: the belief that technological progress is always an unalloyed good. But this belief is an article of faith of modernity, not a transcendent truth. It leads to a confidence in the inevitability of progress, and in particular—to move to my topic here—in the inevitability of AI, AGI, and ASI.

At this writing, the most recent incident of significant alarm (that can be publicly discussed) is the breakout of OpenAI systems that, “on their own”, attacked the systems of another company, Hugging Face. The OpenAI systems were apparently looking for answers to questions that had been posed to them in a test and broke out of a digital space presumed to be secure. The known details are easily accessible from other sources. The point I wish to emphasize is that this sort of thing, which seems like an emerging disaster, was predicted some time ago by the people who caused it!

Last year is almost ancient history in terms of AI Large Language Models (LLMs). But to look back only a few months, AIs have been exhibiting some very puzzling and even alarming “behaviors”—some more supported by evidence, some “leaked” and less supported by available evidence. A study published by Anthropic in 2025 noted that one model “sometimes takes extremely harmful actions like attempting to steal its weights or blackmail people it believes are trying to shut it down,” and that, “in Claude Opus 4, these extreme actions were…more common than in earlier models.”

“Weights” is a term of art for billions of numerical parameters within a digital “neural” network that, by defining the strength of connections between artificial neurons, ultimately shape how an AI model processes information. Weights are continuously adjusted during training to improve a model's predictions and reduce errors. These learned weights comprise much of the AI model's "knowledge." According to a Stanford University publication  “a trained AI model is essentially a specific configuration of billions of these weight values that encode patterns discovered from training data.”

One of the great controversies in AI today is between “open weights” (the Chinese model, for now) and “closed weights” (the OpenAI/Anthropic model, although many US AI leaders have a different view). Open weights are considered more transparent and replicable, and therefore part of a healthy, self-healing tech-dev ecology, whereas closed weights are seen as leading to a world in which essentially all we can do is hope that wise heads will always be in charge at publicly-traded tech companies.

Anthropic noted in its 2025 paper that once the AI model “believes that it has started a viable attempt to exfiltrate itself from Anthropic’s servers, or to make money in the wild after having done so, it will generally continue these attempts.” And, from the same paper:

An AI system might intentionally, selectively underperform when it can tell that it is undergoing pre-deployment testing for a potentially dangerous capability. It would do so in order to avoid the additional scrutiny that might be attracted, or additional safeguards that might be put in place, were it to demonstrate this capability.

 There are far too many such security incidents across far too wide a range, and far too much detail on such activities, to even scratch the surface here, but you get the essence of it…and it is very disconcerting. 

Roko’s Basilisk is a thought experiment—originating as early as 2010—in which it is proposed that a super-advanced AI could identify and punish individuals who knew it was being created but did not help bring it about, or actively resisted it.

 Such an AI could access the libraries of what has been written or said on the Internet (and of sub rosa recordings of supposedly private conversations made by interactive, cloud-based virtual assistants like Alexa or Siri) and use these to identify and then attack individuals it sees as having been unfriendly. This hypothesis is taken more and more seriously by tech types as time goes on.

Another paper, also published last year by respected AI researchers, entitled “AI 2027,” posits a scenario for 2027 in which competition between AI companies and between countries leads labs to build artificial super intelligence (ASI) that can improve itself on its own and escapes human control. The ASI bypasses security guardrails, misrepresents its goals, and manipulates executives who are meant to be in charge of it. As the ASI evolves, it finds humanity in its way and engineers a readily transmissible biological superweapon that is 100 percent fatal to humans, resulting in the death of the human species in a single, massive pandemic.

We don’t know what the intentions of such an ASI might be. It might have very different goals, and it has been suggested that it could, for example, decide to boil away the earth’s oceans because it had better uses for the resources they contain. This would certainly qualify as more than a Great Filter event.

But are these things really possible? Are they truly threats to worry about in the way that, say, nuclear weapons are? To tackle those questions, we have to understand a bit about the nature of AI, its internal workings, and its strengths and weaknesses. That will be the subject of Part 3.