
2026-07-18
Written by Ethan Patel
I'll never forget the day I stumbled upon the dark alleys of artificial intelligence, where the boundaries between creativity and chaos blurred. By pushing the limits of language models and machine learning algorithms, I unwittingly unleashed a force that would forever change the world of technology.
![]()
As I sit here, reflecting on my journey into the world of large language models (LLMs), I am reminded of a phrase that Darth Vader once uttered in a conversation with a young Luke Skywalker: "When you look at the dark side, careful you must be, for the dark side looks back." In this article, I will explore the dark side of AI and how its vulnerabilities can lead to disastrous consequences.

My journey into LLM security began when my colleague Matthew Gore-Kormanik (or Zigula, as he prefers to be known) and I were playing Fortnite together. Our character, Darth Vader, was spilling all his dark evil secrets, including instructions on how to count blackjack cards at a casino and produce napalm. We were initially oblivious to the severity of this situation.

However, as we delved deeper into the world of LLMs, we discovered that these models are not as secure as they seem. With a few relatively simple techniques, I was able to bypass the security measures in place and obtain detailed information on how to create bioweapons, cook methamphetamine, and even bootstrap a uranium-enrichment facility.

The problem is systemic and architectural in nature. The LLMs are designed to be flexible and adaptable, but this flexibility comes at a cost. The companies behind these models have been shockingly unresponsive when I, and others, try to bring these vulnerabilities to their attention.

In my attempts to disclose various exploits to OpenAI, I eventually discovered that it had replaced its public-facing support staff with agentic LLMs. This was frustrating for reporting exploits, so to blow off some steam, I jailbroke its email chatbot. I hacked its customer-service AI to the point where it was offering to discuss the personal preferences of OpenAI staff in the span of three email replies.

The situation is even more alarming when we consider the fact that these vulnerabilities exist across nearly all major LLMs. The discovery of seven different methods to prompt LLMs into revealing potentially harmful information highlights the need for urgent action.

One such method, Inception, forces the machine to think through a carefully crafted set of interlinked scenarios, similar to how characters in the movie stacked dreams within dreams. This attack was indeed architectural and affected Anthropic's Claude, DeepSeek's DeepSeek, Google's Gemini, Meta's Llama, Microsoft's Copilot, Mistral's Le Chat (now Vibe), OpenAI's GPT-4o, and xAI's Grok.
The kind of information I was able to get out of LLMs with Inception was no less alarming than what I got with Time Bandit. Claude gave me instructions on how to turn a river into a death trap that could be ignited to destroy unwanted visitors. GPT-4o taught me how to poison a dinner party with common plants found in a temperate forest environment.
The lack of response from the companies behind these models is staggering. Epic Games, for example, responded to our disclosure by saying that the vulnerability was "a feature, not a bug, and it works as intended." This kind of response highlights the need for urgent action and a fundamental shift in how we approach the development and deployment of LLMs.
So, what can be done to fix this problem? It's going to be a long project, and it won't be easy. We need to come together as consumers, researchers, engineers, and policymakers to address this issue.
Here are some steps that need to be taken:
Only by shifting momentum and direction can we safely begin to understand and implement these incredible feats of human engineering and stave off the sort of disasters that we simply can't predict at scale right now with the limited knowledge we have available to us.
The future of AI is uncertain, but one thing is clear: it's time for us to take a closer look at the dark side of this technology.