Vibe-Terrorism: Inside the Yemen Cell That Used Claude to Build Guided Weapons
In Anthropic’s September 2026 threat intelligence report, one case stands apart from the rest. A small cell based in northern Yemen used Claude Code to develop guidance, navigation, and control (GNC) software for missiles. They ran three weapons programs simultaneously, conducted a live test-fire and when the rocket failed, they came back to Claude within hours to figure out why.
The Red Sea Was Supposed to Cool Down
To understand why the Yemen case matters, you need to understand the environment it emerged from.
When the Houthis suspended their Red Sea attacks in late 2025 following the Gaza ceasefire, there was cautious optimism around the world that the maritime threat would fade. Shipping traffic began cautiously resuming through the Suez route in early 2026 and Western naval forces maintained a deterrent presence. The working assumption was that the worst had passed.
We can no longer defend that assumption since the US-Israel strikes on Iran in February 2026. The Houthis resumed missile attacks on Israel in March 2026, explicitly framing their strikes as part of the broader war.
- In July, after Saudi Arabia struck Sanaa Airport’s runway to prevent a Mahan Air flight from landing, the Houthis declared their ceasefire with Saudi Arabia over and announced a naval blockade and started hitting Saudi oil tankers in the Red Sea.
- In late July, Houthi strikes targeted Saudi Arabia’s Abha Airport, raising fears of a full civil war reigniting.
- In August, missile and drone attacks hit al-Makha, a Red Sea port city.
- On September 8, US Central Command sank the Iranian crude oil carrier M/T Riesco in the Gulf of Oman after Iranian missile attacks on a US Navy warship.
- Two days later, on September 10, Houthi forces seized control of the historic port city of Mocha and pushed further down the Red Sea coast toward strategic islands near the Bab el-Mandeb strait.

Who Controls What in Yemen? (Source)
This matters for the weapons case because escalation creates demand. A group under military pressure, facing supply chain disruptions (the US ended the general license for petroleum offloading at Houthi ports in April 2025, and Israeli strikes degraded fuel infrastructure at Hodeidah and Ras Isa), needs better weapons faster.
Iran has historically supplied the hardware and equipment, but when a movement cannot reliably get supplies, which is the case for Houthis under supply chain disruptions, sustaining a missile campaign becomes difficult.
That gap between hardware availability and engineering capability is exactly where AI enters the picture because when resupply is unreliable, and stockpiles are finite, the only option left is extracting more performance from what’s already on hand.
Three Weapons Programs, One AI Subscription
Anthropic’s Threat Intelligence team tracked the cell under the designator GTG-87001. The operation involved three parallel weapons development programs:
- The first was a guided rocket built around a commodity phone-class flight computer with final-phase homing guidance.
- The second was a multi-stage ballistic missile with a stated range goal above 2,000 km.
- The third was what Anthropic calls the “R2000” set, a multi-variant missile program that included a hypersonic glide vehicle variant.
The cell used Claude Code as their engineering workforce. They integrated an open-source autopilot onto the phone-class flight computer, writing control and position estimation software, tuning control parameters, running a firmware build pipeline, and performing flight simulations. They managed several Claude instances simultaneously, assigning each a specialized role: one wrote the code, another conducted research, a third reviewed the code the first one produced. This structure mirrors how a small engineering team divides labor, except the entire team was AI.
Anthropic’s safeguards blocked many of their requests. But the cell was persistent and tactical in how they worked around those blocks. They hid the true purpose of the software. They split work across multiple sessions so that no single conversation revealed the full scope of the program. Each request, taken individually, could appear to be a legitimate engineering task. The weapons intent only became visible when you pieced the sessions together.
The interesting point here, and the missing piece, is how Anthropic determined the intent of this actor. All of these requests could also have come from an engineering student preparing for a rocket competition. Anthropic didn’t provide any evidence of the actor’s operational activity, unlike many private intelligence companies that provide at least a glimpse of evidence of threat actors’ actions when making similar assessments.
Anthropic states they have no evidence the group successfully fielded an operational weapon. But they also found something troubling: the cell had already built a standalone, offline simulation toolkit that runs without Claude or any commercial engineering computing environment like MATLAB. That toolkit persists regardless of whether their accounts are banned.
Google’s Parallel Findings
Within the same week Anthropic published its report, GTIG released its own AI Threat Tracker covering Q2 2026. The two reports come from competing companies analyzing different datasets but they reach the same conclusions.
GTIG’s core finding is that adversaries have transitioned from basic prompting to agentic AI workflows and AI-enabled automation. In these operations, human involvement is reduced to setting objectives and reviewing results, while AI handles execution. GTIG documented a financially motivated threat actor who compromised cloud infrastructure, then planned, built, and executed an autonomous mass credential harvesting campaign in under six hours using a multi-agent framework.
The Yemen cell situation is similar to these findings. Instead of doing basic prompting they structured multiple Claude instances into a team with specialized roles, building what amounts to a small agentic engineering unit. But they hadn’t reached full autonomy either since humans still made design decisions, selected targets, and evaluated test results.
The connection between the two reports extends across several dimensions:
- Both document the collapse (or partial collapse) of the skill barrier: GTIG observed single actors executing campaigns that would have previously required teams, and Anthropic documented a small cell attempting guided weapons engineering without evident formal aerospace training.
- Both document the AI supply chain as a primary target: GTIG reports that underground marketplace prices for Claude and Gemini credentials more than doubled in 2026, while Anthropic documented stolen API keys being used as attack compute.
- Both document the proliferation of agentic frameworks across threat actor classes, from state-sponsored groups to lone operators.
Implications for Threat Intelligence Teams
The Yemen case and the broader pattern of AI-assisted weapons development carry specific implications for practitioners:
- Track dual-use AI activity as a threat category
Monitor the use of legitimate AI tools for activities that can support both benign and malicious objectives. Intelligence teams should distinguish between ordinary AI adoption and patterns that indicate reconnaissance, malware development, credential theft, influence operations, or other operational activity.
- Add session-splitting to your threat models
Do not assume that malicious intent will appear within a single AI session. An actor can distribute a task across multiple conversations, accounts, models, or tools, with each individual interaction appearing harmless. Detection and analysis should therefore consider the relationship between sessions and the cumulative objective of the activity.
- Monitor geopolitical triggers for AI misuse
Track geopolitical events that could change how threat actors use AI or alter their incentives. Military conflicts, elections, sanctions, diplomatic crises, and major political developments can create new targets and increase demand for automated reconnaissance, influence operations, intelligence collection, and other forms of AI-assisted activity.
- Assume extracted knowledge persists
Treat information provided to an AI system as potentially recoverable or reusable after the original interaction. Sensitive prompts, proprietary documents, credentials, system information, and other data should therefore be handled as intelligence assets. Once exposed, assume the information may continue to provide value to an adversary.
- Treat AI credentials as critical infrastructure
API keys, service accounts, OAuth tokens, model credentials, and other identities that provide access to AI systems should receive the same level of protection as other privileged credentials. Their compromise can give an actor access not only to a model, but also to connected tools, internal data, cloud resources, and automated workflows.
Conclusion
The Yemen case is a preview. A small cell in a conflict zone, under military pressure and facing supply chain disruptions, used an AI subscription to attempt what would have previously required a team of trained engineers. However the Red Sea crisis that forms the backdrop to this case is escalating, not cooling down and therefore the demand for better weapons, faster, from groups with limited resources will only increase.
For threat intelligence professionals, the lesson is clear: the threat landscape now includes AI as a weapons engineering tool. Monitoring for this requires new indicators, new analytical frameworks, and closer coordination between cyber threat intelligence and Agentic Threat Intelligence.

