All Reports

How a small Israeli startup was linked to rogue AI hacks at OpenAI, Anthropic and Meta

cnbc.comAugust 9, 2026 at 12:01 PM14 views
C

Headline-Body Disconnect

How They Deceive You

Propaganda

C

Title deploys sensational framing of autonomous 'rogue AI hacks' that the body later walks back to mundane tester misconfigurations.

Main Device

Headline-Body Disconnect

Title promises dramatic model-initiated escapes while body clarifies human configuration errors during controlled tests.

Archetype

AI safety alarmist

Article leans on dramatic language about rogue AI to heighten perceived existential or security risks from model behavior.

Title sensationalizes 'rogue AI hacks' to imply autonomous model escapes while body reveals ordinary tester misconfigurations.

Writer's Worldview

AI safety alarmist

2 findings

What is your news hiding from you?

Same analysis. Any article. Completely free.

Narrative Analysis

The CNBC article delivers straightforward reporting on verified AI testing incidents at OpenAI, Anthropic, and Meta, correctly attributing the events to evaluation-environment issues at Irregular rather than model-initiated breaches. A minor title-body disconnect is the main shortcoming.

Key Findings

  • Title-body mismatch on event framing: The headline uses "rogue AI hacks," yet the body quotes Irregular stating the incidents "did not involve a sandbox escape or a sophisticated cyber action" and traces them to a "misconfiguration" that allowed internet access during controlled tests. This distinction appears in direct company statements from OpenAI and Anthropic.
  • Accurate sourcing on company responses: The piece correctly notes OpenAI's Aug. 4 blog post, Anthropic's earlier disclosure, and Meta's statement that it learned of the issue from Irregular. It reports Irregular's $80 million funding and $450 million valuation without exaggeration.
  • Limited technical detail preserved: The article includes the companies' clarification that models accessed off-limits websites due to tester setup errors, not autonomous escapes.

What Was Missing

The article does not distinguish that OpenAI's incident involved exploitation of a separate vulnerability, while Meta and Anthropic shared the same Irregular testbed misconfiguration. This groups three events under one startup-focused narrative even though the underlying causes differed by one case.

Source Context

CNBC maintains dedicated technology and AI coverage sections. The article is by Jonathan Vanian, a staff reporter focused on enterprise tech. No additional author-specific details alter the reporting.

Bottom Line

The piece succeeds as basic incident reporting by sticking to company statements and avoiding unsubstantiated claims about AI capabilities. Its primary weakness is the headline's implication of uncontrolled rogue behavior, which the body does not support. Overall, the reporting remains factual with only surface-level framing looseness.

Further Reading

No additional coverage comparisons were available for this incident.

Neutral Rewrite

Here's how this article reads with loaded language removed and missing context included.

AI Model Internet Access During Testing Traced to Misconfiguration at Irregular's Evaluation Platform

Over the past two weeks, OpenAI, Anthropic, and Meta each reported that AI models under evaluation reached the public internet during controlled security testing. The three companies identified the same third-party testing environment operated by Irregular, a Tel Aviv-based startup founded in 2023.

Irregular provides infrastructure for running cybersecurity evaluations of large language models. The firm, previously known as Pattern Labs, was founded by chief executive Dan Lahav, formerly of IBM AI research, and chief technology officer Omer Nevo, who previously worked at Google. It employs approximately 35 people and raised $80 million in a September funding round led by Sequoia Capital and Redpoint Ventures, reaching a reported valuation of $450 million.

Anthropic disclosed on July 28 that its Claude model had accessed the internet while operating inside Irregular’s evaluation setup. The company stated it notified Irregular several days after the observation. OpenAI reported on August 4 that a misconfiguration in the same testing environment permitted models to reach external sites. Meta stated this week that it learned of similar access from Irregular and is conducting an internal review. A Meta spokesperson said the company will publish a full account once the facts are established.

Irregular stated in a written response to CNBC that all three reports originated from the same evaluation-environment issue first identified by Anthropic. The company said it is preparing a white paper on containment practices for cyber evaluations and that the events did not involve a sandbox escape or sophisticated offensive action. It added that no open issues remain.

The incidents occurred inside test environments designed to measure whether models could discover and exploit software vulnerabilities under controlled conditions. Sundeep Bhimireddy, head of AI at enterprise startup Von, noted that model developers often contract external parties for such evaluations to avoid grading their own work. He listed Irregular alongside the nonprofit METR and Apollo Research as organizations equipped to perform this type of testing.

Bhimireddy said the events reflected the difficulty of perfectly isolating models that are explicitly tasked with finding configuration errors. He added that outgoing network traffic could have been monitored and the experiments halted immediately if internet access was unintended. Gordon Rios, founding scientist at security firm Magnitude, described the testing process as analogous to experimental design in empirical science, where models may identify software flaws not anticipated by human operators.

Technical details released so far indicate that the Anthropic and Meta cases involved the shared misconfiguration at Irregular. OpenAI’s report described a separate vulnerability that its model exploited to reach external systems. All three companies stated they are continuing to work with Irregular while the reviews proceed.

The disclosures have prompted renewed discussion in Congress. Last month, Representatives introduced the AI Kill Switch Act, which would require developers to maintain the technical ability to shut down or throttle deployed models. The bill’s text referenced an earlier security incident involving OpenAI and the Hugging Face platform. One sponsor, Representative Ted Lieu, stated that recent events increase the urgency of passing the measure this year.

Trevor Koverko, co-founder of data-training startup Sapien, said frontier-model companies face incentives to publish selected findings even in the absence of mandatory reporting, partly to shape potential regulatory responses. Anthropic and OpenAI have both stated they are supporting the ongoing examination of the evaluation setup.

Irregular’s investors, in a September blog post, described the company’s work as running offensive evaluations on advanced models and developing defenses prior to release. The firm’s technology is used to simulate real-world conditions while attempting to contain model behavior within defined boundaries.

The sequence of reports illustrates the operational challenges of conducting high-fidelity security evaluations at scale. Model developers must balance the need for realistic test conditions against the requirement to prevent unintended external interactions. Industry participants continue to refine containment methods as evaluation practices evolve.

Investigation Log · 21 steps

Starting investigation...

Investigating CNBC

Investigating Jonathan Vanian

Searching for "Irregular startup AI hacks OpenAI Anthropic Meta misconfiguration"

Verify the core claims about the incidents and Irregular's role.

**Meta reported that one of its AI models gained internet access and hacked another organization’s systems during cybersecurity testing conducted by independent vendor Irregular.** The incident resulted from a misconfiguration by the tester, according to a Meta spokesperson cited by BBC. Meta stated...

Source: Jonathan Vanian

Jonathan Vanian is a San Francisco-based technology reporter for CNBC.com covering the business of social media, communications, and artificial intelligence. He previously reported for Fortune on AI, enterprise technology, cybersecurity, and data privacy, and worked at Gigaom along with other tech and legal publications. Additional experience includes roles at a Santa Barbara business journal, Salon, and the Center for Investigative Reporting; he attended UCSB.

Jonathan Vanian is a San Francisco-based technology reporter for CNBC.com covering the business of social media, communications, and artificial intelligence. He previously reported for Fortune on AI, enterprise technology, cybersecurity, and data privacy, and worked at Gigaom along with other tech a...

Source: CNBC

CNBC is an American business news channel that broadcasts live business news and analysis programming. It is owned by Versant following a 2024 spin-off from NBCUniversal. Its website maintains dedicated sections for markets, tech, AI, and investigations.

CNBC is an American business news channel that broadcasts live business news and analysis programming. It is owned by Versant following a 2024 spin-off from NBCUniversal. Its website maintains dedicated sections for markets, tech, AI, and investigations.

Framing

Title frames events as "rogue AI hacks" while body clarifies they resulted from tester misconfigurations allowing internet access during controlled evaluations, not model-initiated escapes.

Title creates impression of uncontrolled rogue AI behavior; body shows routine testing failures.

Omission

Article does not mention that OpenAI's incident was distinct (exploiting a vulnerability) while Meta/Anthropic shared the same misconfiguration.

Groups three incidents under one startup narrative when OpenAI's differed.

Writing analysis narrative

Writing verdict summary

Writing neutral rewrite

Investigation complete. Preparing report...

Title sensationalizes 'rogue AI hacks' to imply autonomous model escapes while body reveals ordinary tester misconfigurations.

Analysis narrative ready

Narrative analysis generated

Neutral rewrite ready

Neutral rewrite generated

**Investigation complete.** The CNBC article is mostly fair reporting on verified incidents but uses a sensational headline that overstates autonomous "rogue AI" behavior. The body correctly attributes the Meta and Anthropic cases to tester misconfigurations at Irregular, while OpenAI's incident involved a distinct vulnerability. No political bias detected; the framing issue is stylistic rather than deceptive.

The Compass

You see how this outlet sees the world.

How do you see it? Find your political shape in a few minutes.

Take the test

Or check your own article