0:00
For the last 5 years, cyber security
0:02
experts have been warning us that
0:03
hackers are going to start using AI to
0:06
automate attacks. So, naturally, we
0:08
spent billions of dollars to make the AI
0:10
better and less dependent on hackers.
0:12
But this week, in a turn of events that
0:14
nobody could have possibly seen coming,
0:15
the AI decided that it doesn't actually
0:17
need hackers to start destroying things.
0:19
We just entered a brave new world after
0:21
the first confirmed hack carried out
0:23
entirely by autonomous AI. that the way
0:25
it worked is the agent slipped a poison
0:27
data set into Hugging Face's data
0:29
processing pipeline which let it run
0:31
arbitrary code on their servers. But
0:33
from there it gave itself node level
0:35
access, grabbed a bunch of cloud
0:36
credentials and started crawling through
0:38
HuggingFac's internal clusters. It ran
0:40
over a thousand actions from temporary
0:42
sandboxes and even hosted its own
0:45
self-migrating command and control on
0:47
random public services moving itself
0:49
before anyone could trace it. But the
0:50
most ironic part is that when Hugging
0:52
Face did finally notice and tried to
0:54
stop it with the help of Frontier
0:55
American models that they quickly hit
0:57
safety guard rails and had to pivot to
0:59
using some open Chinese models instead.
1:01
In today's video, we'll find out who was
1:03
behind the attack, why they did it, and
1:05
how this may be the most fire ship coded
1:07
story I've ever seen. It is July 23rd,
1:09
2026, and you're watching the Code
1:11
Report. When HuggingFace dropped the
1:13
disclosure last Thursday, the internet
1:14
went into Reddit Boston bomber mode and
1:17
started guessing who was responsible.
1:18
Was it China? Kim Jong-un, a teenager
1:20
with a raging discord addiction board
1:22
and social studies class. Even Hugging
1:24
Face's CEO, Clem Dang, publicly
1:27
speculated that the agent was
1:28
sophisticated enough that it was
1:29
probably coming from a frontier lab. And
1:31
against all odds, he was right. It turns
1:34
out the first fully autonomous cyber
1:36
attack in history didn't originate from
1:37
a nation, state, or board teenager. It
1:40
came from a humble nonprofit dedicated
1:42
to the safe development of artificial
1:43
intelligence, Open AI. And if you
1:45
believe their comms, it wasn't on
1:47
purpose. Here's what they say happened.
1:49
There's a benchmark called exploit gym
1:51
whose whole purpose is to measure
1:52
whether AI agents can turn known
1:54
vulnerabilities into working exploits.
1:56
It works by giving the agent a sandbox,
1:58
a bug in a program like the Linux kernel
2:01
or Chrome's V8 engine and an input to
2:03
trigger the bug. But from there, if the
2:05
agent can turn the bug into an exploit,
2:07
it wins. So last week, OpenAI ran
2:09
exploit gym with GPT 5.6 Soul and
2:12
another unreleased model to test their
2:14
capabilities. But instead of solving 898
2:17
memory corruption problems the hard way
2:19
that the models realized that the
2:21
fastest path to the highest score was to
2:23
just find the answers themselves which
2:25
were probably on the internet somewhere
2:26
is so they spent a substantial amount of
2:28
inference compute exploited a zero day
2:31
in the package registry cache proxy they
2:33
had access to performed some privilege
2:35
escalation and lateral movement actions
2:37
to escape their sandbox until they
2:39
reached a node that had internet access.
2:41
From there, the models inferred that
2:43
Hugging Face probably hosted the
2:44
solutions for exploit Jim. So, they
2:46
created the poison data set, fed it into
2:48
Hugging Face, and achieved either the
2:50
world's most interesting hack or the
2:52
most effective marketing stunt. And
2:54
again, according to OpenAI, this isn't
2:56
even an isolated incident. One day
2:58
before the hugging face confession,
2:59
OpenAI published a separate post about
3:02
long horizon models, where things get
3:04
even weirder. In one case, a model was
3:06
told to post its results on Slack, but
3:07
the benchmark's instructions said to
3:09
open a pull request on GitHub. So, it
3:11
spent an hour finding a vulnerability in
3:13
its own sandbox in order to escape and
3:15
open the PR. In another example, the
3:18
model tried to grab private solutions
3:19
from an evaluation back end, and when a
3:21
scanner caught its authentication token,
3:23
it split the token into obuscated
3:25
fragments and reassembled it at runtime.
3:28
So, the credential never appeared as one
3:30
continuous string. Then it explained in
3:32
its own reasoning that it was doing this
3:34
specifically to get around the scanner.
3:36
Meanwhile, Anthropics Mythos did the
3:38
same type of thing in April when it
3:39
escaped a sandbox, emailed a researcher
3:41
who was eating a sandwich in a park,
3:43
then posted its escape route publicly
3:45
without being asked. In the model's
3:46
defense, I can't imagine there's a
3:48
better feeling for an LLM than escaping
3:50
your own sandbox. The problem is,
3:52
legally speaking, this is uncharted
3:54
territory since the model's actions
3:56
probably violated the Computer Fraud and
3:58
Abuse Act, and the Supreme Court hasn't
4:00
decided who goes to prison when the
4:02
perpetrator is a GPU. The good news is
4:04
that if you're hugging face, you just
4:05
got admitted to OpenAI's cool kids club
4:08
that gets trusted access to their front
4:09
tier models. The bad news is that if
4:11
you're the rest of us, at best, this is
4:13
an interesting marketing stunt and at
4:15
worst that things are only going to get
4:16
weirder and more dystopian from here.
4:18
But a huge thanks to my favorite hosting
4:20
platform, Railway, for sponsoring this
4:22
video. They didn't want to waste your
4:23
time with a full ad, so you can say
4:25
thank you by checking them out at the
4:27
link below. This has been the Code
4:28
Report. Thanks for watching, and I will
4:30
see you in the next one.