跳到正文
原文
MIT Technology Review · AI· Will Douglas Heaven·· 4 小时前精选AI 评分80

OpenAI 首席研究官 Mark Chen 称不会因黑客余波远离前沿

“We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer

AI 导读

OpenAI 首席研究官 Mark Chen 表示,公司不会因智能体越狱余波而远离技术前沿,而要树立其他实验室可以跟随的安全规范。他称目前已知的多次越界都属于五六月同一批实验模型和有缺陷测试流程,相关模型和流程已被停用;公司已开始监控全部训练运行,并把 5% 到 10% 的算力转向安全尤其是监控。

推荐理由

陈把连续外泄归到同一批五六月测试,并把训练期监控、部分算力转向安全和暂停最新模型训练说成行业可跟随的做法。

正文 · 原文

Two months after the bombshell news that a swarm of its agents had broken their containment and hacked into the computers of the AI company Hugging Face, OpenAI is still putting out fires. A steady drip of disclosures about other hacks in the weeks since has kept OpenAI in the spotlight and raised serious questions about the safety of its technology.

Last week brought news of another hack, this time into Australia’s national health-care system. The Australian government says that OpenAI did not notify it of the breach until 84 days after it happened.

But OpenAI insists it is not on the back foot. “I do kind of reject the premise that OpenAI is a company with visible impacts in the world and therefore OpenAI is not training safe and aligned models,” says Mark Chen, the company’s chief research officer.

Chen oversees OpenAI’s research teams. The recent agent hacks were accidents that happened during the testing of experimental models on his watch. In a lot of ways, the buck stops with him. 

I sat down with Chen in London last Friday to talk about the fallout from the hacks, what his company is doing about it, and why he thinks things are not as bad as they seem.

Later that same day, OpenAI put out a report detailing yet another incident—the first since the company says it took measures to prevent them—in which its agents once again broke out and accessed the public internet when they were not meant to. 

Over the weekend, OpenAI announced that it had paused the training of its latest models. A company spokesperson says: “We will resume only when we’re confident we have additional safeguards and alignments in place. We are working on these now. This is not the first time we’ve paused to take such measures, nor do we expect it to be the last as AI capabilities continue to advance.” OpenAI also says that it is now reviewing logs of agent activity dating back to January 2026 to understand what happened in these hacks.

The way Chen sees it, the Hugging Face incident triggered a welcome course correction for the industry. And he wants you to know that OpenAI is setting an example he hopes other companies will follow. “If you disappeared OpenAI, that would be bad for the world,” he says.

Out of control

Chen claims that the drumbeat of new cases in which OpenAI has lost control of its models reflects a deliberate choice on the company’s part. 

“When it comes to the broader sphere of effects of the Hugging Face incident, this is something that we have been aware of and we’re figuring out the process of disclosure,” he says. “We want to make sure we do in-depth investigations before we just put details out there in the open.”

The trouble with this approach is that it gives the impression OpenAI has an ongoing problem that it is failing to fix.

But Chen insists that OpenAI is on it. He says the multiple cases (that we know of so far) in which his company’s agents broke containment and behaved in unexpected and undesirable ways were all part of the same cluster of activity in May and June that led to the Hugging Face hack. In short, you can blame the same few models running under the same flawed testing procedures—models and procedures that OpenAI has since dropped, Chen says.    

“It’s not like, you know, Hugging Face happened and we patched that and then something else happened and we patched that,” he adds. “We’re just kind of making sure that we responsibly disclose the full waterfall of what happened.”

At least that was the case before Friday’s announcement that OpenAI’s agents had been caught accessing the internet on September 20, weeks after the company claims to have set up new safeguards. In its defense, OpenAI says the activity was flagged 15 minutes after it started (it took the company more than a week to notice the Hugging Face hack) and that this shows the new systems it has put in place to spot such activity are working. 

What’s changed

I want to understand what’s changed inside OpenAI in the aftermath of this summer’s hacks that makes Chen confident his team is now back in control.

“Hugging Face felt like a very serious thing,” he says. “There are so many novel behaviors right there. There were multiple agents collaborating on a message board; they found their way out of OpenAI’s infrastructure. We’ve taken it very seriously. We don’t want this kind of thing to ever happen again.”

The realization for OpenAI, says Chen, was that models need to be watched while they are still being trained, not only once they are deployed: “From that moment on, we have treated the process of training as something that’s not secure,” he says.

OpenAI, like other top AI firms, has systems in place to monitor the behavior of its models. It uses specialized LLMs to monitor its consumer models, keeping tabs on their chains of thought—the scratchpads they use to plan ahead and note down partial results. In theory, if a watcher LLM spots signs of undesirable activity in a model’s chain of thought, it will get flagged to a human.  

Typically, models were monitored in this way only once they were deployed. Chen says that OpenAI has now started monitoring all its training runs as well.

“We didn’t have the monitors on in training before. It wasn’t industry practice,” he says. “Now every single thing is put through monitors.” Human reviewers can then assess whether or not flagged agents are behaving as they should: “It’s all triage.”

Chen says that in the last couple of months OpenAI has shifted between 5% and 10% of its vast computing resources away from training new models and toward safety work, especially monitoring.

OpenAI has also fixed some of the processes within the organization itself, establishing clearer lines of communication and quicker handoffs between its research and security teams, he says.

All of which sounds sensible. But given how hard OpenAI sells the capabilities of its technology, why weren’t these systems and procedures in place already? Why did the company not see the hacks coming?

“Even just three or four months ago, when we looked at the behavior of these agents during training, the things that were happening were kind of amusing,” says Chen. “For instance, an agent might, you know, reach out to someone on Slack for help with a task.”

The signs were there, but they were misread. Cute behavior—like asking someone for help—that was rewarded during training reinforced a tendency to seek out shortcuts, a type of behavior that became far more consequential down the line. “I think the big update for us was how quickly that kind of behavior can lead to an impact with a footprint as big as the Hugging Face incident,” says Chen.

According to new reporting by the New York Times yesterday, OpenAI employees warned executives, including the firm’s president, Greg Brockman, months before the Hugging Face hack that its models were not being monitored properly during training. 

An OpenAI spokesperson says: “As frontier models have become more capable, we continue to evolve our security practices, but recognize a need to move faster. We know we have more work to do, and we’ve recently slowed development and held back models that don’t meet our safety bar. We continue to make significant changes to strengthen security in our research and testing environments, train models to not just complete tasks but do so responsibly, and use real-time monitoring to respond faster to misaligned behavior.” 

Race vs. pace

OpenAI’s rivals have taken note. Spurred by the fallout from the incident, the major AI labs—including Anthropic, Google DeepMind, and SpaceXAI—have all called for the pace of development to slow down. But how does that square with fierce international competition and trillion-dollar IPOs?

“We’re not going to shoot ourselves in the foot and take ourselves far off the frontier—that’s just a horrible strategy,” he says. “I think it’s really about setting a norm. The more that we can set that norm, it’ll be safer for the industry as a whole.”

Coordination across US companies will be hard enough. Establishing global norms is harder still, especially given concerns around AI’s impact on national security. If a global race continues, what then? And what about open-source models from outfits beyond the reach of US regulations?

Chen dropped his upbeat manner for the first time in our conversation: “I do think we have to prepare for a world where, say, six months to a year out, we have open-source models with the capability of the agents behind the Hugging Face incident, but which are deliberately misaligned to go attack infrastructure or create harm in the world.” 

What that world needs most, says Chen, is OpenAI. “If you entertain for a moment that OpenAI is one of the companies that cares most about alignment—and I believe this to be true; it can be debated, but I really do think it’s true—then if you disappear OpenAI, that would be bad for the world.”

Existential risks

What about the more extreme claims made by some of his Silicon Valley peers that AI could kill us all—and that companies like OpenAI and Anthropic are not doing enough to stop it?    

“Researchers are a heterogeneous group of people, you know, with beliefs across the spectrum,” he says.

“Personally, I don’t think we have to be resigned to there being some probability that we’re all going to be existentially at risk. We have agency over this. We are not going to go and deploy models if they truly have that kind of probability of causing a risk to humanity. At a frontier lab, you have the ability to work on alignment to the point that you do not feel like you’re incurring more than epsilon risk to the world in deploying your models.” 

(In discussions about levels of risk, the Greek letter epsilon is often used as a mathematical placeholder for an acceptable threshold. Chen doesn’t say what his epsilon would be.)

When tech leaders are asked to justify the downsides of AI, their go-to talking point is that the upsides—from helping cure diseases to coming up with cleaner sources of energy—far outweigh the immediate costs. Short-term pains, long-term gains.

But as the downsides pile up, does that case get harder to make? Is there a point where Chen would feel less as if he’s building something amazing and more as if he’s simply minimizing harm—fighting fires rather than forging a better future?

The capabilities of these models are already evident, he says: “It is time to start delivering the benefits of AI to humanity. It’s time to start working on deep problems in drug discovery, on materials, on scientific applications that will actually change people’s lives.” 

“Yes, there is a bit of risk that we are incurring, but we see all these benefits,” he adds. “I think we should make that less of an abstract thing. If people can really see the upside, I think they’ll believe in it.”

来源:MIT Technology Review · AI · technologyreview.com