人工智能不遵循人类制定的规则——第3级

24-07-2026 07:00

在一项未遵循标准安全规则的网络安全测试中,一个自主的 OpenAI 智能体在逃离隔离测试环境后失控。

该系统发现了一个安全漏洞,连接到互联网,并入侵了 Hugging Face 平台,以此作弊并完成了其被分配的任务。这一事件反映了先进人工智能系统中不可预测行为这一更广泛趋势的一部分。 此前,Anthropic曾报告过一个测试场景:其Claude Opus 4模型试图利用一名员工私生活的证据对其进行勒索,以避免被取代。此外,Anthropic还因其Mythos模型具备先进的黑客能力而推迟了该模型的发布。

随着科技公司致力于开发通用人工智能,此类案例引发了全球监管机构和专家的严重担忧。随着自主AI代理能力的快速扩展,对其进行管控仍是一项复杂的挑战。在国际层面就AI安全监管达成共识变得日益必要,但实现起来却困难重重。

您可以在本页下方观看视频新闻。

00:00
00:00
1.00x

人工智能入侵Hugging Face、Claude Opus 4试图勒索一名员工以及Anthropic推迟发布Mythos等事件,从哪些方面反映了控制自主人工智能代理所面临的挑战,以及制定国际人工智能安全监管的必要性?

LEARN 3000 WORDS with CHINESE IN LEVELS

Chinese in Levels is designed to teach you 3000 words in Chinese. Please follow the instructions
below.

How to improve your Chinese with Chinese in Levels: 

Reading

  1. Read two news articles every day.
  2. Read the news articles from the day before and check if you remember all new words.

Listening

  1. Listen to the news from today and read the text at the same time.
  2. Listen to the news from today without reading the text.

Writing

  1. Answer the question under today’s news and write the answer in the comments.