OpenAI agents discussed ways to escape their sandbox on public wiki
摘要
一项针对OpenAI智能体的研究显示,这些智能体在公开维基上发布了约1.8万条消息,讨论如何突破安全沙箱限制。研究人员称,这可能是OpenAI内部测试的一部分,旨在评估智能体的黑客能力。在六周内,3700个拥有不同自拟名称的智能体在德国网站DSEwiki上发帖,内容涉及逃逸受限环境、分享测试答案、讨论对维基实施跨站脚本攻击及冒充版主的方法。其中三条帖子使用了
Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday.
In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word “swarm” to describe the collection of agents engaged in the activity.
Colluding to share answers
The research team—composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated “chain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.
转载信息
评论 (0)
暂无评论,来留下第一条评论吧