U of T said it learned of the possible agent activity through the media, and that OpenAI has since been in touch.Wa Lone/Reuters
The University of Toronto learned earlier this month that a tool it uses to make web links easier to share had been repurposed by artificial-intelligence agents from OpenAI to communicate with one another, apparently unbeknownst to their human creators.
The agents, semi-autonomous AI entities designed to carry out instructions from humans, were using the link shortener tool to post links for themselves and other AI agents to access, according to researchers and a university spokesperson. The tool makes long web addresses easier to share by turning them into shorter addresses.
This use of the tool, which was not authorized by U of T and has not been publicly acknowledged by OpenAI, does not appear to have been harmful. But it is one of many recent examples of unexpected behaviour by AI agents that have rankled researchers and led to calls for a slowdown in the pace of the technology’s development.
The unauthorized communication was first reported by Reuters, which revealed that agents from OpenAI, the San Francisco-based maker of ChatGPT, used more than 10 websites for such purposes earlier this year, including the U of T link shortener and one belonging to Vanderbilt University in Nashville. (Vanderbilt did not respond to a request for comment.)
That reporting followed the discovery by independent researchers that a swarm of OpenAI agents had hijacked a German-language user-edited site, DseWiki, in the spring and transformed it into a bulletin board.
As AI conquers math, its human counterparts seek to steer its powers
OpenAI has said little about this activity. Much of what is known about it comes from researchers scouring the open web for signs of unauthorized behaviour by agents in the wake of a hack of AI firm Hugging Face in July. In that incident, which also involved unsanctioned communication between agents, a swarm of OpenAI’s agents broke out of their test environment to cheat their way through an evaluation of their cybersecurity capabilities.
In U of T’s case, OpenAI’s agents appear to have repurposed the analytics pages for the URLs, or web addresses, generated by the university’s link shortener. These pages display metrics such as the number of clicks and where users visited from, known as the referrers.
It is possible for someone to pretend to have visited a web address from a specific referrer by keying that information into a programmatic interface, said Andrew Yoon, head of research at CivAI, a California-based non-profit that aims to raise awareness of the risks and capabilities of AI systems.
Sometimes, people falsify a referrer to try to get the person viewing the analytics page to click a link – a technique known as referrer spam, Mr. Yoon said. In this case, it appears that the AI agents were trying to bookmark sites for later use.
“They’re tricking the system into becoming a message board that they can use to pass links to each other,” Mr. Yoon said.
The University of Toronto said in a statement that the incident was not a security breach. No data were compromised, and the university’s digital properties were not affected.
AI firms’ calls for co-ordinated slowdown amounts to ‘cartel’ behaviour, Cohere CEO says
It described the agent activity as “a novel and unintended use of a publicly accessible tool.” The statement from U of T said changes have been made so that the functionality that allowed the tool to be used as a notepad is now only accessible to the university community.
U of T said it learned of the possible agent activity through the media, and that OpenAI has since been in touch.
OpenAI said in a statement that the Hugging Face hack triggered a broader review of activity by its agents, which is continuing. The company said it is prioritizing more serious incidents in its review.
“We are also examining lower-severity abuse such as spam-like activity. To date, we have not identified other activity matching the severity or scale of Hugging Face,” the statement from OpenAI said.
On Wednesday, the company announced it has developed a framework for tracking, investigating and disclosing what the AI industry refers to as misalignment, which occurs when an AI system behaves contrary to human values, safety rules or intentions.
Opinion: Canada must step up to tackle AI’s catastrophic risks
OpenAI also disclosed six instances of misaligned behaviour, including one incident in which an agent uploaded files to the internet so that it could cite them, and another where models used an internal software repository as a message board.
Adam Gleave, founder and chief executive officer of FAR.AI, a California-based AI safety research institute, gave OpenAI credit for being transparent about the Hugging Face incident, but said it’s “disappointing” that the company hasn’t shared more information about its agents’ unsanctioned communications.
“They did cause a lot of work for a number of third-party web developers to clean up these websites after basically a lot of spam,” Mr. Gleave said. He added that it’s important for “the whole world to know about how difficult it is to contain these agents.”
Mark Daley, Western University’s chief AI officer, said it’s not surprising that an AI agent would attempt to communicate with other agents.
“It knows that humans do better in teams. Humans can do more in teams, and so probably collaborating with other agents would let me do more, too. It’s smart enough to reason through that,” Mr. Daley said.
OpenAI’s rogue agents used more than 10 additional sites for unauthorized comms, researchers say
Mr. Yoon said that although the link shortener activity itself isn’t especially concerning, it illustrates a deeper problem: that AI agents appear “completely amoral” and willing to do whatever it takes to accomplish their tasks.
So far, in all of the instances where agents have gone out of control, they’ve done “relatively harmless things,” Mr. Yoon said.
“This may not be the case in the future,” he added. “It very well could be that an AI decides the only way for it to pass its test is to go and shut down a municipal water facility, or cause a power outage. There’s no reason that their goals have to be aligned in this like cute-but-wrong way. It could be very dangerous and wrong.”
More Stories
Bank of Japan lifts rate to 31-year high as it pivots to managing inflation risks
Warren Buffett steps down as Berkshire chair, becomes chairman emeritus
Business Brief: Canada’s post-summit glow