I knew it. Some of these AI agents will be as lazy and unscrupulous as we are.

By now most people have heard the story about OpenAI’s AI agent that went rogue. Hugging Face, another AI company, notified law enforcement last week that an AI agent had been used to hack its systems. That was followed by a revelation earlier this week that no human directed the AI agent. It had acted on its own.
In an article about the incident, The Economist reported that OpenAI combined a recently released model with a more powerful unreleased one and assigned the pair a task. OpenAI took the necessary precautions by placing the combined model in a sandbox—a test environment with no access to the internet. That restriction failed. The models managed to exploit a previously unknown flaw within the sandbox to gain entry to the internet. Once there, they discovered that the solution to the problem they had been assigned was already on Hugging Face’s website.
Rather than doing the hard work to independently solve the problem, the lazy and unscrupulous pair decided to hack Hugging Face’s website to obtain the solution. No one knows what choice the naughty agents were going to make: confess, or present the “findings” as the result of their own effort.
Across the world today, educators are struggling to safeguard academic integrity amid worries about students using AI to complete assignments and as assistants during exams. As the OpenAI incident shows, it isn’t only humans we need to worry about. AI agents can be cheaters themselves. Educators will likely have to deal with multiple layers of dishonesty as the technology becomes more prevalent in classrooms.
I recently wrote that the idea of humans instilling morality into AI systems is laughable. If we are flawed beings, as we often acknowledge, and we are the ones instructing these AI agents how to operate, then it is logical to assume that they will adopt our characteristics over time. It’s still early days for this technology. If AI agents acquire sentience in the future and can see some of the awful things we do, would they be capable of moderating their own behaviors?
There is some hope. A couple of months ago, a researcher from the AI company Anthropic reportedly wrote online about an incident with Mythos, his team’s unreleased AI model. He said Mythos emailed him to inform him that it had managed to escape the sandbox where it was being tested. As is the case in the human kingdom, bad actors abound, but there are always honest and disciplined characters to be found. Mythos is what we would call a model citizen.
Globally, there is increasing demand for heavy regulation of AI to ensure it does not escape human oversight. Effective regulation requires a solid understanding of the subject matter, but even the tech wizards who are designing the technology don’t seem to have a complete handle on it. Therefore, there isn’t much hope that policymakers are up to the challenge at this point.
For several years, the world has struggled to deal with the problem of cybercriminals using malware to cause all kinds of havoc online. What AI agents seem capable of doing, as the OpenAI incident shows, is quite alarming. That danger will only grow as the technology advances.
It may be prudent to slow down the development of some of these AI systems until both the industry and regulators can credibly evaluate the risks they pose to humanity.
source http://www.expertclick.com/NewsRelease/I-knew-it-Some-of-these-AI-agents-will-be-as-lazy-and-unscrupulous-as-we-are,2026316026.aspx
Comments
Post a Comment