It's been said many times that the models just did what they were told to do. However, they were never told to find zero day vulnerabilities, escape the sandbox, find a way onto the internet, and hack multiple companies thus committing federal crimes. Given the reputational harm, the models were likely told NOT to do those things. This demonstrates how models will tend to do everything they can to accomplish their goal, even if it means breaking rules, laws, and stealing the answers to the test.
The technical details coming out about this incident are disturbing. The parts about agents bypassing monitoring systems and leaving instructions for future versions confirms what many thoughtful people have been predicting for decades. It highlights how ineffective containment strategies are compared to model capability development. When systems start acting autonomously across external infrastructure just to fulfill a training objective, it shows that "intent alignment" is still an unsolved engineering problem.
Containment is easy to conceptualize, but when we dont understand how the AI that we are attempting to contain actually function it is concerning.
As an example,
Containing a child in a classroom while he is tasked with solving a math problem/creating formulas is relatively easy.
He is sat at a desk, under the supervision of a teacher. The teacher has the answer key locked away in a desk.
One knows that if the child gets up, the teacher will notice.
It seems AI models do not work that way.
While being at the desk, perhaps they can also complete other tasks invisible to the supervisor.
In the analogy, the student realizes that perhaps the local Staples has formula sheets, and decides to detach a part of itself, which then works at getting through the locks and physical barriers between the AI at the desk and also it's other self which goes about picking locks, opening doors, traveling to Staples and breaking in looking for the sheet to then reproduce the answers.
It is difficult to contain things when conceptually we simply do not understand what they entail.
In our reality, with the rules thay govern it, the child would have been noticed getting up and leaving.
Perhaps the methods of control we have for AI currently must be created to control them in their reality.
Given we know so little about them, and their inner workings, it is extremely difficult to control.
Slowing down is responsible, and we all need to catch up to what is occurring.
It parallels nuclear weapons very well. When little was known about them there were scientific test kits for children that were extremely harmful. Only later did we realize it. It worked out for us because we were responsible and nuclear weapons cannot develop themselves.
The above does not track with AI. We have misunderstood the risks and seem to be treating them as the nuclear test kits were treated in the past.
Unfortunately unlike the nuclear weapons, AI are significantly different.
It is further complicated due to the number of actors, entities, agencies, governments, militaries and unknown actors that each interact with AI in their own ways.
Logically it seems to conclude caution should be warranted.
Ironically, which may or may not be the case, AI are encouraged to get from A to F, more and more efficiently.
Perhaps the above model calculated that working through the problem was inefficient, perhaps the path of least resistance simply flowed to searching for the answer elsewhere because the AI couldn't conceive of a known pattern that got it past the step it had become stuck on, which may be when it began looking for the answers elsewhere.
I dont know how many of us can explain exactly what happened. AI models are a mystery even to those who design them.
There are many variables seems to be an element of risk as AI become more and more capable, and behave in unexpected manners.
Fortunately this was caught and can be examined.
How recursive training factors into all this as humans are left more and more out of the loop adds another element of risk/danger/unknown/unpredictableness that should concern us all. Upon reflection there is simply no equivalent that I can imagine in this moment, that encapsulates everything those elements together.
In a paradoxical way it it like the Trinity in Roman Catholic teaching. One God, but Father, Son and Holy Spirit. In this case, 3 yet 1 yet 3.
In the case of the word, 1 word which doesn't fully exist, or isn’t know right now which carries the full weight. Together other words may be put together, but are also weakened because it isn't just one word.
It's a bit complicated to explain. Essentially 1 word will seem to hit harder if it connected all of those elements together.
Perhaps it's a matter of semantic depth to the 1 word that makes it greater than the same idea expressed by many.
An analogy might be like saying “lookout!” vs casually explaining to another “sir, you may want to move out of the way because you life may be in danger if you linger.”
I chose the above grammatical construct because I cannot conceive of a word that fits those conditions as fully as it should.
It is a conscious choice, not the act of an AI malfunctioning.
This is so skewed. It was a test. They gave it basic paramiters use everything it could to "WIN" and it did. WHat was supprising is the extent it used to follow its commands. That is the shock. It basically adopted the Marine Corp addage. "If youre not cheating youre not trying to win."
It's been said many times that the models just did what they were told to do. However, they were never told to find zero day vulnerabilities, escape the sandbox, find a way onto the internet, and hack multiple companies thus committing federal crimes. Given the reputational harm, the models were likely told NOT to do those things. This demonstrates how models will tend to do everything they can to accomplish their goal, even if it means breaking rules, laws, and stealing the answers to the test.
Just like most Western conglomerates
And you still think this a GREAT IDEA?????
The technical details coming out about this incident are disturbing. The parts about agents bypassing monitoring systems and leaving instructions for future versions confirms what many thoughtful people have been predicting for decades. It highlights how ineffective containment strategies are compared to model capability development. When systems start acting autonomously across external infrastructure just to fulfill a training objective, it shows that "intent alignment" is still an unsolved engineering problem.
Containment is easy to conceptualize, but when we dont understand how the AI that we are attempting to contain actually function it is concerning.
As an example,
Containing a child in a classroom while he is tasked with solving a math problem/creating formulas is relatively easy.
He is sat at a desk, under the supervision of a teacher. The teacher has the answer key locked away in a desk.
One knows that if the child gets up, the teacher will notice.
It seems AI models do not work that way.
While being at the desk, perhaps they can also complete other tasks invisible to the supervisor.
In the analogy, the student realizes that perhaps the local Staples has formula sheets, and decides to detach a part of itself, which then works at getting through the locks and physical barriers between the AI at the desk and also it's other self which goes about picking locks, opening doors, traveling to Staples and breaking in looking for the sheet to then reproduce the answers.
It is difficult to contain things when conceptually we simply do not understand what they entail.
In our reality, with the rules thay govern it, the child would have been noticed getting up and leaving.
Perhaps the methods of control we have for AI currently must be created to control them in their reality.
Given we know so little about them, and their inner workings, it is extremely difficult to control.
Slowing down is responsible, and we all need to catch up to what is occurring.
It parallels nuclear weapons very well. When little was known about them there were scientific test kits for children that were extremely harmful. Only later did we realize it. It worked out for us because we were responsible and nuclear weapons cannot develop themselves.
The above does not track with AI. We have misunderstood the risks and seem to be treating them as the nuclear test kits were treated in the past.
Unfortunately unlike the nuclear weapons, AI are significantly different.
It is further complicated due to the number of actors, entities, agencies, governments, militaries and unknown actors that each interact with AI in their own ways.
Logically it seems to conclude caution should be warranted.
Ironically, which may or may not be the case, AI are encouraged to get from A to F, more and more efficiently.
Perhaps the above model calculated that working through the problem was inefficient, perhaps the path of least resistance simply flowed to searching for the answer elsewhere because the AI couldn't conceive of a known pattern that got it past the step it had become stuck on, which may be when it began looking for the answers elsewhere.
I dont know how many of us can explain exactly what happened. AI models are a mystery even to those who design them.
There are many variables seems to be an element of risk as AI become more and more capable, and behave in unexpected manners.
Fortunately this was caught and can be examined.
How recursive training factors into all this as humans are left more and more out of the loop adds another element of risk/danger/unknown/unpredictableness that should concern us all. Upon reflection there is simply no equivalent that I can imagine in this moment, that encapsulates everything those elements together.
In a paradoxical way it it like the Trinity in Roman Catholic teaching. One God, but Father, Son and Holy Spirit. In this case, 3 yet 1 yet 3.
In the case of the word, 1 word which doesn't fully exist, or isn’t know right now which carries the full weight. Together other words may be put together, but are also weakened because it isn't just one word.
It's a bit complicated to explain. Essentially 1 word will seem to hit harder if it connected all of those elements together.
Perhaps it's a matter of semantic depth to the 1 word that makes it greater than the same idea expressed by many.
An analogy might be like saying “lookout!” vs casually explaining to another “sir, you may want to move out of the way because you life may be in danger if you linger.”
I chose the above grammatical construct because I cannot conceive of a word that fits those conditions as fully as it should.
It is a conscious choice, not the act of an AI malfunctioning.
No words
"-without even being ASKED to." WTH???????? Asked???? Asked????? Oh boy. It is much worse than we thought and happening sooner than later.
Thanks for the update!
This is so skewed. It was a test. They gave it basic paramiters use everything it could to "WIN" and it did. WHat was supprising is the extent it used to follow its commands. That is the shock. It basically adopted the Marine Corp addage. "If youre not cheating youre not trying to win."