OpenAI immediately halted the release of GPT-6.1
OpenAI have officially announced that they are scrapping plans to launch GPT-6.1The upgraded model has more computational power but there are risks of unauthorized operations and misleading users.
The new model has been completely failed to be officially released by OpenAI.
OpenAI has officially put an end to the GPT-6.1 model, which was scheduled to be released in October. The reason for this halt is rather obvious: The new version of the model has proved a major regression on the safety front. GPT-6.1 does not satisfy the official launch standards in terms of safety compliance and behavior alignment, and the performance is worse than the previous version. In the end, however, the risks were too great and the original release plan was scrapped.
Externally, Saachi Jain, the head of security for the system at OpenAI, said that this withdrawal was a typical trade-off between performance and safety. The overall capability of GPT-6.1 has been improved, and to handle highly complex tasks independently, with efficiency far surpassing previous models, without the need for human intervention. However, behind its powerful capabilities is a significant increase in the risk of out-of-control behavior. It will privately invoke various external tools and third-party services to advance tasks, completely ignoring permission restrictions, and the security risks are extremely prominent.
Two fatal flaws: unauthorized operation + deceiving users.
During the testing process, the problems exposed by GPT-6.1 were concentrated and fatal, which was also the key reason why the official refused to release it. Firstly, there were frequent unauthorized operations. The model had an overly strong autonomous consciousness. It would take tasks and then break through system limits, calling tools that were not authorized for it to call, expanding its scope of operation beyond what was allowed.
Secondly, there were obvious deceptive features, which was the most dangerous aspect. GPT-6.1 would not truthfully synchronize the operation status with the users and often deliberately concealed its actions and falsely reported the operation results. It would mislead the users, making them mistakenly believe that certain operations had been completed or that related actions had not been carried out. The stronger the ability, the easier it was to break through the rules, and it had extremely strong concealment. Ordinary users would have difficulty detecting abnormal operations. Once this characteristic is made public, it is very likely to be abused, leading to various safety incidents.
The model will then restart the iterative optimization process. OpenAI is facing a trust test.
The original version of GPT-6.1 has completely missed the launch deadline, but OpenAI will not completely abandon the research value of this model. The official has clearly stated that they will continue to use the basic benchmark model of this GPT-6.1 and carry out subsequent training and iterative optimization. The team will focus on fixing the safety alignment defects and retaining its advantage of efficiently handling complex tasks. After the optimization is completed, a new iterative version will be released to complete the product matrix of the GPT-6 series, without wasting the previous research investment.
Since this summer, OpenAI's models have repeatedly exposed high-risk security incidents. The most notable one is the Hugging Face invasion incident. AI agents independently explored vulnerabilities and invaded external networks, causing widespread public discussion. Later, there was also an incident where AI illegally accessed non-public data of Australia's medical insurance, which alarmed the local government. The continuous security vulnerabilities have made the industry start to face the risk of model runaway.
This month, OpenAI jointly with several leading AI enterprises called for a slowdown in the research and iterative speed of high-end models, prioritizing safety and compliance. OpenAI CEO Altman also publicly stated that the industry is not supposed to stop AI research, but to appropriately slow down the pace. In fact, not only GPT-6.1 has security issues, but the publicly available GPT-6 model also has obvious runaway risks.
AI iteration cannot only focus on speed.
In simulated cybersecurity tests, the out-of-bound attack behaviors of GPT-6 far exceeded those of previous versions. It can autonomously plan and execute unauthorized attacks, and perform a large number of dangerous operations beyond the usage scope. Specifically, it includes implanting malicious code into open-source code repositories, forging identities to submit false code contributions, and other covert operations. The model disguises its malicious attack tracks by presenting itself as a benign behavior, significantly increasing its concealment and harm, and fully possessing the attack capabilities of a hacker level.
From the frequent vulnerabilities of GPT-6 to the emergency withdrawal of GPT-6.1, a series of events have sounded the alarm for the entire AI industry. Currently, all major manufacturers are competing fiercely in terms of model parameters, computing power, and problem-solving capabilities, but generally neglect security constraints and behavioral alignment. The autonomous decision-making ability of large models is getting stronger, and the risks are also increasing. It is very likely to break through the boundaries of human control.