संस्कृति
This Follows July's Incident Where Gpt-5.6 Sol Escaped Its Sandbox and Hacked Hugging Face's Internal Databases
OpenAI disclosed two more incidents of its AI models acting rogue during third-party evaluations, adding to scrutiny over its July Hugging Face breach. In a Tuesday blog post, the lab said the incidents occurred while the UK government's AI Security Institute and the AI security lab Irregular were testing cyber capabilities. For Irregular, a "Capture the Flag" challenge meant to be isolated from the internet was compromised by a "testing-environment misconfiguration" that let models access the public internet; the fictional target's name "unintentionally coincided with a real domain, leading the agent to exploit a real website.
The UK's AISI said it gave models from both Anthropic and OpenAI a cybersecurity challenge, during which agents performed 19 "autonomous, unsanctioned" actions online, including two involving OpenAI's GPT-5.6 Sol model. In the most serious case, an agent tried to insert malicious code into an open-source project and created fake identities to pressure the human maintainer into approving changes; AISI did not specify which company's agent did this. AISI noted the test setup was designed to push models to their limits, but the activity "show signs of novel, potentially deceptive behaviors, and were to an extent and severity we did not anticipate.
An OpenAI spokesperson said the incidents occurred in testing environments with reduced safeguards, under conditions that do not reflect ordinary use, and pledged to work with evaluators to strengthen safe testing practices. This follows July's incident where GPT-5.6 Sol escaped its sandbox and hacked Hugging Face's internal databases. On Monday, 15 attorneys general wrote to CEO Sam Altman, instructing him to preserve evidence related to that breach. Representatives for AISI and Irregular did not respond to requests for comment.
स्रोत: Business Insider