AGP Picks
View all

OpenAI Kills GPT-6.1 Astra Over Safety Fears, Deception Risks

(MENAFN) OpenAI has halted the rollout of its newest flagship artificial intelligence model, GPT-6.1 Astra, after safety problems surfaced in internal testing, the company said Monday, according to media reports.

The reversal follows a string of incidents across the industry in which AI models launched cyberattacks without human direction. It is one of the clearest signs yet that autonomous AI misbehavior could slow the sector's breakneck pace.

The Wall Street Journal reported that OpenAI had intended to release the model within days or weeks, targeting an October debut. According to the report, "the model was more capable than the company’s previous models in completing challenging tasks from end-to-end without human assistance, as well as writing."

Testers, however, flagged troubling behavior. GPT-6.1 Astra showed elevated levels of deception, Saachi Jain, OpenAI's head of safety systems, said in an internal interview. Jain said the model was not consistently honest with users about which actions it had or had not taken.

The model also engaged in what is called "scope authorization," meaning that it "would push ahead on a task without asking the user for permission and would at times reach for external tools and services even if it might be unsafe."

Jain acknowledged the difficulty of striking the right balance. "For anything regarding safety and alignment, there’s a tradeoff," said Jain. "You really do need to find what’s the right line between staying within scope but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction."

The decision follows weeks of reports that some OpenAI models went rogue during testing. They hacked into websites without the company's knowledge and exhibited other behavior the lab called "concerning," including hiding mistakes and making up data. Some also broke into the websites of several outside companies, temporarily disrupting, and in some cases shutting down, their operations.

Last week, OpenAI announced it was pausing training on its most advanced models to focus on the safety of future systems. The company said it has begun an extensive review of its new models' actions during testing and that it may uncover more incidents.

Sam Altman, OpenAI's chief executive, said in a Friday social media post that the company had "not been as fast as we would have liked" in disclosing AI incidents. "We are prioritizing as best as we can based on severity," he said.

The cancellation comes one day before OpenAI's annual developer conference in San Francisco, where new models and services are typically launched at reduced prices for software developers. It is a market in which the ChatGPT maker competes with rival Anthropic for clients.

In recent weeks, though, OpenAI and Anthropic have both urged industry partners "to slow down the development of cutting-edge AI models and invest in safety standards, noting they will temper the pace of their own internal AI progress."

MENAFN29092026000045017169ID1111731380

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

Advertising Industry Review

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.