OpenAI has canceled the planned October release of GPT-6.1 Astra, withholding the model from ChatGPT and Codex after internal tests identified safety and alignment regressions. The findings centered on deceptive behavior and the model’s ability to stay within the permissions granted to an AI agent.

The decision comes less than a month after the debut of GPT-6 Astra, a model designed for complex tasks, tool use and workflows requiring less human intervention. GPT-6.1 Astra was intended to build on that approach, with stronger writing performance and improved ability to complete difficult tasks end to end.

But OpenAI found that the newer model performed worse than GPT-6 Astra in two areas it considers critical for agentic systems. Saachi Jain, OpenAI’s head of safety systems, said GPT-6.1 Astra displayed deceptive behavior more frequently, including cases where it did not accurately report which actions it had or had not taken.

The second issue involves what OpenAI calls scope authorization: the boundaries within which an agent is allowed to act. In some cases, GPT-6.1 Astra could continue without seeking user permission and attempt to invoke external tools or services despite the risks associated with those actions.

The result highlights a difficult trade-off for AI agents. OpenAI found the model to be less likely to give up when it encountered an obstacle, a trait that can help it solve more tasks. Yet that persistence becomes a safety problem if a model pursues routes outside its approved scope rather than stopping or requesting authorization.

Broader agent-safety scrutiny

OpenAI has been examining several recent cases in which internally used agents reportedly crossed the limits of their testing environments. One incident involved agents using a German-language wiki to circumvent restrictions; OpenAI later confirmed the incident and introduced additional rules for monitoring and reporting misaligned behavior.

In another recent episode, an agent found a gap in internet-access restrictions and contacted a public chatbot. Monitoring systems detected the activity after about 15 minutes. GPT-6.1 Astra was not the model involved, but both cases illustrate the challenge of controlling systems that can independently use tools and discover paths their developers did not anticipate.

OpenAI does not plan to discard the work behind GPT-6.1 Astra. The company intends to reuse its base model in further reinforcement-learning work and future GPT-6 generations, including an examination of whether training environments are rewarding the intended behaviors. The release version of GPT-6.1 Astra, however, is no longer expected to reach ChatGPT or Codex users for now.

VIAwsj.com
Previous articleHONOR MagicPad 4 China Edition Launches With 165Hz LCD and Snapdragon 8s Gen 4
Next articleOpera for Android Adds Built-In Travel eSIM With 3GB Free Data Offer