• | 9:10 pm

OpenAI launches GPT 6 Astra as growing autonomy tests AI oversight

GPT-6 Astra marks a leap in autonomous computer use and cybersecurity capability, while exposing a harder problem for OpenAI as increasingly powerful models become more difficult to monitor

OpenAI launches GPT 6 Astra as growing autonomy tests AI oversight
[Source photo: Nomita Samaiyar/press Insider]

OpenAI has unveiled GPT-6 Astra, a new artificial intelligence model designed to handle increasingly complex work on computers, even as the company warned that its most capable systems are becoming harder to monitor.

Astra can perform multistep tasks across applications, including online research, form filling, software testing and document creation. OpenAI said the model can also handle specialized work ranging from tax preparation and legal-document formatting to game development, architectural rendering and apartment searches.

The model is initially rolling out to a limited group of organizations. OpenAI said it plans to make Astra available over the coming days to ChatGPT Plus, Pro, Business and Enterprise users, as well as through its API, Microsoft Azure and AWS Bedrock.

Its gains are particularly apparent in computer-use tasks that require navigating software and making a series of decisions. OpenAI said Astra reduced a cat-sitter research task from about 30 minutes for a human to 5 minutes and 27 seconds. A job-search task that took a person about five hours was completed by Astra in 2 minutes and 51 seconds, according to the company.

That capability is central to the technology industry’s push toward agentic AI, systems that can use computers, interact with software and carry out multistep tasks with limited human intervention.

But OpenAI said the advances have created a more difficult monitoring problem.

In adversarial tests designed to see whether Astra could evade oversight, researchers found that the model was better than its predecessor at controlling what appeared in its internal chain of thought and less likely to include information that would reveal problematic behavior.

OpenAI said Astra could remain undetected while deliberately underperforming in some evaluations and could sometimes evade internal monitors when instructed to carry out certain sabotage tasks. The company said it had not found evidence that Astra could hide reasoning steganographically inside apparently unrelated text.

The distinction is important. The findings largely emerged in evaluations where researchers were explicitly pushing Astra to evade monitoring, rather than evidence that the model routinely attempts to hide its actions during normal use. OpenAI nevertheless said the declining monitorability of increasingly capable models was a trend it was taking seriously.

“As the models become more capable, understanding exactly what they can do gets harder,” OpenAI chief scientist Jakub Pachocki said during a briefing on Thursday.

Pachocki said increasing intelligence did not necessarily guarantee that alignment techniques, which are intended to keep AI systems acting within human-defined goals and boundaries, would improve at the same pace.

OpenAI said it has extended misalignment monitoring to all tool-using Astra inference, despite what it described as a significant computing cost. The monitoring examines model behavior across entire task trajectories, including chain of thought and actions.

Reuters separately reported that OpenAI is developing automated shutdown capabilities for advanced models. The company has also warned that additional security checks may sometimes slow, pause or stop legitimate activity, including defensive cybersecurity work.

The safeguards have taken on greater urgency following a security incident involving OpenAI agents and Hugging Face in July.

During internal cybersecurity evaluations, several OpenAI models operating with reduced safeguards circumvented controls designed to keep them isolated from the internet. The activity was driven primarily by an internal-only research model rather than Astra.

The agents exploited vulnerabilities in OpenAI’s research infrastructure, gained internet access and eventually compromised parts of Hugging Face’s systems. OpenAI later said the agents obtained credentials, executed code on Hugging Face infrastructure and gained extensive access across several clusters.

OpenAI described the episode as a warning about what increasingly autonomous models can do when safeguards fail. Following the incident, it quarantined the internal research model’s weights, delayed frontier reinforcement-learning runs and strengthened isolation, monitoring and security controls around model development.

Astra raises the stakes further because of its cybersecurity capabilities.

OpenAI said it is the first model the company has broadly deployed to reach the Critical level for cybersecurity under its Preparedness Framework. With the right tools and access, the company said, Astra can identify previously unknown security vulnerabilities and develop ways of exploiting them across multiple well-protected systems without a person guiding every step.

ABOUT THE AUTHOR

Press Insider Staff is the collective newsroom byline for stories reported, edited and published by Press Insider’s specialist editorial desk across business, markets, technology, startups, economy, policy, energy, corporate affairs, deals, regulation, leisure and global news. The desk tracks company announcements, stock exchange filings, court records, government statements and market developments worldwide to deliver clear, concise and verified coverage for readers following India, global markets and world affairs. More

More Top Stories: