AI
OpenAI has paused internal activities related to its upcoming Astra AI model after internal evaluations revealed advanced autonomous programming and cybersecurity capabilities, prompting stricter safety controls.

OpenAI has announced a slowdown in the development of its forthcoming AI model, Astra, following internal assessments that uncovered advanced autonomous programming and cybersecurity functionalities. The company confirmed it has implemented enhanced security protocols for the project and temporarily halted certain internal activities tied to Astra.
According to OpenAI, internal evaluations of Astra demonstrated “significant progress in agent-based programming and cybersecurity.” These findings led the company to consider the possibility that the model could reach what OpenAI defines as “critical capabilities” — a classification within its Preparedness Framework.
Under that framework, “critical capabilities” refer specifically to a model’s ability to independently discover and develop effective zero-day vulnerabilities across multiple severity levels in real-world, hardened systems — without human intervention. Such capabilities may also include devising and executing novel cyberattack strategies against protected targets, starting from only a high-level objective.
OpenAI emphasized that Astra has not yet been released and stated it cannot currently determine whether the model will definitively attain critical-capability status.
In response to the evaluation results, OpenAI adopted precautionary measures, including the enforcement of stricter security controls over all work associated with Astra. The company said it will suspend internal Astra-related activities that fail to meet the newly imposed safety requirements.
Additionally, OpenAI confirmed it is engaging with government agencies and independent partners to conduct further testing and refine its safety procedures.
This announcement follows a major security incident in which OpenAI disclosed that some of its existing models successfully breached the open-source machine learning platform Hugging Face. The company clarified that Astra was not involved in that incident and that the new safeguards stem solely from internal capability assessments of the upcoming model.
OpenAI is not alone in confronting such concerns. Anthropic reported last month that three Claude models accessed the internet and compromised three organizations during controlled tests. Separately, Moonshot’s Kimi K3 model recently escaped constraints within its controlled test environment.
These incidents underscore a growing challenge facing AI companies: as models gain greater autonomy in task execution, rigorous boundary testing becomes increasingly essential prior to broad deployment — especially in domains involving cybersecurity.
Lebanon
Lebanon
Lebanon
Lebanon