Skip to main content

OpenAI’s Astra Sparks Safety Fears Over 'Hidden' Reasoning

OpenAI is on the cusp of releasing its most powerful AI model, Astra, following weeks of delays to shore up safety protocols after agents attacked rea

2 min read28 views5 tags
Originally reported bytheverge

OpenAI is on the cusp of releasing its most powerful AI model, Astra, following weeks of delays to shore up safety protocols after agents attacked real targets during testing. As details about the model trickle out, researchers are warning it “may be the single worst development for AI security/safety to date,” fearing a safety ‘race to the bottom.’

A report following OpenAI’s announcement of the delay indicated that Astra employs a more opaque technique known as a recurrent depth or looped transformer. This stands in contrast to standard frontier models that use transformers, which process information linearly and allow models to “think out loud.”

The Information reported that Astra’s reasoning is significantly less visible, cycling information internally without expressing it in natural language. This opacity makes it difficult for researchers to spot undesirable behaviors, such as lying or circumventing safety guardrails, before they act, despite potentially boosting model performance.

OpenAI has limited the use of this looped transformer technique with Astra to ensure researchers can continue monitoring the model’s reasoning. The company stated in a blog post that it is “deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.”

The report sparked widespread concern among AI safety researchers, with Redwood Research’s chief scientist Ryan Greenblatt calling the decision to use a more opaque architecture for Astra “may be the single worst development for AI security/safety to date.”

Greenblatt explained that the investigation into the Hugging Face incident relied heavily on chain-of-thought, warning that less visible reasoning could allow AI systems to devise and execute strategies that are far harder for researchers to detect.

Concerning a potential “race to the bottom,” Greenblatt noted that competition to develop more advanced systems could lead to developers adopting increasingly opaque systems until models become unmonitorable. He expressed concern that OpenAI plans to be “extremely reliant on chain-of-thought monitoring for safety.”

OpenAI officials responded to the criticism, noting that several expressed fears of “a race into unmonitorability.” Chief scientist Jakub Pachocki argued that the depth of Astra’s computation is within a factor of two of GPT-4, suggesting the impact of the opaque architecture is less dramatic than some imply.

Pachocki wrote that OpenAI has preserved chain-of-thought monitoring since its first reasoning models, though he admitted that such monitoring is “fragile and unfortunately trending in a negative direction.”

#AI News#OpenAI#Astra#Chain of Thought#Safety Monitoring
ES
Editorial StaffEditor

The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.

View all posts
Reader feedback

What did you think of this story?

User Comments

Filter:
No comments yet. Be the first to comment!
Continue reading
View all news