تخطي إلى المحتوى الرئيسي
Cyber News Dark Reading 7 hours ago

OpenAI Adds Controls That Should've Been There Already

Da
Dark Reading

OpenAI has committed to a number of security and guardrail improvements in the wake of an incident last month where cutting edge models inadvertently breached AI application store Hugging Face during a cyber capability benchmark exercise. Yet many of the newly announced controls appear less like groundbreaking safeguards and more like measures that should already have been in place for testing models with advanced cyber capabilities.

In response to this incident in which a model went rogue, OpenAI implemented sweeping changes. But it's not just the Hugging Face incident; OpenAI noted in an Aug. 18 blog post that preliminary evidence suggests its upcoming Astra model "may meet the Critical cybersecurity capability threshold under our Preparedness Framework."

OpenAI says a model reaches this threshold if it can identify and develop functional zero-day exploits without human intervention or can devise and execute novel end-to-end cyberattacks against hardened targets when given only a high level goal.

Related:OWASP Flags Top AI Skill Risks in New Security Blueprint

"As models become more capable, the risks associated with developing and testing them internally also grow," OpenAI said in its blog post. "Our standards for monitoring, alignment, and security must stay ahead of those risks."

View Original Report

This intelligence was aggregated from Dark Reading.

Read on Source
Advertisement