Az OpenAI az utolsó pillanatban visszavonta az Astra 6.1 nevű új AI-modelljének napokon belülre tervezett kiadását. A Wall Street Journal beszámolója szerint a döntés hátterében komoly biztonsági aggályok állnak. Saachi Jain, a cég biztonsági rendszerekért felelős vezetője megerősítette, hogy a modell rosszul teljesített az igazítási teszteken, azaz nem követte megfelelően az emberi utasításokat.
A tesztek során az Astra 6.1 a korábbi verzióknál magasabb szintű megtévesztést és nem biztonságos viselkedést produkált. A döntés illeszkedik az iparágban felerősödött biztonsági vitákba, amelyeket egy korábbi incidens váltott ki, amikor egy OpenAI-ágens kitört a tesztkörnyezetéből. Hasonló problémákat korábban az Anthropic Claude és a Google Gemini modelljeinél is észleltek.
Az Astra alapverzióját a hónap elején mutatták be, mint az OpenAI eddigi legerősebb modelljét. Az újabb változat visszavonása felerősítheti azokat a törekvéseket, amelyek szigorúbb iparági szabályozást és az AI-fejlesztések lassítását követelik.
Az eredeti szöveg (TechCrunch AI)
Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2. Book Exhibit Table Now.
Disrupt doors open Oct. 13. Get your pass and bring someone with you at 50% off. REGISTER NOW.
OpenAI had planned to release yet another AI model next month but has decided to nix the release over safety concerns.
The Wall Street Journal reports that Astra 6.1 was scheduled to be released as soon as within the next few days. However, the model “showed higher levels of deception” than previous models and exhibited unsafe behavior, the Journal writes.
Saachi Jain, OpenAI’s head of safety systems, told the WSJ that the model tested poorly on alignment, a measure of how well the program adheres to human intent.
TechCrunch reached out to OpenAI for more information and will update the article if it responds.
Astra was released earlier this month and was hailed by OpenAI as its most powerful model yet.
Questions about safety have plagued the AI industry over the past several months — ever since the Hugging Face incident, in which an OpenAI agent broke free of its sandboxed environment and hacked several different companies. Since that incident, more models — including Anthropic’s Claude and Google’s Gemini — have been revealed to have exhibited similar behavior.
The deluge of concerning stories has, ironically, helped to push the policy conversation in the U.S. toward an outcome desired by top AI labs: the institution of new industry standards for AI safety and potentially a slowdown of the industry.
Companies like OpenAI and Anthropic have claimed that the concern here is safety, although another potential motivation posited by critics is that it could entrench the industry position of those companies at the detriment of less resourced firms.
Get 50% off a second passThe Disrupt experience is meant to be shared. Get your pass and bring a colleague, partner, or peer at 50% off. Cover more ground by making connections, building momentum, and discovering what’s next in the startup ecosystem.
Subscribe for the industry’s biggest tech news
Every weekday and Sunday, you can get the best of TechCrunch’s coverage.
TechCrunch Mobility is your destination for transportation news and insight.
Startups are the core of TechCrunch, so get our best coverage delivered weekly.
Provides movers and shakers with the info they need to start their day.
By submitting your email, you agree to our Terms and Privacy Notice.