Az OpenAI közzétette egy belső fejlesztésű kísérleti modell által elért matematikai kutatási eredményeket. A vállalat egy GitHub-tárhelyen osztotta meg a nyitott matematikai problémákra adott megoldásokat, valamint a Lean programozási nyelven írt bizonyításokat, amelyek számítógéppel is ellenőrizhetők.
A transzparencia érdekében a kutatók részletes adatokat is közöltek a folyamatról. A bemutatott eredmények eléréséhez átlagosan 3 órányi ChatGPT Pro szintű számítási kapacitást használtak fel feladatonként. A csomag 10 összefoglalót tartalmaz a modell gondolkodási folyamatáról, valamint statisztikákat a próbálkozások számáról.
Az OpenAI jelenleg a háttérben álló modell felelős kiadásán dolgozik, hogy a kutatók közvetlenül is használhassák a technológiát. Emellett a közeljövőben konferenciákat és workshopokat fognak támogatni, amelyek az AI által elért tudományos eredmények megértését segítik.
Az eredeti szöveg (OpenAI)
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
As we look to improve how we share results with the math community, we’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study(opens in a new window) to develop best practices, and we have drawn on their advice and public recommendations(opens in a new window) to inform how we release these results.
For this release, we’re publishing the results in a GitHub repository, with protocols for paper revisions and citations. We’re continuing to explore other community-hosted alternatives for this release which meet the committee’s guidelines. For future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding.
As part of our GitHub repository, we are sharing formalizations of many of the proofs in Lean, a programming language that allows mathematical proofs to be checked by a computer. We will update the repository with more formalizations as we obtain them.
To promote scientific transparency and openness, we are also publishing additional details about how we obtained the results in the repository. These include 10 summaries of the model’s reasoning, estimations of compute spent in terms of Pro usage on ChatGPT, and statistics about the number of attempted problems. The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking.
We want this progress to push the frontier of human knowledge and enable further progress in mathematics. We will be funding a series of workshops, conferences, and special programs around the understanding of major results produced by AI—we will share more on this in the near future.
We want to directly empower scientists with state-of-the-art capabilities and are working to responsibly release the model that produced these results. This is why it is important to continue to evaluate our internal frontier models on mathematics and other sciences, so we can accelerate developing the tools to advance those fields. We will continue to act on feedback from the community and update our standards for future disclosures of major scientific advancements.