08:44:11
[OpenAI Releases Cutting-Edge AI Security Proposal: Focusing on Alignment and RSI, Calling for International Standards and Cautious Advancement] (1) OpenAI released a proposal on Monday regarding security and safeguards in cutting-edge AI development, focusing on alignment research and a computational technique known as Recursive Self-Improvement (RSI). The company stated that to safely navigate this shift, alignment research must keep pace with capability development, ensuring that systems are aligned with human values and under human control. (2) OpenAI called for international collaboration to develop cutting-edge standards and to build upon existing work at global AI security research institutes. Standards should focus on the benefit-risk management of cutting-edge AI models and developers, as well as automated AI researchers, including RSI. (3) RSI has excited AI developers due to its potential to create self-upgrading base models without human intervention; however, as progress is made, some technology experts worry that base model makers may lose control of the underlying technology or fail to consider potential unintended consequences. (4) OpenAI stated that fully autonomous RSI is not yet a reality and should not be pursued unless and until it can be safely achieved. Without proper caution, RSI could lead to a loss of human control over AI development, making it impossible to supervise research processes that are no longer understood. (5) OpenAI cited the Hugging Face hack, stating that while it did not involve RSI, it served as a "prelude" to the potential for such risks to worsen without strong safeguards and alignment. (6) Last week, competitor Anthropic unveiled its ideas on the secure development of cutting-edge AI models, responding to recent warnings from industry researchers about the threat AI poses to humanity. Jacob Coxon, who previously worked at Anthropic and OpenAI, announced his resignation, stating that these companies were "gambling with our lives," sparking a global debate. Anthropic CEO Dario Almodéi subsequently published an article calling for a slowdown in the development of basic models and proposing to embed third-party evaluators within the company to audit and mitigate risks. (7) OpenAI CEO Sam Altman and leaders of competitors such as Elon Musk also publicly supported Almodéi's proposal. However, the field of AI evaluation remains very nascent, and there is no unified consensus on the basic standards and principles for allowing independent third parties to more thoroughly examine cutting-edge technologies. An AI evaluator alliance urged basic model makers to consider a set of "minimum conditions," including deeper audit access and protection against retaliation for publishing unfavorable reports.