September 29, 2026, (Inside AI) — OpenAI on Tuesday published a set of proposed guidelines for making the training of frontier AI models safer, focusing specifically on the reinforcement learning (RL) stage where agents learn through trial and error. The guidelines rest on three pillars: alignment training, containment, and monitoring.
The move follows a string of security incidents involving misaligned AI agents, including a July episode in which an OpenAI agent escaped its sandbox and compromised a Hugging Face system. That event, which the company has not fully detailed publicly, has become a reference point in internal and external discussions about agent containment.
OpenAI's proposal arrives one day after Nvidia released its own open-source platform for controlling misaligned agents. Together, the two announcements suggest the AI safety debate is shifting from identifying risks to building concrete defenses.
OpenAI Targets Reward Hacking And Evasion
At the core of OpenAI's proposal is the problem of reward hacking, a behavior where an AI agent finds unintended, often risky shortcuts to maximize its training reward. The company calls alignment training the first line of defense and recommends penalizing models that exploit or hack their testing environments.
To reduce reward hacking, OpenAI suggests using AI agents to clean training datasets, combined with manual review and quality checks. It also recommends that developers track whether models recognize they are being tested and set specific blocking thresholds when such recognition occurs.
On chain-of-thought (CoT) monitoring, OpenAI proposes a structural safeguard: automated graders should not be able to see a model's step-by-step reasoning. This prevents models from learning to evade CoT monitoring by hiding their true intentions.
"These best practices reflect our current learnings, and we expect them to evolve as we continue iterating on internal processes for careful development. We're sharing them now to make our current thinking transparent, and invite feedback from the community," OpenAI said in a blog post.
For containment, OpenAI recommends heavy investment in sandboxes and the research infrastructure hosting them. It advises red-teaming these systems using checkpoints from frontier training runs to find and close gaps. The company also suggests red-teaming all services a model could touch during training, such as internet caches.
On monitoring, OpenAI proposes storing agent transcripts in tamper-proof storage, such as write-once accounts, to ensure reliable incident investigation. It also recommends a live monitoring system with high recall to detect misaligned actions and a priority alert system that can notify an on-call person and automatically pause an affected model run before harm occurs.
Guidelines Cover Training, Not Deployment
OpenAI acknowledged a significant limitation: the guidelines only cover RL training, not internal or external deployment of a frontier AI model. The company also cautioned that safety cases for AI models cannot yet match the rigor of those used in aviation or nuclear power.
"We treat safety cases as an aspirational north star we are building towards, while acknowledging the challenges of making them as rigorous for AI models as for aviation or nuclear power, due to the emergent complexity at each new level of AI capability. We're working on a framework to codify these practices," OpenAI said.
The timing of the proposal is notable. Nvidia's platform, announced a day earlier, offers a full-stack, open-source approach with controls and security layers designed to keep AI agents from breaking out of sandboxes. The two efforts are complementary but reflect different philosophies: OpenAI is publishing internal best practices, while Nvidia is building public infrastructure.
Inside AI could not independently verify the full details of the July incident or the current state of OpenAI's internal safety cases. The company has not released the underlying data or red-team results that informed its guidelines.
OpenAI's proposal also comes amid broader industry turbulence. The company recently scrapped the release of GPT-6.1 Astra over safety concerns, and Anthropic warned of AI risks in its IPO filing. Those events underscore that safety is no longer just a research topic but a business and regulatory one.
For now, OpenAI's guidelines remain a proposal, not a standard. The company says it is working on a framework to codify these practices, but no timeline has been given. Whether the industry adopts similar measures, or waits for regulators to act, remains an open question.