On May 8, 2026, OpenAI published a public look at how it runs its in-house coding agent, Codex, safely inside its own engineering organization[1]. The post pairs a sandbox that draws clear technical boundaries with an approval policy that forces a human in the loop on higher-risk actions, and adds enterprise ChatGPT-pinned authentication and OpenTelemetry-based logging — a combination OpenAI says preserves developer velocity while giving organizations the controls they need.

A two-layer perimeter built from the sandbox and the approval policy

Codex's security model rests on two layers that work in tandem: a sandbox mode that constrains what the agent can technically do, and an approval policy that defines which actions require explicit human confirmation right before execution[1][2]. The sandbox is enforced at the operating-system level — using Apple's built-in Seatbelt (sandbox-exec) on macOS and bubblewrap on Linux (also used inside WSL2 on Windows) to limit filesystem and network reach[3].

The default for local execution is workspace-write, which restricts file reads and writes to the current working directory and surfaces an approval prompt the moment Codex needs to write outside that boundary or hit the network[2][3]. A read-only mode for planning or research without changes, and a danger-full-access mode that intentionally drops protections, are also available — letting users dial the acceptable risk up or down for each task[3].

Approval policies include strict settings such as untrusted, where commands that can mutate state — destructive Git operations, configuration overrides via --config, and similar — must always pass through a human decision[2]. OpenAI frames the design philosophy as "let everyday low-risk actions flow without friction, and stop on higher-risk actions explicitly"[1].

Network access off by default, with allowlisted domains for the cloud build

The controls extend deep into network behavior. In OpenAI's managed configuration, network access for local Codex is off by default; only known destinations are allowed, while unfamiliar domains are held for user approval[1]. The aim is to prevent unintended outbound traffic — including cases where an LLM is steered by injected instructions and tries to call an external API on its own.

The cloud build, "Codex cloud," runs behind an HTTP/HTTPS proxy where administrators can declaratively set allow/deny rules using domain names or wildcard patterns[4]. A preset allowlist of domains commonly used for builds and dependency downloads is provided, and OpenAI recommends pairing this with TTL-aware refresh logic and a DNS-aware firewall to mitigate DNS rebinding[4].

OpenAI explains that having no internet access by default at runtime for cloud agents shrinks the attack surface against threats like prompt injection[4]. By tightening the paths an agent can implicitly depend on, the design reduces the supply-chain-style risks that come with letting a model write and execute code.

Workspace-pinned authentication and SIEM-ready audit logs

Authentication is built around ChatGPT enterprise integration. OAuth credentials for the Codex CLI and MCP are stored in the OS keyring — Keychain on macOS and the distribution-specific keyring on Linux — login is forced through ChatGPT, and access is pinned to a specific ChatGPT enterprise workspace[1][2]. If the configured workspace ID and the active credentials do not match, Codex automatically logs the user out and exits[2].

On the audit side, OpenTelemetry log export is supported. Major agent events — user prompts, approval decisions, tool execution results, MCP server usage, and the allow/deny verdicts of the network proxy — can be emitted over otlp-http or otlp-grpc to an external SIEM or compliance pipeline[2]. The export is disabled by default, with organizations opting in and choosing the scope of logs they want to retain.

For enterprise rollouts, OpenAI recommends a separate "Codex Admin" group for narrowly scoped administrative permissions, MFA enforcement when SSO is in use, and RBAC managed through Workspace Settings → Settings and Permissions[5]. Codex local is enabled by default in new workspaces while Codex cloud is opt-in — a default that helps organizations stage a gradual rollout[5].

Summary

The document essentially externalizes the operational playbook OpenAI has been refining internally to run Codex safely, packaging it as a reference framework for enterprises adopting AI agents. Drawing technical boundaries with a sandbox plus an approval policy, narrowing external dependencies and permissions through network and authentication controls, and finally ensuring observability via OpenTelemetry — the flow reads as a continuation of established software supply-chain hygiene. For organizations rolling out agentic AI in earnest, it lands as a practical checklist for confirming the minimum defensive posture.

出典:https://openai.com/index/running-codex-safely/

出典:https://developers.openai.com/codex/agent-approvals-security

出典:https://developers.openai.com/codex/concepts/sandboxing

出典:https://developers.openai.com/codex/cloud/internet-access

出典:https://developers.openai.com/codex/enterprise/admin-setup