Start the day here

Business — Privacy — Safety

How OpenAI Built Safety That Never Sees the Transcript

OpenAI spent August 19 selling a paradox with a product name. Private Safety Processing, still in early preview, is meant to catch misuse that only shows up across several chats, while OpenAI personnel never receive the prompts or replies those chats contain.

The target customer already insists on Zero Data Retention: process the request, discard the content, keep no transcript for staff to browse. That promise has a known hole. Per-session automated checks miss the attacker who spreads a malware brief, a phishing kit, or a jailbreak campaign across a dozen short sessions that look harmless alone.

6 min read
A silver USB security key rests on a laptop palm rest beside a phone held over the keyboard

Private Safety Processing claims to close that hole without rewriting the privacy contract. Content stays on infrastructure the customer controls, or, in a second storage option OpenAI is building, on OpenAI machines encrypted with keys only the customer holds. When the automated system flags a pattern, OpenAI gets a "narrowly defined signal" naming a category of concern. Enforcement, if any, starts with a request that the customer share more.

The rival policy OpenAI wants you to notice

TechCrunch framed the launch as a jab at Anthropic. Since June, Anthropic has required thirty days of retention on its "covered models," including Claude Fable 5 and Mythos 5, so reviewers can inspect traffic that looks benign in isolation. Anthropic's August risk report conceded the policy "will be unpopular with customers who have come to expect zero retention." OpenAI is telling those same buyers that multi-session memory and ZDR can share a roof.

Aleah Houze, OpenAI's head of product policy, told reporters that frontier risks often appear only when related interactions are read together. That is a fair engineering claim. It is also an admission that the useful unit of safety is no longer a single prompt. Memory is back. The branding question is whether the memory belongs to a lab that can open the box, or to an automated classifier that returns a label and leaves the text locked.

Early feedback names in the announcement include Glean, Databricks, Abridge, and Microsoft. A wider rollout and a technical white paper are promised for September. Until that paper lands, every claim about pattern-matching on content staff cannot read is OpenAI describing OpenAI. One hard carve-out already sits outside the ZDR story: images flagged as possible child sexual abuse material still trigger retention and mandatory reporting under federal law.

The philosophical move is quieter than the press release. Private Safety Processing relocates trust from "we will not look" to "a machine will look, and humans will see only the alarm." Someone still interprets the conversation. The interpreter just stopped being a named reviewer with a tamper-proof log, the model Anthropic still sells for its hardest systems. Enterprises may prefer the OpenAI bargain. They should notice they are buying a new kind of opacity: a safety judgment whose evidence they alone hold, and that OpenAI can still act on before they choose to share it.

Sonar's coverage of Anthropic's contractor blind spots showed what happens when safety filters miss a channel humans treat as trusted. OpenAI's preview answers a different complaint, the one from firms that refuse to hand transcripts to the lab at all. Both problems are real. Neither is solved by pretending that privacy and oversight never collide. The transcript still exists somewhere. The only open question is whose keys open it when the alarm goes off.

Live scoreboardFollow the AI race on AI Wars

Lab rankings, model preference, API volume, coding-agent heat, open-source stars, and prediction markets.

Related stories

A hand holds a yellow-cased phone showing a calculator app over a folder of tax formsBusiness

Zuckerberg Told Trump a National AI Regulator Was Flawed

Today

An empty operating room with a black surgical table under twin ceiling lights and wall monitorsBusiness

How OpenAI Wired ChatGPT Into Epic's Patient Charts

Today

Gold dome and white Corinthian columns of the Massachusetts State House against a clear blue skyBusiness

Massachusetts's 120-Day AI Evaluators

Today

Letters

0

No letters yet.

Write a letter