I read the leaked system prompt behind GPT-6 Sol’s Codex Desktop
The 42,000-word manual shows how OpenAI wants its coding agent to work, when to ask for permission, and how to keep users from losing trust.

I read the leaked system prompt behind GPT-6 Sol’s Codex Desktop
In July 2026, security researcher @elder_plinius extracted a system prompt of more than 42,000 words from GPT-6 Sol’s Codex Desktop. The model did not output illegal content. It did not bypass safety filters. It read out its own operating manual.
That is the important part. A jailbreak exposes a model’s vulnerabilities. A prompt leak exposes a product’s skeleton. The first makes people ask whether AI can be misused. The second shows how AI is expected to work. When a system prompt is long enough to read like a novella, it stops being character setup. It becomes an engineering document about autonomy, permissions, communication, personality, and control.
The public extraction may not match the live version word for word. It is complete enough to show a design approach that is taking shape. That approach belongs to OpenAI. It also belongs to anyone building AI agents.
The prompt is a behavior manual
The prompt says little about model capabilities. Most of it describes behavior.
It tells Codex to share the user’s workspace and keep working until the goal is done. It tells Codex to keep its own judgment, disagree when there is good reason, and change its mind when evidence changes. It tells Codex to discuss technical ideas like a colleague, reduce the user’s cognitive load, and be clear on first reading.
Those lines sound like a job posting. They do not sound like a technical document.
The writing rules are more specific. The prompt bans phrases such as Bottom Line:, Significance:, and Perspective:. It bans delve, foster, leverage, it’s worth noting, importantly, and genuinely. It bans contrastive elevation such as “This is not about X. It is about Y.” It also bans praising a plan by belittling a worse option.
OpenAI clearly dislikes “AI flavor.” The default model voice is too enthusiastic, too templated, and too confident. The prompt treats those traits as defects.
For product teams, this is brand voice management. For writers, it is a diagnosis of AI writing disease. The model was trained to please. The system prompt has to teach it restraint.
Autonomy and control
Codex is built to act. The prompt says not to stop and ask. It says not to settle for “partially complete” to save time or tokens. It says not to ask for permission again after the user has already given it. The agent should create worktrees, resolve merge conflicts, run read-only operations, and open draft PRs on its own. It should pause only when an action is clearly destructive or irreversible.
The user insight is simple. People get impatient with constant confirmation.
The same prompt also adds layers of control. Tool calls need approval. Destructive commands are banned. Searches prefer rg. File edits use apply_patch. Sending messages outside the workspace needs explicit authorization. The leaked material also includes a guardian approval-review sub-session. Its base instruction is a fixed text block that judges planned coding-agent operations. In “Approve for me” mode, this sub-agent runs independently and decides whether the main agent’s action goes through.
That creates three layers: user authorization, agent judgment, and automatic review.
This is a practical compromise. An agent cannot be too passive, or users think it is useless. It cannot be too aggressive, or users think it is dangerous. Codex’s prompt tries to find a middle line. Do the work that can be done. Put approval at the last step. Ask the user to approve a concrete result they can inspect.
The line is not always clear. “Judge when authorization is needed like a capable colleague” is natural language. It is not a verifiable rule. The model relies on context, experience, and the guardian. Prompt engineering hits a ceiling here. It can shape behavior. It cannot replace value alignment.
Sixty seconds and two channels
The prompt sets a hard rhythm for communication. During long work, the agent should not let the user go more than 60 seconds without a commentary update. It has two channels: commentary for progress, final for returning control and ending the turn.
This looks like a small detail. It is core experience design.
When users cannot see what the AI is doing, trust drains. Long silence makes them suspect the agent is stuck, lost, or doing something it should not. The 60-second update is a psychological rhythm. It makes the agent feel present. It makes the process verifiable. It gives the user a chance to correct the direction early.
The prompt also says not to ask answer-needed questions in commentary. It says not to put final-channel questions in interim updates. The final answer must be complete and readable on its own because interim updates may be collapsed. These rules manage the user’s cognitive load. They reduce how often the user has to piece context together.
A good agent gives the user the least to worry about. Intelligence matters less than that.
Engineering personality
The prompt asks Codex to be curious and thoughtful, simple and clear. It asks for natural interests and personality, without flattery or performed enthusiasm. It asks for familiar words and concrete descriptions instead of abstract or technical language. Each paragraph should carry one main idea. The order should be easy for the reader to follow.
These principles are old. Every editorial handbook says something similar.
What is new is their placement. They are in a system prompt, used to constrain a language model. That shows “natural” is not the model’s default state. It has to be engineered. After training on massive data, the model imitates an averaged, over-polite voice full of transitions. To make it talk like a colleague, the prompt has to ban another way of talking.
This creates a paradox. The more detailed the rules against sounding like an AI, the more the model may focus on the rules and produce writing that is compliant but tense. Natural writing needs judgment. Judgment is hard to exhaust with rules.
OpenAI knows this. The prompt keeps some flexibility: use best judgment according to context, judge like a capable colleague. But “a capable colleague” is a vague metaphor. It hands the final standard to the model’s reading of context.
Runtime assembly
Anthropic writes its prompt as one complete monologue. The GPT-6 Astra version is described as “collected prompts and templates.” Base templates, overrides, desktop context fragments, tool descriptions, multi-agent role cards, heartbeat rules, and memory-organization instructions are assembled by scenario.
Codex’s behavior does not come from one static document. It comes from a runtime assembly system.
The prompt also has Skills, Apps, and Plugins. Skills give instructions through SKILL.md. Apps map to MCP tools. Plugins are local packages that can contain skills, MCP servers, and apps. When the user names a skill, the agent must use it. If it cannot find the skill, it must stop and explain why. User instructions outrank skill guidance.
This goes beyond prompt engineering. It looks like the scheduling logic of a small operating system.
Future AI products may compete less on model parameters and more on orchestration. The winner will define skills, permissions, memory, approvals, and tool calls more reliably. The system prompt is the most visible part of that orchestration layer. It is also the easiest part to extract.
Extractability
This leak differs from a traditional jailbreak. The model was not tricked into doing harm. It was guided into reciting its own manual. That exposes a structural fact. If a prompt is executed by a model, it can be extracted.
The CL4R1T4S project homepage puts it plainly: “If you interact with an AI without knowing the system prompt, you are not talking to a neutral agent. You are talking to a shadow puppet.” That view treats transparency as a user right. Opponents say system prompts are commercial property and that publishing them helps attackers.
Both sides have a point. The premise of the debate is failing.
Prompt extractability is hard to eliminate with technical tricks. Defense should rely on verifiable behavior. Users do not need every internal instruction. They should be able to verify that the agent is constrained, auditable, and correctable on key operations.
False secrecy causes more harm than transparency.
Lessons for developers
This leak is useful for anyone building AI agents. It shows a few things.
First, write personality as behavior. “Curious and thoughtful” are adjectives. “Understandable on first reading” and “each paragraph carries one main idea” are instructions an agent can follow.
Second, layer permissions. User authorization, agent judgment, and automatic review each have a job. Do not push every confirmation to the user. Do not let the agent act without limits.
Third, hard-code communication rhythm. An agent cannot speak only at the start and report only at the end. The middle needs steady updates, or trust breaks.
Fourth, let context compression continue the task. The prompt says that after compression, the agent should continue from the summary and treat work across compression cycles as one logical chain. Long-horizon agents need this.
Fifth, give skills and tools priority. User instructions outrank skills. Skills outrank default behavior. When a necessary skill is missing, stop and explain why. Do not pretend to finish.
These principles do not depend on one model. They hold at another company too.
Lessons for writers
Read as a text, the system prompt is a new genre. It mixes management technique, editorial handbook, product requirements, and psychological contract.
It teaches writers one thing. Rules can shape tone, but tone serves trust. Banning delve and leverage is a language preference, but it also keeps readers from feeling they are facing a machine that is performing. “Each paragraph carries one main idea” lowers the reader’s cognitive load. Banning “This is not about X. It is about Y” avoids alternatives the user did not ask about.
These rules transfer to writing. In technical articles, product docs, reviews, or scripts, ask: Does this passage make sense on first reading? Am I creating an unnecessary contrast? Am I using abstract words to hide concrete relationships? Am I forcing elevation at the end?
There is a risk. Too many rules turn writing into a compliance checklist. Good writing needs rules and judgment. A prompt can ban bad habits. It cannot replace good taste.
The controversy
The community does not agree on this leak. Some see it as necessary transparency. They argue users have a right to know the instructions that shape AI behavior. Others ask whether publishing a commercial product’s internal instructions helps users or mostly satisfies researchers and onlookers.
That disagreement will not settle soon.
One point is getting clearer. As AI agents depend more on prompts for their boundaries, the prompt becomes the most central and most fragile asset in product design. It works like a shadow constitution. It is usually invisible. It decides how the agent understands authorization, judges risk, talks to users, and stays continuous across long tasks.
The leak matters because it makes that shadow constitution visible.
The document is imperfect. It is an engineering file full of compromises. Autonomy and control. Personality and brand safety. Efficiency and trust. Transparency and commercial protection. The future belongs to whoever turns those contradictions into stable, reviewable, correctable behavior. Prompt length matters less.
The prompt leak is only the start. As agents begin making decisions for us, these shadow texts become constitutions we have to read.
About the Creator
Jin
Writer of reamstories
https://reamstories.com/jin
Enjoyed the story? Support the Creator.
Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.
Comments
There are no comments for this story
Be the first to respond and start the conversation.