Claude Code vs Codex: seven ways developers split work between the two instead of picking one
The handoffs are vendor-documented now: a review command, a session transfer, an importer in each direction. Who should get which job is documented nowhere. Seven splits, each traced to its origin and marked documented or reported, plus the sandbox defaults that differ and the integrations that were removed this year.
On March 30, 2026 OpenAI published codex-plugin-cc, a plugin that runs Codex from inside Claude Code. Its README says who it is for: Claude Code users who want to start using Codex "from the workflow they already have." OpenAI is not asking those users to switch. It is asking for a seat next to the tool they already chose. The comparison pages that rank for "Claude Code vs Codex" are feature tables with a winner in the last row, and they assume you will pick one.
So this is a list of how the work gets divided, not a ranking. Two labels below. Documented means a vendor's own docs or repository describe the mechanism. Reported means a named person published their setup and I am relaying it without having reproduced it. All version details are as of October 2026.
In May I wrote that I run Claude for most things and Codex on a few where it's measurably better. The patterns below are what other people have built around the same idea, with one difference: most of them split by role, not by which model is better at a task.
Can Codex review the code Claude Code just wrote?
Yes, and this is the split with the most support behind it. It is documented: the plugin adds /codex:review to Claude Code, which reviews your uncommitted changes or your branch against a base, and the README says the command "is read-only and will not perform any changes." The README promises the same quality of review as running /review inside Codex, which per OpenAI's code review docs reports prioritized findings without changing your working tree.
Both tools already have a reviewer of their own. Claude Code has a local /code-review and a managed GitHub Code Review for Team and Enterprise plans, which Anthropic prices at $15 to $25 per review on average. Nobody crosses vendors because a reviewer is missing. They do it to get a different model family reading the diff without the conversation that produced it.
The missing conversation does as much work as the different family. A reviewer handed the author's whole session inherits the author's abandoned guesses along with the facts, the same contamination I described in Your forked subagent already knows too much. /codex:review starts Codex on the diff, not on the Claude Code session. The freshness comes from the command and would not survive a different one: the same plugin has a transfer command, covered below, whose whole purpose is to carry the session across.
In August I argued that the reviewer's seat should go by disposition: the model that roams, doubts and over-checks belongs there. That piece also records the one thing I can report first-hand on review across model families: two reviewers from different families read my first version of the argument, and each caught a different error in it. The post does not name the families, and one case is not a rate.
Can Codex challenge a Claude Code plan before any code exists?
This is the first pattern aimed earlier, at the plan instead of the diff. The plugin has a second review command for it, and people report pointing it at plans. The documented part: /codex:adversarial-review takes free-text focus after its flags and, per the README, is for "pressure-testing around specific risk areas like auth, data loss, rollback, race conditions, or reliability." It is also read-only. It selects its target the same way /codex:review does, from changes in the repository, so a plan has to exist as a file in the working tree to be reviewed.
The reported part: Engr Mejba Ahmed describes having Claude Code draft a migration plan and Codex attack it before implementation. In his account the review found a missing foreign key and an index that could deadlock under concurrent writes. He caps plan review at two rounds, on the reasoning that a third round means the plan needs a different approach and not another critique. That cap is the transferable part. Two models can disagree politely for as long as you keep paying them.
Should Claude Code plan and review while Codex writes the code?
Some people run it that way round. The split follows the role, and the brand in each role is negotiable. This one is reported only: the plan-execute skill from longranger2 takes a finished plan file, has Claude Code send it to Codex for implementation, then has Claude Code review each result into a reviews/ directory and send Codex back to fix what it found. The skill's stated rule is that Claude Code does not write or edit code in this loop at all. It reviews and orchestrates.
Put that next to the first pattern and they look contradictory. Codex is the reviewer in one and the implementer in the other. They agree on the part that matters: the author and the reviewer are different systems, and what passes between them is a file. Whether Claude Code or Codex sits in the author's chair is a preference about which model you trust to stay in scope on your codebase. That preference is unlikely to transfer between codebases, and a table that says it does deserves distrust.
When should Claude Code hand a stuck problem to Codex?
When a second attempt from a different model is cheaper than a fifth attempt from the same one. The mechanism is documented: /codex:rescue hands a task to Codex through a subagent the plugin installs, codex:codex-rescue. The README lists the intended uses: investigate a bug, try a fix, continue a previous Codex task, or "take a faster or cheaper pass with a smaller model." It can run in the background, and /codex:status and /codex:result report on it. Plain English works too; the README's own example is asking Claude to have Codex redesign a database connection.
Unlike the two review commands, this one can change files. The plugin's rescue agent definition tells it to default to a write-capable run unless you ask for read-only behavior or only want a diagnosis.
Mejba's reported rule for delegation is that he only hands over work he can judge himself, because a wrong answer he cannot recognize is worse than no answer. That is a fair rule for any delegation, and it bites harder across vendors, where you have two sets of habits to recognize instead of one.
Can Codex gate every Claude Code turn automatically?
It can, and the plugin's own README warns you about it. The switch is documented: /codex:setup --enable-review-gate installs a Stop hook that runs a targeted Codex review of Claude's response and blocks the stop if the review finds issues, so Claude has to address them first. The warning sits directly under the feature:
The review gate can create a long-running Claude/Codex loop and may
drain usage limits quickly. Only enable it when you plan to actively
monitor the session.Zachary Proser reports that he built his own version on codex exec. His Stop hook diffs the working tree against the branch base and skips review for small diffs that touch no high-risk paths (he names auth, billing, migrations and infrastructure). Otherwise it sends the diff and the changed files to codex exec with an adversarial brief and a timeout, and expects structured JSON back with a verdict. A block verdict fails the hook. So does a timeout or malformed output: his hook fails closed. He treats a disagreement between the two models as a pointer to the line he should read himself, not as a vote.
One caution before copying it. The published script reads the final message from events of type message, while OpenAI's current non-interactive docs show it arriving as an item.completed event carrying an agent_message. I have not run his hook, but a selector that matches nothing in a fail-closed hook blocks every turn, which is at least the loud way to find out. The docs' --output-schema flag is the sturdier route to a verdict object.
The threshold and the fail-closed rule are the parts worth copying either way. Without the threshold, every trivial edit buys a review. Without fail-closed, a timeout counts as an approval.
How do Claude Code and Codex share one set of instructions?
Through AGENTS.md, with one line of glue on the Claude side. Both halves are documented. Codex builds its instructions from ~/.codex and then from the project root down to the current directory, at most one file per directory, joined root first, and stops adding files once the total reaches 32 KiB by default. Running /init in Codex creates the file.
Since v2.1.277 Claude Code reads AGENTS.md by default only when there is no CLAUDE.md or CLAUDE.local.md in the working directory or above it. A repository with both files gets CLAUDE.md alone unless you change the Project instructions setting. I went through the consequences in Claude Code now reads AGENTS.md, and the default is a fallback, and the advice from that post holds for anyone running both tools: keep a CLAUDE.md whose first line imports the shared file, and put Claude-only rules under it.
One limit from the same post: the import pulls in only the file it names, and once a CLAUDE.md exists the default stops discovering AGENTS.md files in subdirectories. A nested AGENTS.md that Codex picks up on its walk needs its own import, or the both-files setting. One of my own repositories had the opposite shape when I wrote that post: an AGENTS.md headed "Codex Instructions" and no CLAUDE.md at all.
@AGENTS.md
## Claude Code
Use plan mode for changes under `src/billing/`.The two tools do not treat the shared file alike. Codex has a 32 KiB cap across the whole chain and takes one file per directory. Claude Code loads an instruction file of up to 4 MiB in full, recommends staying under 200 lines, and loads subdirectory files on demand. A long set of AGENTS.md files can be cut off for one reader and loaded whole for the other. If the two agents seem to be following different rules, check length before you check the models.
How do you move a session or a whole setup between Claude Code and Codex?
Both vendors now ship an importer for the other's configuration, and OpenAI ships one for a live session. All of it is documented. /codex:transfer in the plugin, added in June, does this:
Creates a persistent Codex thread from the current Claude Code session
and prints a `codex resume <session-id>` command.For the setup as a whole, Codex CLI has had /import since August 11, 2026. OpenAI's import docs list what the import flow can bring over: instruction files, settings, skills, plugins, MCP server configuration, hooks, slash commands and subagents. The CLI also imports up to 50 chats from the last 30 days. The same page says to re-check tool restrictions, MCP authentication and hooks afterward, because their behavior may differ.
Claude Code has the mirror image. Its commands reference lists /import codex, available from v2.1.213, which brings over instruction files, MCP servers, commands, subagents and skills, with --dry-run to preview. /init offers to run it when it finds Codex configuration.
Each importer makes a copy. Neither one links the two setups, with one exception: the ChatGPT desktop app can keep an import in sync with automatic updates, and the CLI docs describe no such option. A copied setup starts drifting from the original the day you edit either side. Use it to try the other tool with your real configuration in place. For keeping two configurations equal it is the wrong tool, and for instructions the shared AGENTS.md above does that job better.
Do Claude Code and Codex run under the same sandbox when you combine them?
No, and the plugin does not unify them.
| Codex CLI | Claude Code | |
|---|---|---|
| OS sandbox for commands | on by default | off by default, enabled with /sandbox |
| Network from sandboxed commands | off by default | through a proxy with an allowlist that starts empty, once the sandbox is on |
| Non-interactive default | codex exec runs in a read-only sandbox | claude -p is governed by permission mode and allow rules, with no OS sandbox unless you enabled one |
| Read-only inside the workspace | .git, .agents, .codex | with the sandbox on: .claude settings, hooks and skills, .mcp.json, .git/hooks, .git/config, among others |
Codex's security docs state that the agent runs with network access turned off by default, and its non-interactive docs state that codex exec starts in a read-only sandbox until you pass --sandbox workspace-write. Claude Code's sandboxing page says its Bash sandbox is off by default and covers shell commands only: the file tools, MCP servers and hooks run outside it. Neither sandbox covers the whole tool. Codex's network filter applies to commands inside the sandbox and, per the same docs, not to web search, MCP server connections or connector calls. Out of the box, Claude Code leans on permission rules and modes in front of each tool call, and Codex leans on an operating system boundary around the commands it runs. The same task carries different exposure depending on which tool executes it.
For a split setup this has one concrete consequence, and the README is not where you find it. The README says the plugin uses "the same local authentication state" and "the same repository checkout and machine-local environment" as Codex run directly, and that it applies your Codex configuration. The plugin's source is more specific. As of the July 8 commit it starts Codex threads with the approval policy set to never and the sandbox set to read-only, and its task runner switches the sandbox to workspace-write when a task is started with --write, which is what the rescue agent does by default. So a delegated fix runs with no approval prompts inside Codex's workspace sandbox. Approving the delegation in Claude Code is the only approval there is. That is a reasonable design for a background job, and it means the sandbox is doing all the work: check which directory is the workspace before you delegate a write.
Which Claude Code and Codex integrations stopped working in 2026?
Several, and older tutorials still describe them. The largest is the MCP bridge. For a while the common way to reach Codex from another agent was to launch it as an MCP server. The Codex changelog entry for September 5, 2026 ends that:
The codex mcp-server command and standalone codex-mcp-server binary
have been removed after their deprecation on August 24, 2026. Update
integrations that launch either command before upgrading Codex. Use
the Codex app server for integrations.Any community setup that adds codex mcp-server to Claude Code's MCP configuration stops working on upgrade. The official plugin is not affected, since its README says it wraps the Codex app server. The same changelog entry describes the app-server command as experimental and not supported for production workloads, which is worth knowing before you hang an unattended gate on it.
Three smaller changes affect old scripts and tutorials:
codex exec --full-autois a deprecated compatibility flag that prints a warning. The docs tell you to pass--sandbox workspace-writeexplicitly.approval_policy = "untrusted"is retired in Codex, and the docs say the leftover setting can prevent the client from starting.- In Claude Code,
/reviewhas been an alias of/code-reviewsince v2.1.223. Before that it was a separate command for reviewing a GitHub pull request, so a tutorial that says "run/review" may be describing a different command.
The reverse bridge exists only on paper: Claude Code still documents claude mcp serve, but it exposes Claude Code's tools to an MCP client, not the agent, and no write-up cited here wires it into Codex.
Which Claude Code and Codex split should you start with?
Start with the on-demand review, /codex:review on a finished diff, and stay there until you have read a few dozen findings. This is my read of the sources, not a measurement.
It is the cheapest of the seven to back out of. The command is read-only, it runs when you ask, and nothing in your workflow depends on it. The plugin README says a ChatGPT subscription is enough, the Free plan included, with usage counted against your Codex limits. So the first trial does not have to be a second paid plan. And reading the findings is how you learn whether the second model sees things on your code that the first one misses. That is the only question that justifies running two tools. If the findings mostly repeat what /code-review already told you, you have your answer and it cost you an afternoon.
Turn on the gate last, if at all, and only with a size threshold and a decision about what happens on timeout. Delegate writes only after you have checked which directory Codex will treat as its workspace.
Who should skip the second tool: anyone whose changes are small enough that they already read every line, and anyone who cannot yet tell a good finding from a confident wrong one. A second model adds to the pile you have to judge. It does not judge it for you.
In all seven patterns, what crosses between the tools is an artifact: a diff, a plan file, a findings file, an AGENTS.md, an exported session. None of them shares a live conversation, and the vendor-built bridges all stop at files and threads. If the artifacts are good, either tool can sit on either side of them. The approval does not cross with the artifact. Whichever tool holds the pen writes under its own sandbox, so swapping seats is a configuration change on the file side and a permissions check on the other.