Commit Graph
3735 Commits
Author SHA1 Message Date
pakrym-oaiandGitHub dbd2857f4b [codex] Add optional IDs to response items (#28812)
## Why

`ResponseItem` variants do not have a consistent internal ID shape: some
variants carry required IDs, some carry optional IDs, and some cannot
represent an ID at all. The existing fields also use inconsistent serde,
TypeScript, and JSON-schema annotations. A single enum-level access path
is needed before history recording can assign and retain IDs.

This PR establishes that internal model only. It intentionally does not
generate or serialize IDs; allocation and wire persistence are isolated
in the stacked follow-up.

## What changed

- Give every concrete `ResponseItem` variant an `Option<String>` ID
field.
- Apply the same internal-only annotations to every ID field:
`#[serde(default, skip_serializing)]`, `#[ts(skip)]`, and
`#[schemars(skip)]`.
- Add `ResponseItem::id()` and `ResponseItem::set_id()` as the shared
accessors.
- Preserve IDs when history items are rewritten for truncation.
- Adapt consumers that previously assumed reasoning and image-generation
IDs were required.
- Regenerate app-server schemas so the hidden fields are represented
consistently.

The serde catch-all `ResponseItem::Other` remains ID-less because it
must remain a unit variant.

## Test plan

- `cargo check --tests -p codex-core -p codex-api -p codex-rollout-trace
-p codex-image-generation-extension`
- `just test -p codex-protocol`
- `just test -p codex-app-server-protocol`
- `just test -p codex-api -p codex-rollout-trace -p
codex-image-generation-extension`
- `just test -p codex-core event_mapping`
2026-06-17 18:27:43 -07:00
Owen LinandGitHub 38211a3ff8 [codex] trace tools build latency (#28782)
Add more tracing spans around tool building.
2026-06-17 14:53:54 -07:00
Adam Perry @ OpenAIandGitHub 5867b529ae unified-exec: preserve PathUri through exec-server (#28681)
## Why

It should be possible for app-server to handle "foreign" OS paths in
unified_exec working directories, allowing e.g. a Linux app-server to
run processes on e.g. a Windows exec-server.

## What

Convert the core unified_exec cwd values to use `PathUri`.

Adds fallible path conversion in several places to try to minimize the
scope of this change. The only time this change suppresses errors from
converting `PathUri` to an `AbsolutePathBuf` is when the turn is
configured with no sandboxing at all to allow us to make progress
testing without sandboxing.

Future changes to apply_patch and sandboxing will clean up these error
paths.

A tool's cwd is resolved from joining a model-provided workdir to the
environment's cwd. When using `AbsolutePathBuf::join()`, an
absolute-path workdir would overwrite the environment's cwd and we would
resolve permissions/sandboxing against the model-provided path. This
change extends `PathUri::join()` to also treat an absolute rhs as an
override of the base/lhs.

This also removes some coverage from the remove_env_windows tests until
a follow-up converts foreign paths in command exec events correctly.

## Breaking Changes

When using `AbsolutePathBuf::join()` for workdir resolution, we ended up
resolving tilde-prefixed paths against the app-server's `$HOME`, e.g.
`~/foo/bar` becomes `/home/anp/foo/bar`. It's difficult to do this with
`PathUri` joining, so after offline discussion this PR no longer
implements it.

A quick check of some power users' rollouts suggests that models don't
actually generate home-prefixed absolute working directories for their
spawns, so this shouldn't have any real blast radius.
2026-06-17 19:36:16 +00:00
Jeremy RoseandGitHub 7dc7096ae1 [codex] Restore thread recency with compatible migration history (#28671)
## Summary

- Revert #28655, restoring the thread `recencyAt` behavior introduced by
#27910.
- Move `threads_recency_at` to migration 0039 so it no longer collides
with `external_agent_config_imports` at version 0038.
- Repair databases that already applied the recency migration as version
38 by moving the matching migration-history row to version 39 before
SQLx validation. The current version-38 migration can then apply
normally.

## Validation

- `just test -p codex-state
migrations::tests::repairs_recency_migration_that_was_applied_as_version_38`
- `just test -p codex-state -p codex-rollout -p codex-thread-store -p
codex-app-server-protocol -p codex-tui`: 3,439 passed; six TUI tests
could not open the machine's existing read-only incident database at
`~/.codex/sqlite/state_5.sqlite`.
- `just fix -p codex-state`
- `just fmt`
- Verified that state migration versions are unique.
2026-06-17 18:52:18 +00:00
jifandGitHub 1391d786bc Scope command approvals by execution environment (#28738)
## Why

Command approval cache keys included the command and working directory,
but not the execution environment. An approval for `/workspace` locally
could therefore be reused for the same command and path on an executor.

## What changed

- Include the selected environment ID in shell and unified-exec approval
cache keys.
- Carry that ID through the normal command approval request so clients
can show which environment is being approved.
- Expose the environment through app-server as a required nullable
`environmentId` and show it in the inline TUI approval prompt.
- Keep older recorded approval events compatible when the environment is
absent.

For example, `echo ok` in local `/workspace` and `echo ok` in executor
`/workspace` now produce different approval keys and separate prompts.

## Scope

This PR does not change network approvals, Guardian review actions, MCP
elicitation, full-screen TUI rendering, or environment-ID validation.
Remote `shell_command` execution itself remains in #28722; this PR only
makes its approval key environment-aware.
2026-06-17 19:52:43 +02:00
iceweasel-oaiandGitHub ef75171f18 Run fs helper through Windows sandbox wrapper (#28359)
## Why

This is the final PR in the Windows fs-helper sandbox stack and contains
the actual bug fix.

The exec-server filesystem helper is a direct-spawn path: it asks
`SandboxManager` for a `SandboxExecRequest`, then launches the returned
argv itself. That works on macOS and Linux because the transformed argv
is already a self-contained sandbox wrapper. On Windows, the transformed
request carried `WindowsRestrictedToken` metadata, but the direct-spawn
fs-helper runner still launched the helper argv directly.

That means Windows filesystem built-ins backed by the fs-helper could
run with the parent Codex process permissions instead of the configured
Windows sandbox. This PR makes the direct-spawn transform produce a
self-contained Windows wrapper argv before fs-helper launches it.

## What Changed

- Added `SandboxManager::transform_for_direct_spawn()` for callers that
launch the returned argv themselves.
- Wrapped Windows restricted-token direct-spawn requests with `codex.exe
--run-as-windows-sandbox` and then marked the outer request as
unsandboxed, matching the macOS/Linux wrapper argv shape.
- Updated `exec-server/src/fs_sandbox.rs` to use the direct-spawn
transform for fs-helper launches.
- Materialized the inner `codex.exe --codex-run-as-fs-helper` executable
into `.sandbox-bin` so the sandboxed user can run it.
- Carried runtime workspace roots through `FileSystemSandboxContext` as
`PathUri` values so `:workspace_roots` policies resolve correctly
without sending native client paths over exec-server JSON.
- Preserved wrapper setup identity environment needed by Windows sandbox
setup without changing the serialized inner helper environment.

## Verification

- `just bazel-lock-update`
- `just bazel-lock-check`
- `just test -p codex-sandboxing transform_for_direct_spawn_windows`
- `just test -p codex-exec-server fs_sandbox::tests`
- `just fix -p codex-windows-sandbox -p codex-sandboxing -p
codex-exec-server -p codex-core -p codex-file-system`

Local note: `just fmt` completed Rust formatting, but this workstation
still fails the non-Rust formatter phases because uv cannot open its
cache and the local buildifier/dotslash path is missing.
2026-06-17 10:00:42 -07:00
Alex ZamoshchinandGitHub c78911e37f [ez][codex-rs] Support apps._default.default_tools_approval_mode (#27965)
[from codex]

## Summary

- add `default_tools_approval_mode` to `[apps._default]` and expose it
through app-server v2 `config/read`
- apply it after managed, per-tool, and per-app approval settings,
before the built-in `auto` fallback
- document the precedence, regenerate config/app-server schemas, and add
unit plus end-to-end approval coverage

## Configuration

```toml
[apps._default]
default_tools_approval_mode = "prompt"
```

The effective precedence is managed requirements, tool-specific
`approval_mode`, app-specific `default_tools_approval_mode`,
`apps._default.default_tools_approval_mode`, then `auto`.

## Test plan

- `just write-config-schema`
- `just write-app-server-schema`
- `just write-app-server-schema --experimental`
- `just test -p codex-core app_tool_policy`
- `just test -p codex-core mcp_turn_metadata`
- `just test -p codex-config`
- `just test -p codex-app-server-protocol`
- `just test -p codex-app-server config_read_includes_apps`
- `just fix -p codex-config -p codex-core -p codex-app-server-protocol
-p codex-app-server`
- `just fmt`
2026-06-17 08:50:39 -07:00
jifandGitHub 0318381762 Replace SkillsManager with SkillsService (#28705)
## Why

Host skill discovery was still exposed as a manager even though it is a
process-owned service shared by sessions, the app-server catalog, and
file-watcher invalidation. The skills extension also consumed an ad hoc
loaded-skills wrapper instead of a named immutable snapshot.

## What changed

- replace `SkillsManager` with concrete `SkillsService`
- make the service cache and return immutable `HostSkillsSnapshot`
values
- migrate the skills extension host provider to the snapshot boundary
- migrate app-server catalog, watcher, and invalidation paths to the
service

This keeps the service limited to host discovery, caching, roots, and
invalidation. Catalog rendering and invocation remain extension
responsibilities for the next stacked change.
2026-06-17 17:01:06 +02:00
jifandGitHub 5935a90619 app-server: keep the model cache warm (#28699)
## Why

The app server is long-lived, but its shared model cache otherwise
refreshes only when a caller needs it. Once the five-minute cache
expires, starting a thread or calling `model/list` can wait for
`/models` on the request path.

Refresh the cache in the background before it expires so foreground
callers normally use fresh local state.

## What changed

- Start an app-server worker that refreshes models immediately and then
every three minutes using the existing models-manager API.
- Hold only a weak reference to the models manager between refreshes, so
the worker does not extend its lifetime.
- Stop scheduling refreshes when the app-server lifecycle handle is shut
down or dropped. A refresh already in progress is allowed to finish.
- Adjust affected app-server test fixtures to distinguish the background
`/models` probe from the connection they are testing.

The existing models-manager cache, refresh strategies, auth handling,
ETag behavior, and concurrency semantics are unchanged.

## Testing

-
`models_refresh_worker::tests::refreshes_immediately_periodically_and_stops_when_dropped`
-
`suite::v2::remote_control::listen_off_honors_persisted_remote_control_enable`
-
`suite::v2::attestation::attestation_generate_round_trip_adds_header_to_responses_websocket_handshake`
2026-06-17 16:18:39 +02:00
jifandGitHub 45f603302c Add join key for MAv2 inter-agent messages (#28561)
## Summary
This keeps inter-agent communication on the existing raw response item
path and adds a join key for MAv2 tool calls.

MAv2 `spawn_agent`, `send_message`, and `followup_task` now stamp the
originating tool call id into `ResponseItemMetadata.source_call_id` on
the raw `ResponseItem::AgentMessage`. App-server clients can join that
raw item back to the existing tool/activity event by call id, while
using the raw agent message's existing sender, receiver, and content
fields.

No new app-server `ThreadItem` or notification type is added.

## Tests
- `just fmt`
- `just write-app-server-schema`
- `just test -p codex-protocol`
- `just test -p codex-app-server-protocol`
- `just test -p codex-core
multi_agent_v2_spawn_returns_path_and_send_message_accepts_relative_path`
- `just test -p codex-core
multi_agent_v2_followup_task_completion_notifies_parent_on_every_turn`
- `just fix -p codex-protocol`
- `just fix -p codex-app-server-protocol`
- `just fix -p codex-core`
2026-06-17 14:48:56 +02:00
Won ParkandGitHub 1315198853 [codex] Persist built-in image results reported as generating (#28656)
## Why

#27920 stopped persisting image-generation items unless their status was
`completed`, preventing failed standalone extension items with empty
results from being saved. Built-in image generation can instead emit a
terminal `response.output_item.done` containing a complete base64 PNG
while the item status remains `generating`. In that case, app-server
emits no `savedPath`, so Codex Apps can render the inline image but
cannot expose a file artifact.

## What changed

- Persist image-generation items whenever `result` contains image data.
Failed terminal items still have empty results and remain unpersisted.
- Update the existing built-in image-generation integration test to
cover a terminal `generating` item and verify both `saved_path` and the
written PNG bytes.

## Validation

- Confirmed with a raw built-in websocket trace: the image progressed
through `in_progress`, `generating`, and `partial_image`, then emitted
one `response.output_item.done` with `status: "generating"` and a
complete PNG result.
- `just test -p codex-core builtin_image_generation_call_persisted` is
currently blocked before test execution by a pre-existing compile error
in `thread-store/src/thread_metadata_sync.rs:171`.
2026-06-17 06:03:00 +00:00
pakrym-oaiandGitHub 172b2218a5 core: remove redundant TurnContext and Prompt fields (#28638)
## Why

`TurnContext` had accumulated dead fields and cached projections of
values already owned by its per-turn `Config` or `ModelInfo`. Keeping
both copies made ownership unclear and allowed artificial split-brain
states, such as a compatibility hash differing from the model metadata
it came from.

`Prompt` similarly carried a write-only personality after personality
selection had already been materialized into its base instructions.

This makes the canonical owner explicit: configuration-backed values
come from `config`, model-derived values come from `model_info`, and
prompts contain only data consumed by request construction.

## What changed

- Remove the unused `ghost_snapshot`, `codex_self_exe`, and
`thread_source` fields.
- Remove duplicate `comp_hash`, `truncation_policy`, `features`,
`shell_environment_policy`, `codex_linux_sandbox_exe`, `compact_prompt`,
and `tool_mode` fields.
- Read those values directly from `TurnContext::config` or
`TurnContext::model_info` at their consumers.
- Remove the write-only `Prompt::personality` field and its constructor
assignments.
- Preserve review-turn inheritance of the parent turn's shell policy,
Linux sandbox executable, and compact prompt through the review config.

## Testing

- `cargo check -p codex-core --tests`
2026-06-16 22:17:24 -07:00
pakrym-oaiandGitHub cb15c64760 Revert thread recencyAt for sidebar ordering (#28655)
## Why

Revert #27910 to remove the newly introduced thread `recencyAt`
persistence and API behavior from `main`.

## What changed

This reverts commit `fac3158c2a783095768076489815f361fa9b0db4`,
including the state migration, thread-store propagation, app-server API
surface, generated schemas, and related tests.

## Validation

Not run before opening; relying on CI for the initial fast signal.
2026-06-16 21:39:30 -07:00
Ahmed IbrahimandGitHub 0a3ad4c4ba [codex] Test code-mode variable truncation (#28471)
## Summary

Code mode has two separate truncation points: the nested tool result
returned to JavaScript and the code-mode output later recorded for the
model. These tests now verify those behaviors independently.

- Report whether `result.output` was truncated before printing it.
- Verify omitted or sufficiently large nested limits produce `Variable
truncated: False`, while allowing the printed value to be truncated
downstream.
- Verify an explicit nested limit produces `Variable truncated: True`
when the command output exceeds it.
- Use a token-policy model fixture so downstream truncation is visible
as `…N tokens truncated…`.
- Align the explicit nested-truncation expectation with the warning
header.

This PR changes test coverage only; runtime truncation behavior is
unchanged.

## Validation

- `env -u CODEX_SANDBOX_NETWORK_DISABLED RUST_MIN_STACK=8388608 cargo
test -p codex-core --test all code_mode_exec -- --nocapture` (8 passed)
2026-06-16 20:14:23 -07:00
Adam Perry @ OpenAIandGitHub 3d02c443bc [codex] core: restore absolute turn context cwd (#28629)
## Why

#28152 jumped the gun on moving the rollout format to store URIs, and
would likely break compat with some features that don't go through the
same types as the core logic.

## What

Make `TurnContextItem.cwd` an `AbsolutePathBuf` again, remove test added
for `PathUri` serialization in rollouts. Also drops a bunch of error
paths that are no longer needed.
2026-06-16 19:05:26 -07:00
Jeremy RoseandGitHub fac3158c2a Add thread recencyAt for sidebar ordering (#27910)
## Summary

Add a server-owned `recencyAt` timestamp and `recency_at` thread-list
sort key for product recency ordering while preserving the existing
meaning of `updatedAt` as the latest persisted thread mutation.

This is the server-side alternative to #27697. Rather than narrowing
`updatedAt`, clients can sort the sidebar by `recency_at` and continue
treating `updatedAt` as mutation time.

Paired Codex Apps PR:
[openai/openai#1024599](https://github.com/openai/openai/pull/1024599)

## Contract

- `recencyAt` initializes when a thread is created.
- A turn start advances `recencyAt` monotonically.
- Commentary, agent output, tool results, token/accounting updates, turn
completion, archive, unarchive, resume, and generic metadata writes do
not advance it.
- `updatedAt` retains its existing behavior and continues to advance for
persisted thread mutations.
- Current servers populate `recencyAt`; the response field is optional
in generated TypeScript so clients connected to older servers can fall
back to `updatedAt`.
- Filesystem-only fallback uses existing updated/mtime ordering when
SQLite is unavailable.

## Persistence and compatibility

Migration 0038 adds second- and millisecond-precision recency columns,
backfills them from the existing updated timestamp, creates list
indexes, and includes an insert trigger so older binaries writing to a
migrated database seed recency without causing later mutations to
advance it.

Generic metadata upserts preserve existing recency values. Turn-start
updates use a dedicated monotonic touch, and process-local allocation
keeps millisecond cursor values unique. State DB list, search, read,
filtered-list repair, rollout fallback propagation, and app-server
conversions all carry the new field.

## API

`Thread` responses include:

```ts
recencyAt?: number
```

`thread/list` and `thread/search` accept:

```json
{ "sortKey": "recency_at" }
```

Generated TypeScript and JSON schemas are included.

## Validation

- `just test -p codex-state` — 146 passed
- `just test -p codex-rollout` — 69 passed
- `just test -p codex-thread-store` — 81 passed
- `just test -p codex-app-server-protocol` — 231 passed
- Focused app-server list ordering, response mapping, archive/unarchive,
and resume lifecycle tests passed
- Scoped `just fix` for state, rollout, thread-store,
app-server-protocol, and app-server
- `just fmt`
- `git diff --check`
- Independent correctness, simplicity, elegance, security, and
test-quality reviews; actionable ordering, lifecycle, query-projection,
and timestamp-uniqueness findings were addressed
2026-06-16 17:06:22 -07:00
canvrno-oaiandGitHub f0cb96bcb1 PAC 1 - Add system proxy feature config surface (#26706)
## Summary

Introduces the default-off `respect_system_proxy` feature flag used to
gate first-class system PAC/proxy support for Codex-owned native
clients.

With the feature disabled or absent, behavior remains unchanged. This PR
establishes the configuration and managed-requirement surface; proxy
discovery and request routing are implemented by follow-up PRs.

## Configuration

User configuration uses the standard boolean feature form:

```toml
[features]
respect_system_proxy = true
```

Managed feature requirements use the corresponding boolean key. The
effective runtime configuration is exposed as a boolean and defaults to
`false`.

## Implementation

- Registers `respect_system_proxy` as an under-development, default-off
feature.
- Resolves user configuration and managed feature requirements into
`Config.respect_system_proxy`.
- Provides bootstrap resolution for startup paths that must evaluate the
feature before full configuration loading completes.
- Uses the standard feature CLI and config-editing behavior.
- Excludes `features.respect_system_proxy` from project-local
configuration.
- Updates the generated configuration schema.

## End-user behavior

- No networking behavior changes when the feature is absent or disabled.
- Enabling the feature makes the boolean available to the native
proxy-routing implementation in follow-up PRs.
- Repository-local configuration cannot enable the feature.

## Test coverage

Covers scalar configuration and CLI override resolution, managed
requirement constraints, bootstrap resolution, and project-local
filtering.
2026-06-16 16:54:37 -07:00
Alex DaleyandGitHub a397b59887 [codex] [4/4] Simplify recommended plugin install schema (#28403)
## Summary
- Simplify recommendation-context `request_plugin_install` arguments to
`plugin_id` and `suggest_reason`.
- Derive plugin type and install action from the matched candidate while
preserving Codex-owned elicitation metadata.
- Keep the legacy list-backed schema unchanged and accept resumed calls
that still use `tool_id`.

## Stack
- #28399
- #28400
- #27704
- This PR

## Validation
- `just test -p codex-tools -p codex-core request_plugin_install` (25
passed)
- `just fix -p codex-tools -p codex-core`
- `just fmt`
- `git diff --check`
2026-06-16 23:44:42 +00:00
Adam Perry @ OpenAIandGitHub 4c79527e31 core: render remote environment cwd natively (#28152)
## Why

Model-visible `<environment_context>` should match the environment of
the executor, not of the app server.

Stacked on #28146.

## What

- Keep selected environment cwd values as `PathUri` while building
environment context.
- Render cwd text using the path convention represented by the URI, with
the canonical URI as a fallback.
- Preserve compatibility with legacy `TurnContextItem.cwd` values when
reconstructing and diffing context.
- Extend the Wine-backed remote Windows test to assert that the model
sees `powershell` and `C:\windows`.
2026-06-16 16:17:47 -07:00
Alex DaleyandGitHub a34da3b295 [codex] [3/4] Activate endpoint plugin recommendations (#27704)
Summary\n- Await endpoint recommendation selection while constructing
each authenticated turn, removing the first-turn cache race.\n- Snapshot
and filter endpoint candidates once per turn, then use that same set for
the bounded contextual user fragment, tool exposure, and exact install
validation.\n- Keep recommendation selection ephemeral: do not persist
recommendation state in or gate resumed threads on prior context.\n-
Hide the legacy list tool in endpoint mode and preserve legacy discovery
unchanged when the endpoint is disabled or unavailable.\n- Keep remote
plugin and connector app identities out of model-visible context and
attach them only to Codex-owned elicitation metadata.\n\nStack\n- 3/4,
based on #28400.\n- Endpoint client and cache: #28399.\n- Generalized
suggestion presentation: #28400.\n- Install-schema follow-up:
#28403.\n\nValidation\n- \n- \n- \n- \n- Full : 2,649 passed and 88
environment-dependent tests failed because this sandbox cannot write ,
nest Seatbelt, or locate auxiliary test binaries.
2026-06-16 23:04:07 +00:00
Alex DaleyandGitHub 587487df9e [codex] [2/4] Generalize plugin suggestion presentation (#28400)
Summary
- Add list-backed and developer-context presentations for plugin
suggestion candidates.
- Let tool planning, install validation, and request-tool copy follow
the selected presentation.
- Keep every production caller on the existing list-backed presentation,
preserving the current list tool, request schema, connector behavior,
and model-visible copy.
- Leave developer-context presentation latent until the final PR in the
stack.

Stack
- 2/3, based on #28399.
- Follow-up: #27704 activates endpoint recommendations.

Validation
- `just test -p codex-core request_plugin_install`
- `just test -p codex-core spec_plan`
- `just fix -p codex-core`
- `just fmt`
- `git diff --check`
2026-06-16 22:44:10 +00:00
Adam Perry @ OpenAIandGitHub f8850cab1d app-server: preserve target-native environment cwd (#28146)
## Why

app-server may run on a different OS from the selected exec-server
environment. Parsing that environment’s cwd with the Codex host’s path
rules prevents thread startup.

## What

Carry environment cwd values as `LegacyAppPathString` at the app-server
boundary and `PathUri` internally. Existing tool-call schemas and
relative-path behavior stay host-native; remaining local-only consumers
convert explicitly and leave follow-up TODOs.

The Wine integration test verifies app-server can start a thread and
complete an ordinary turn with a Windows environment cwd from Linux.

## Validation

- `bazel test //codex-rs/core/tests/remote_env_windows:smoke-test
--test_output=errors`
- focused app-server environment-selection and protocol schema tests
- scoped Clippy for `codex-core` and `codex-app-server-protocol`
2026-06-16 21:42:28 +00:00
Adam Perry @ OpenAIandGitHub 322b83de5e Clarify model-generated and legacy app path types (#28577)
## Why

`ApiPathString` kind of implies that it can be used anywhere we pull a
path out of JSON, but it's not really appropriate for tool arguments
when the model might generate relative paths.

Prefer `String` for model-generated paths and we can handle the
conversion per feature for now and define a shared abstraction later if
it makes sense.

# What

Rename `ApiPathString` to `AppLegacyPathString` to clarify its role.

Expand the `path-types` skill to tell the model to leave tool args as
bare strings.
2026-06-16 20:47:43 +00:00
Adam Perry @ OpenAIandGitHub a50671e748 [codex] test exec relative additional permissions (#28587)
## Why

Review caught some would-be regressions in changes to unified_exec that
weren't surfaced in CI.

## What

Add coverage for requesting permissions through unified exec when there
are additional permissions. Previously this flow was only tested against
shell_command.
2026-06-16 20:45:57 +00:00
Channing CongerandGitHub e93516e259 code-mode: extend test coverage to lock in cell lifecycle (#28468)
This PR establishes the intended behavior as an executable contract
before a refactor of the cell runtime begins. It also fixes cases where
a second observer or termination request could replace an existing
response channel and leave the original caller unresolved.

### Behavior codified
- A cell can yield output and subsequently resume to completion.
- A caller can run a cell until it has no immediately runnable work,
receive its accumulated output and outstanding tool-call IDs, and then
resume the same cell when the awaited work is available.
- Each cell admits one active observer:
   - a second observer receives an explicit busy error
   - the existing observer remains registered and is not displaced
- A natural result (conclusion of the js module) that has already
reached the cell controller wins over a later termination request.
- Otherwise, termination preempts execution and resolves both:
  - the active observer, if present
  - the caller requesting termination
- Repeated termination requests are rejected while termination is
already in progress.
- Terminal responses are sent only after outstanding callback work has
been handled:
- natural completion drains notifications and cancels outstanding tool
calls
- termination cancels and drains both notification and tool callbacks.
- Cell removal and cell_closed notification happen after callback
cleanup
2026-06-16 13:34:16 -07:00
Adam Perry @ OpenAIandGitHub 4b7351700f [codex] re-enable absolute workdir integration test (#28581)
## Why

In #28146 I missed the invariant that an absolute `exec_command` workdir
must override the environment cwd. The existing integration test would
have caught that regression, but it was ignored as flaky.

## What

Re-enable `unified_exec_respects_workdir_override`.

## Validation

`just test -p codex-core unified_exec_respects_workdir_override`
2026-06-16 20:19:41 +00:00
pakrym-oaiandGitHub 7baf7e467e [codex] Route MCP file uploads through environment filesystem (#27923)
## Why

Codex Apps tools can mark arguments with `openai/fileParams`, but the
execution path resolved and opened those files directly on the host.
That bypassed the selected turn environment and prevented annotated file
arguments from working with remote environments.

## What changed

- resolve annotated file arguments against the primary turn environment
- read file metadata and contents through that environment's sandboxed
`ExecutorFileSystem`
- reject files over the 512 MiB limit from metadata before reading or
transferring them
- retain the buffered upload-size check as defense in depth
- make the OpenAI upload API accept a filename and buffered contents
instead of owning local filesystem access
- describe the model-visible argument as a path in the primary
environment

This builds on #27927, which added `size` to internal filesystem
metadata.

## Testing

- `just test -p codex-api upload_openai_file_returns_canonical_uri`
- `just test -p codex-mcp
tool_with_model_visible_input_schema_masks_file_params`
- `just test -p codex-core mcp_openai_file`
- `just test -p codex-core
codex_apps_file_params_upload_environment_files_before_mcp_tool_call`
2026-06-16 11:27:46 -07:00
Ahmed IbrahimandGitHub 952656356a [codex] Warn clearly when code mode output is truncated (#28467)
## Summary

- make `formatted_truncate_text` prepend `Warning: truncated output
(original token count: N)` above the existing `Total output lines`
header
- update direct formatter, unified-exec, user-shell, and code-mode
expectations
- add core unit coverage that runs in Bazel without requiring the
skipped V8-backed code-mode integration suite

## Validation

- `cargo test -p codex-utils-output-truncation -- --nocapture` (17
passed)
- `cargo test -p codex-core --lib
truncated_text_output_starts_with_warning -- --nocapture`
- `cargo test -p codex-core --test all
clamps_model_requested_max_output_tokens_to_policy -- --nocapture` (2
passed)
- `cargo test -p codex-core --test all
unified_exec_formats_large_output_summary -- --nocapture`
- `cargo test -p codex-core --test all
user_shell_command_output_is_truncated_in_history -- --nocapture`
- Bazel CI exercises the shared formatter and downstream integration
expectations
2026-06-16 10:37:06 -07:00
pakrym-oaiandGitHub a4711b88dd [codex] exec-server: stream files in chunks (#28354)
## Why

`fs/readFile` buffers the entire file in one response, which makes large
remote reads expensive and prevents callers from applying backpressure.
We need an opt-in streaming path with bounded block sizes while
preserving the existing single-call API for small and sandboxed reads.

## What changed

- Add `ExecServerClient::stream`, returning a named `FileReadStream`
that implements `futures::Stream` and yields immutable 1 MiB byte
blocks.
- Add internal `fs/open`, `fs/readBlock`, and `fs/close` RPCs.
`fs/readBlock` accepts an explicit offset and length.
- Keep unsandboxed files open between block reads, cap open handles per
connection, and clean them up on EOF, error, stream drop, explicit
close, or connection shutdown.
- Reject platform-sandboxed streaming opens instead of turning the
one-shot sandbox helper into a persistent server. Existing `fs/readFile`
behavior is unchanged.

## Testing

- `just test -p codex-exec-server`
- Integration coverage for 1 MiB chunking, exact block-boundary EOF,
sandbox rejection, and continued reads from the opened file after path
replacement.
- Handle-manager coverage for non-sequential offsets, variable block
lengths, the 128-handle limit, and capacity release after close.
2026-06-16 09:50:55 -07:00
jifandGitHub 1b24ba912a core: surface terminal subagent errors to parent agents (#28375)
## Why

When a subagent exhausts its retries, it emits an `Error`, but the
generic task lifecycle then emits `TurnComplete(None)`. That completion
used to overwrite the subagent's `Errored` status with
`Completed(None)`, so the parent received an empty completion
notification.

This made a failed child look indistinguishable from a child that
completed without an answer. In unattended or long-running multi-agent
work, the root could silently continue without knowing that delegated
work failed or how to restart it.

## Behavior

Before, a terminal stream failure was reduced to an empty completion:

```text
<subagent_notification>
{"agent_path":"/root/worker","status":{"completed":null}}
</subagent_notification>
```

Now the parent receives the actual terminal error, bounded to 1,000
tokens, together with an actionable recovery hint:

```text
<subagent_notification>
{
  "agent_path": "/root/worker",
  "status": {
    "errored": "stream disconnected before completion: stream closed before response.completed"
  },
  "next_action": "This agent's turn failed. If you still need this agent, use `followup_task` to give it another task."
}
</subagent_notification>
```

The notification remains queue-only: it does not wake the root or replay
the failed request. The root sees it at the next sampling boundary and
can use `followup_task` to start a new turn for that agent.

## What changed

- Added terminal-error precedence to the [agent status
reducer](https://github.com/openai/codex/blob/e95fcfe2bb6a02f1a75650afa20048859f556511/codex-rs/core/src/agent/status.rs#L23-L34),
so a closing `TurnComplete` cannot erase an immediately preceding
`Errored` status.
- Made MultiAgentV2 completion forwarding use the retained session
status instead of re-deriving `Completed(None)` from the final event.
- Extended the [subagent notification
fragment](https://github.com/openai/codex/blob/e95fcfe2bb6a02f1a75650afa20048859f556511/codex-rs/core/src/context/subagent_notification.rs#L6-L60)
with a `next_action` for terminal errors and a hard cap on model-visible
error text.
- Kept successful completions and interrupted turns unchanged.

## Verification

- Added a status-reducer test proving that `Errored` survives the
trailing `TurnComplete`.
- Added an integration test that exhausts a subagent's stream retries
and verifies the exact `agent_message` delivered to the parent,
including the error and `followup_task` guidance.
- Re-ran the existing successful-completion and interrupted-turn
notification tests.
2026-06-16 14:34:54 +02:00
jifandGitHub ef8eb8bdd9 [tests] Keep Apps out of generic core test harness (#28508)
## Summary

- disable the stable Apps feature in the generic `test_codex()`
integration-test harness
- keep Apps-specific tests explicit: their builders re-enable Apps and
point it at a local mock server

## Why

Generic tests that use dummy ChatGPT auth were also enabling the
host-owned `codex_apps` MCP server. That made unrelated tests contact
`chatgpt.com` and wait for MCP startup, causing the Bazel timeouts
observed on #28368.

The generic harness should be hermetic and should not start an external
service that the test did not request. This is test-only; production
Apps behavior is unchanged. The broader optional-MCP startup behavior is
being handled separately in #28407.

## Testing

- `just test -p codex-core -E
'test(pre_sampling_compact_runs_when_comp_hash_changes) |
test(model_switch_to_smaller_model_updates_token_context_window) |
test(codex_apps_file_params_upload_local_paths_before_mcp_tool_call)'`
- `just fix -p codex-core`
- `just fmt`
2026-06-16 13:07:43 +02:00
jifandGitHub 5b22a8e5b1 feat: render typed envelopes for multi-agent v2 messages (#28368)
## Why

Multi-agent v2 messages need a consistent, model-visible envelope that
identifies what kind of interaction occurred, who sent it, and which
agent it targets. Previously, encrypted deliveries exposed only
`encrypted_content`, while child completion used the legacy
`<subagent_notification>` shape. That meant the client could not
consistently present `NEW_TASK`, `MESSAGE`, and `FINAL_ANSWER` using the
same format.

This change adds the routing envelope as plaintext while keeping task
and message payloads encrypted. No new Responses API field is required:
an encrypted delivery is represented as an `input_text` header
immediately followed by its existing `encrypted_content` item.

Every envelope now follows this shape:

```text
Message Type: <NEW_TASK | MESSAGE | FINAL_ANSWER>
Task name: <recipient agent path>
Sender: <author agent path>
Payload:
<message payload>
```

## Message types

### `NEW_TASK`

`NEW_TASK` is used when the recipient should begin a new turn, including
an initial `spawn_agent` task and a later `followup_task`.

For a root agent spawning `/root/worker`, the request contains a
plaintext envelope followed by the encrypted task:

```json
{
  "type": "agent_message",
  "author": "/root",
  "recipient": "/root/worker",
  "content": [
    {
      "type": "input_text",
      "text": "Message Type: NEW_TASK\nTask name: /root/worker\nSender: /root\nPayload:\n"
    },
    {
      "type": "encrypted_content",
      "encrypted_content": "<encrypted task payload>"
    }
  ]
}
```

Conceptually, the model receives:

```text
Message Type: NEW_TASK
Task name: /root/worker
Sender: /root
Payload:
Review the authentication changes and report any regressions.
```

### `MESSAGE`

`MESSAGE` is used for a queued `send_message` delivery. It communicates
with an existing agent without starting a new turn.

For `/root/worker` reporting progress to the root agent, the request
contains:

```json
{
  "type": "agent_message",
  "author": "/root/worker",
  "recipient": "/root",
  "content": [
    {
      "type": "input_text",
      "text": "Message Type: MESSAGE\nTask name: /root\nSender: /root/worker\nPayload:\n"
    },
    {
      "type": "encrypted_content",
      "encrypted_content": "<encrypted message payload>"
    }
  ]
}
```

Conceptually, the model receives:

```text
Message Type: MESSAGE
Task name: /root
Sender: /root/worker
Payload:
The protocol tests pass; I am checking the resume path now.
```

### `FINAL_ANSWER`

`FINAL_ANSWER` is emitted when a child agent reaches a terminal state
and reports its result to its parent. Completion payloads are already
available locally, so the complete envelope is represented as plaintext
rather than as a plaintext header plus encrypted content.

For `/root/worker` completing work for the root agent, the request
contains:

```json
{
  "type": "agent_message",
  "author": "/root/worker",
  "recipient": "/root",
  "content": [
    {
      "type": "input_text",
      "text": "Message Type: FINAL_ANSWER\nTask name: /root\nSender: /root/worker\nPayload:\nNo regressions found."
    }
  ]
}
```

The model-visible form is:

```text
Message Type: FINAL_ANSWER
Task name: /root
Sender: /root/worker
Payload:
No regressions found.
```

Errored, shut down, and missing agents also use `FINAL_ANSWER`, with a
terminal-status description in the payload.

## What changed

- Render `NEW_TASK` or `MESSAGE` in
`InterAgentCommunication::to_model_input_item`, based on whether the
encrypted delivery starts a turn.
- Replace the multi-agent v2 `<subagent_notification>` completion
payload with a model-visible `FINAL_ANSWER` envelope.
- Document `Task name`, `Sender`, and `Payload` consistently in the
multi-agent developer instructions.
- Prevent local-only history projections from treating an encrypted
message's plaintext header as the complete assistant message.
- Preserve rollout-trace interaction edges when an agent message
contains both plaintext and encrypted content.

Legacy multi-agent behavior remains unchanged.

## Verification

- `just test -p codex-protocol`
- `just test -p codex-rollout-trace`
- `just test -p codex-web-search-extension`
- `just test -p codex-core
encrypted_multi_agent_v2_spawn_sends_agent_message_to_child`
- `just test -p codex-core
plaintext_multi_agent_v2_completion_sends_agent_message`
- `just test -p codex-core
multi_agent_v2_followup_task_completion_notifies_parent_on_every_turn`
- `just test -p codex-core
multi_agent_v2_completion_queues_message_for_direct_parent`
2026-06-16 11:46:59 +02:00
pakrym-oaiandGitHub 1e015884c5 [codex] Use local environment for user shell commands (#28163)
## Why

User shell commands still read the legacy turn cwd and session shell
even though execution context is now owned by selected turn
environments. App-server also defines `thread/shellCommand` as a
local-host escape hatch, so it must use an available local environment
even when a remote environment is primary.

## What changed

- Add `ResolvedTurnEnvironments::local()` to find the selected local
environment.
- Resolve the user shell command cwd and shell from that local
`TurnEnvironment`.
- Emit the standard `shell is unavailable in this session` error when no
selected local environment or resolved local shell is available.
- Add an integration test covering `/shell` without a local environment.

## Test plan

- `just test -p codex-core
user_shell_command_without_local_environment_emits_error`
2026-06-16 04:55:20 +00:00
pakrym-oaiandGitHub e752f7b4ae [codex] Use expect in integration tests (#28441)
The workspace denies `clippy::expect_used` in production. Although
`clippy.toml` allows `expect` in tests, Bazel Clippy compiles
integration-test helper code in a way that does not receive that
exemption, which encouraged verbose `unwrap_or_else(... panic!(...))`
and equivalent `match`/`let else` forms.

This allows `clippy::expect_used` once at each integration-test crate
root (including aggregated suites and test-support libraries), then
replaces manual panic-based Result and Option unwraps with
`expect`/`expect_err`. Standalone `tests/*.rs` files remain their own
crate roots. Intentional assertion and unexpected-variant panics remain
unchanged, and the production `expect_used = "deny"` lint remains in
place.

The cleanup is mechanical and net-negative in line count.
2026-06-15 21:53:47 -07:00
pakrym-oaiandGitHub 08901fc8e1 [codex] Add interruptible sleep tool (#28429)
## Why

Models sometimes need to pause briefly while waiting for external work,
but using a shell command for that delay ties the wait to a process and
does not naturally resume when new turn input arrives.

## What changed

- add a built-in `sleep` tool behind the under-development `sleep_tool`
feature
- accept a bounded `duration_ms` argument, matching the millisecond
convention used by unified exec
- end the sleep early when either steered user input or mailbox input
arrives
- include elapsed wall-clock time in completed and interrupted outputs
- emit a dedicated core `SleepItem` through `item/started` and
`item/completed`
- expose the sleep item as app-server v2 `ThreadItem::Sleep` and retain
it in reconstructed thread history
- regenerate the configuration schema for the new feature flag
- regenerate app-server JSON and TypeScript schema fixtures

## Test plan

- `just test -p codex-core sleep_tool_follows_feature_gate`
- `just test -p codex-core any_new_input_interrupts_sleep`
- `just test -p codex-app-server-protocol`
- `just test -p codex-app-server
sleep_emits_started_and_completed_items`
2026-06-15 21:39:21 -07:00
pakrym-oaiandGitHub 022f1221e8 [codex] Bind shell snapshots to retained thread environments (#28421)
## Why

Shell snapshots are currently session-scoped even though shell and cwd
are properties of a selected turn environment. That makes snapshot
refresh depend on separate session-cwd plumbing, prevents retained
environments from retaining their snapshot work, and can make snapshot
construction use a different shell than command execution.

This follows #27955 by making the retained thread-environment service
own environment snapshot lifecycles. Session configuration remains the
requested selection state, while `ThreadEnvironments` remains the source
of successfully resolved environments.

## What changed

- Configure the shell-snapshot builder before initial environment
resolution.
- Start each local environment snapshot task when its `TurnEnvironment`
is built and retain that shared task while environment ID and cwd still
match.
- Inherit retained environment snapshots into spawned child threads.
- Carry the selected `TurnEnvironment` through shell runtimes so
snapshot construction and command execution use the same
environment-specific shell and cwd.
- Load project instructions and warm plugins/skills after initial
environment resolution.
- Continue decoding invalid UTF-8 instruction files lossily without
emitting a startup warning.
- Keep requested selections in `SessionConfiguration`; failed or
duplicate resolutions only affect the resolved environment snapshot.

## Validation

- `cargo check -p codex-core --tests`
- `just test -p codex-home instructions` (6 passed)
- Focused environment, instruction, shell-snapshot, and user-shell tests
(84 passed)
- Focused shell-snapshot, user-shell, and unified-exec tests (126
passed; two event-timing tests passed on retry)
2026-06-15 20:10:53 -07:00
Adam Perry @ OpenAIandGitHub 1fe89de576 Run core integration tests against a Wine-backed Windows executor (#28401)
## Why

We want to exercise a linux app-server against a windows exec-server
without having to repeat every test case. This approach has slight
precedent in the remote docker test setup.

## What

Run the shared `codex-core` integration suite against Windows
exec-server behavior from Linux. This makes cross-OS path and shell
regressions visible while keeping unsupported cases owned by individual
tests.

- Add `local`, `docker`, and `wine-exec` test environment selection with
legacy Docker compatibility.
- Extend `codex_rust_crate` to generate a sharded Wine-exec variant
using a cross-built Windows server and pinned Bazel Wine/PowerShell
runtimes.
- Teach remote-aware helpers about Windows paths and track temporary
incompatibilities with source-local `skip_if_wine_exec!` calls and
follow-up reasons.
2026-06-16 00:38:41 +00:00
guinness-oaiandGitHub d5b4b98370 Add a toggle for realtime startup context (#28405)
## Summary
- Add `includeStartupContext` to realtime start requests so callers can
explicitly skip Codex startup context while keeping the backend prompt
- Thread the new flag through protocol types, request processing, and
realtime session config
- Update app-server docs and coverage for the new default and opt-out
behavior

## Testing
- Added protocol serialization coverage for `includeStartupContext`
- Added realtime integration coverage for starting a session with
startup context disabled
2026-06-15 17:14:22 -07:00
Alex DaleyandGitHub 515f705426 [codex] Fix missing response item metadata in tests (#28415)
Summary
- Add the two missing `metadata: None` initializers after #28355 made
response-item metadata required.
- Restore test compilation for `codex-core` and `codex-api` on main.

Validation
- `git diff --check`
- `just fmt` (Rust formatting passed; unrelated Python formatter steps
could not use the sandboxed shared `uv` cache)
- Focused crate tests are running after PR creation.
2026-06-16 00:08:39 +00:00
Adam Perry @ OpenAIandGitHub 46f17930b6 Use PathUri in filesystem permission paths for exec-server (#28165)
## Why

Progress towards letting app-server and exec-server run on different
platforms, specifically for sandbox configuration.

## What

- Make the filesystem path containment hierarchy generic, defaulting to
`AbsolutePathBuf` for now.
- Have clients specify `AbsolutePathBuf` or `PathUri` directly where
needed.
- Use `PathUri` throughout exec-server filesystem protocol and trait
boundaries.
- Implement `From` for conversion to path URIs and `TryFrom` for
fallible conversion to absolute paths through the generic type
hierarchy.
2026-06-15 23:55:23 +00:00
guinness-oaiandGitHub 1d8ff89aa3 Add realtime speech append control (#27917)
## Why

Realtime voice harness tuning needs app-side control over what backend
Codex text is spoken. Backend orchestrator text is written for a reading
UI, so automatically speaking every preamble, progress update, or final
assistant message can make the realtime voice model too chatty.

For experimentation, clients need two simple controls: keep app/client
text-item injection on the existing item-create path, and add an
explicit speakable path that app code can call only when it wants
realtime to speak. Automatic Codex output also needs an opt-in way to
switch from the protocol's default speakable path to regular realtime
items, with a caller-provided prefix so prompt wording can be tuned
outside core.

The default remains unchanged: if a client omits the new start fields
and never calls `appendSpeech`, automatic backend output continues down
the existing speakable path for the selected realtime protocol.

## What Changed

- Adds experimental `thread/realtime/appendSpeech` for app-provided
speakable text.
- Keeps existing `thread/realtime/appendText` as the item-create API for
app-provided realtime text items.
- Adds `codexResponsesAsItems` / `codex_responses_as_items` on
`thread/realtime/start` to send automatic Codex responses with
`conversation.item.create` instead of the protocol's default speakable
output path.
- Adds `codexResponseItemPrefix` / `codex_response_item_prefix` so
clients can prepend experiment instructions to those automatic Codex
response items.
- Keeps literal `conversation.handoff.append` routing scoped to the v1
speakable path; v2 default speech uses its item/function-output plus
`response.create` behavior.
- Removes the earlier public silent-context API and hardcoded
silent-context prefix.
- Updates realtime tests to cover default automatic speakable behavior,
opt-in automatic item-create behavior, and explicit `appendSpeech`
behavior.

## Validation

- `cargo check -p codex-core -p codex-app-server -p codex-api`
- `just test -p codex-app-server realtime_conversation`
- `just test -p codex-core realtime_conversation` (50/51 passed in the
filtered parallel run; the lone failure passed when rerun in isolation)
- `just test -p codex-core
conversation_mirrors_assistant_message_text_to_realtime_handoff`
- `just test -p codex-api
e2e_connect_and_exchange_events_against_mock_ws_server`
- `just fix -p codex-core`
- `just fix -p codex-app-server`
- `cargo build -p codex-cli`
2026-06-15 16:15:58 -07:00
pakrym-oaiandGitHub 9728992fab [codex] retain resolved environments across turns (#27955)
## Why

Selected execution environments are thread-scoped resources, but startup
and turn construction repeatedly resolved their IDs and working
directories. That discarded existing environment handles and shell
metadata even when a selection had not changed.

Session configuration updates also need to affect future turns without
changing the resolved environment set already captured by a running
turn.

## What changed

- Create a `ThreadEnvironments` service inside `Codex` from the spawned
`EnvironmentManager` and raw environment selections, then store it on
`SessionServices`.
- Split service construction from `update_selections`, allowing session
configuration updates to mutate the resolved set in place.
- Retain an existing `TurnEnvironment` when its environment ID and
working directory match; resolve only added or changed selections and
remove selections that are no longer present.
- Normalize duplicate IDs by keeping the first selection and skip
individual selections that fail to resolve instead of rejecting the
entire update.
- Give each `TurnContext` a cloned `TurnEnvironmentSnapshot`, so later
session configuration updates affect future turns without rewriting an
active turn.
- Reuse the service-owned environment manager and resolved snapshot for
startup work, MCP initialization, and child-thread spawning instead of
flowing resolved environments through spawn arguments.

## Test plan

- `cargo check -p codex-core --tests`
- `just test -p codex-core environment_selection`
- `just test -p codex-core turn_environments`
- `just test -p codex-core
session_update_settings_does_not_rewrite_sticky_environment_cwds`
- `just test -p codex-core
default_turn_does_not_overlay_legacy_fallback_cwd_onto_stored_thread_environments`
2026-06-15 16:15:07 -07:00
felixxia-oaiandGitHub db8927a8e2 Deflake realtime handoff steering test (#28300)
## Summary
- keep the realtime mock websocket open for the handoff steering test
after scripted responses
- avoid racing the mock server close before the standalone handoff
append is observed, which was showing up as a Windows timeout in CI

__Details__:
Failures in samples seem to be caused by:
1. The mock websocket sends conversation.handoff.requested.
2. The mock immediately closes the websocket because
start_websocket_server(...) defaults to close_after_requests: true.
3. On Windows, that close often surfaces as os error 10053 / 10054.
4. The realtime stream shuts down before the routed handoff finishes
creating/steering the follow-up request.
5. The test waits for the expected follow-up event and times out.

The PR changes only step 2: for this test, the mock websocket stays open
after sending the scripted handoff event. The same handoff event is
still sent, and the test still asserts the important steering behavior:
1. first Responses request has the original prompt
2. first request does not contain realtime delegation
3. second Responses request does contain the realtime delegation

## Validation
- `just fmt`
- `just test -p codex-core --test all
suite::realtime_conversation::inbound_handoff_request_steers_active_turn`

## Recent CI failures with the same signature

-
https://github.com/openai/codex/actions/runs/27538033492/job/81392362858
  - 2026-06-15, `[codex] update multi-agent v2 prompts`
- same test failed after `conversation.handoff.requested`; websocket
read failed with `os error 10053`

-
https://github.com/openai/codex/actions/runs/27543877820/job/81412200651
- 2026-06-15, `feat: dispatch queued user messages through core idle
extensions`
  - same test failed; websocket read failed with `os error 10054`

-
https://github.com/openai/codex/actions/runs/27544342375/job/81413801641
  - 2026-06-15, `[codex] Make marketplace loading capability aware`
  - same test failed; websocket read failed with `os error 10053`
2026-06-15 23:50:11 +01:00
Matthew ZengandGitHub 1fbaac1e50 [codex] Reuse Apps policy evaluation across MCP tool exposure (#27813)
## Summary

- move `AppToolPolicyEvaluator` and the Apps config/requirements policy
logic from `codex-core` into `codex-connectors`
- resolve one immutable policy snapshot per exposure build and reuse it
across every Codex Apps MCP tool
- keep core as a thin adapter from MCP metadata to connector-owned
policy input while preserving the call-time defense-in-depth check

## Why

`build_mcp_tool_exposure` evaluates every Codex Apps tool on each
sampling request. The old path rebuilt effective Apps configuration for
every tool, and the policy implementation lived in the already-large
core crate even though it is connector-specific.

The connector-owned evaluator keeps the expensive config merge/decode
out of the loop and gives core only the effective policy result it
needs.

## Performance

With the real 557-tool Apps corpus, `build_mcp_tool_exposure` measured
3.74 ms and 3.33 ms after the extraction (3.54 ms mean). The original
path measured 807 ms mean, so the final result retains the 99.6%
reduction.

## Validation

- `cargo check -p codex-connectors -p codex-core`
- `just test -p codex-connectors` — 15 passed
- `just test -p codex-core --lib connectors` — 35 passed
- `just test -p codex-core --lib mcp_tool_exposure` — 5 passed
- `just test -p codex-core --lib mcp_tool_call` — 72 passed
- `just bazel-lock-update`
- `just bazel-lock-check`
- `just fix -p codex-connectors`
- `just fix -p codex-core`
- `just fmt`
2026-06-15 15:24:33 -07:00
AbhinavandGitHub d7f298fe20 Respect blocking PostToolUse hooks in code mode (#28365)
## Summary

Make blocking hook behavior reliable for tools invoked from code mode.

Previously, a `PostToolUse` hook could block a completed tool result,
but code mode would still return the original typed result to
JavaScript. The hook appeared blocked in hook telemetry while the
running script continued with the result.

This change:

- rejects the nested JavaScript tool promise when `PostToolUse` blocks
- normalizes `decision: "block"` and exit code 2 to the same blocking
behavior
- surfaces the hook feedback as the rejected promise's error
- adds end-to-end coverage for the relevant PreToolUse and PostToolUse
interactions

## Hook semantics in code mode

| Hook behavior | Code-mode result |
|---|---|
| PreToolUse block | Reject the promise before the tool executes |
| PreToolUse `updatedInput` | Execute the rewritten invocation and
return its result |
| PostToolUse `decision: "block"` | Execute the tool, then reject the
promise with the hook reason |
| PostToolUse exit code 2 | Same behavior as `decision: "block"` |
| PostToolUse `continue: false` | Preserve the existing feedback-only
behavior; do not reject the promise |

## Test coverage

Added or strengthened end-to-end coverage proving that:

- a PreToolUse block rejects the JavaScript promise before execution
- a PreToolUse input rewrite executes only the rewritten command
- JavaScript receives the rewritten command's result
- PostToolUse `decision: "block"` rejects after the command executes
- PostToolUse exit code 2 has the same behavior
- the hook observes the original completed tool response
- the blocked original result does not reach JavaScript
- existing direct-mode replacement behavior remains intact
- `continue: false` without a reason produces deterministic fallback
feedback
2026-06-15 15:12:26 -07:00
Eric NingandGitHub 709f19e111 [codex] Add created-by-me remote plugin marketplace (#28203)
## Summary
- add the `created-by-me-remote` marketplace backed by paginated
`scope=USER` plugin directory and installed-plugin requests
- include USER plugins in installed-plugin caching, bundle sync, and
stale-cache cleanup without client-side discoverability filtering
- expose the marketplace through app-server v2 and regenerate the
protocol schemas

## Testing
- `cargo build -p codex-app-server --bin codex-app-server`
- production-auth `plugin/list` smoke test for `created-by-me-remote`
(returned the expected USER plugin as installed and enabled)
- `just test -p codex-core-plugins` (221 passed)
- `just test -p codex-app-server-protocol` (231 passed)
- `just test -p codex-app-server suite::v2::plugin_list::` (37 passed)
- `just fix -p codex-core-plugins -p codex-app-server-protocol -p
codex-app-server`
- `just fmt`
2026-06-15 22:07:07 +00:00
Owen LinandGitHub 040dafa32d feat(core): add metadata field to ResponseItem (#28355)
## Description

This PR adds an optional `metadata` field to `ResponseItem` for
Responses API calls. Only mechanical plumbing, no actual values
populated and sent yet. Turns out just adding a new field to
`ResponseItem` has quite a large blast radius already.

This change is backwards compatible because `metadata` is optional and
omitted when absent, so existing response items and rollout history
without it still deserialize and requests that do not set it keep the
same wire shape. For provider compatibility, we strip out `metadata`
before non-OpenAI Responses requests so Azure and AWS Bedrock never see
this field.

My followup PR here will actually make use of it to start storing and
passing along `turn_id`: https://github.com/openai/codex/pull/28360

## What changed

- Added `ResponseItemMetadata` with optional `turn_id`, plus optional
`metadata` on Responses API item variants and inter-agent communication.
- Preserved item metadata through response-item rewrites such as
truncation, missing tool-output synthesis, compaction history
rebuilding, visible-history conversion, rollout/resume, and generated
app-server schemas/types.
- Strip item metadata from non-OpenAI Responses requests while
preserving it for OpenAI-shaped requests.
- Updated the mechanical fixture/test construction churn required by the
new optional field.
2026-06-15 15:05:28 -07:00
mchen-oaiandGitHub af99f6a72f core: cache the tool search handler per session (#27258)
## Why

Tool router construction rebuilds the deferred-tool BM25 index during
session initialization and before each sampling continuation, even when
the searchable tool metadata is unchanged. Local profiling measured
`append_tool_search_executor` at roughly 113 ms per continuation, making
repeated index construction the largest measured router-building cost.

## What changed

- Add a session-scoped `ToolSearchHandlerCache` so continuations and
user turns can reuse the existing handler.
- Key reuse on the complete ordered `Vec<ToolSearchInfo>`, rebuilding
when searchable text, loadable tool specs, source metadata, or ordering
changes.
- Build handlers outside the cache lock and recheck before publishing
them, avoiding holding the mutex during index construction.

## Verification

- `cache_reuses_identical_search_infos_and_rebuilds_changed_inputs`
covers exact cache reuse and invalidation when the ordered search
metadata changes.
- Local rollout profiling showed the initial router build populating the
cache and unchanged later continuations reusing it:
  - uncached: 118 ms median across 14 spans from 3 rollouts
  - cached: 4 ms median across 12 spans from 3 rollouts
2026-06-15 14:48:30 -07:00
iceweasel-oaiandGitHub fbbe7706d6 Add hidden Windows sandbox wrapper entrypoint (#28358)
## Why

This is the second PR in the Windows fs-helper sandbox stack. The
fs-helper path needs a Windows sandbox launcher that has the same
argv-shaped contract as macOS `sandbox-exec` and `codex-linux-sandbox`,
but this PR only introduces that hidden launcher. It does not route
fs-helper through it yet.

The hidden launcher still needs to be policy-complete before later
direct-spawn callers use it. In particular, it has to carry the same
Windows sandbox policy details that the existing spawn paths already
understand: proxy enforcement, read/write root overrides, and
deny-read/deny-write overrides.

## What Changed

- Added the hidden `codex.exe --run-as-windows-sandbox` arg1 dispatch
path.
- Added `windows-sandbox-rs/src/wrapper.rs`, which parses the wrapper
argv, launches the requested command through the shared Windows sandbox
session runner from PR1, and forwards stdio.
- Added `create_windows_sandbox_command_args_for_permission_profile()`
so later direct-spawn callers can build the wrapper argv consistently.
- Made the wrapper argv round-trip the full Windows sandbox policy
surface it needs later: workspace roots, environment, permission
profile, sandbox level, private desktop, proxy enforcement, read/write
root overrides, and deny-read/deny-write overrides.
- Carried `proxy_enforced` through the shared Windows session request so
proxy-managed executions continue to use the offline/elevated sandbox
identity.
- Added wrapper argument round-trip coverage for the full policy fields.

## Verification

- `just test -p codex-windows-sandbox windows_wrapper_args_round_trip`
- `just test -p codex-arg0`
- `just test -p codex-core exec::tests::windows_`
- `just fix -p codex-windows-sandbox -p codex-core -p codex-cli`

Local note: the full `just fmt` command still fails on this workstation
in non-Rust formatter setup (`uv` cache access denied and missing
`dotslash`/buildifier), but the Rust formatter phase completed.
2026-06-15 21:30:32 +00:00
iceweasel-oaiandGitHub e7a9988d1a Add Windows unified exec yield floor (#27086)
## Why

The Windows `unified_exec` experiment regressed at the turn level in a
way that points to premature backgrounding / extra command cycles rather
than individual responses getting heavier:

- `codex_local_tool_calls_per_turn` was up about 20.7%.
- `codex_local_blended_tokens_per_turn` was up about 4.1%, and
`codex_local_output_tokens_per_turn` was up about 4.0%.
- `codex_local_response_latency_per_turn` was up about 8.3%.
- The primary activity metrics also moved down: `codex_turns` about
-6.6%, `codex_dau` about -1.0%, and `codex_local_hourly_active_users`
about -3.0%.

At the same time, the per-response metrics moved in the other direction:
blended tokens per response, output tokens per response, and latency per
response were all lower in test. That suggests the bad turn-level shape
is largely about extra tool/model cycles, not each response being slower
or more expensive on its own.

Local Windows benchmarking showed the likely mechanism: shell-wrapped
commands pay a large PowerShell startup/teardown tax before the actual
command has much time to run. In the benchmark, the PowerShell wrapper
added roughly 0.7-1.0s versus direct exec:

- Windows PowerShell: about 740ms p50 / 800ms p90 overhead versus direct
exec.
- PowerShell 7 (`pwsh`): about 930ms p50 / 980ms p90 overhead versus
direct exec.

The model commonly asks for a 1s initial yield. On Windows, that can
spend nearly the whole window waiting on PowerShell machinery, so
otherwise-short commands are more likely to return as background
sessions and require follow-up polling/tool calls.

This is intentionally a temporary unlock. It gives Windows closer to the
same useful post-shell command window as other platforms while we work
on reducing the PowerShell tax directly, for example with persistent
PowerShell workers or conservative direct-exec paths for commands that
do not need shell semantics.

## What changed

- Adds a Windows-only 2s floor to `unified_exec`'s initial
`yield_time_ms` clamp.
- Keeps larger model-requested waits unchanged, including the existing
10s default.
- Keeps the existing 30s max clamp.
- Leaves non-Windows behavior unchanged.
- Adds platform-gated tests for both the Windows floor and the
non-Windows clamp behavior.

## Verification

- `just test -p codex-core unified_exec`
2026-06-15 13:56:18 -07:00