[codex] Add use_responses_lite 'override' logic (#26487)

## Summary

- add a defaulted `ModelInfo.use_responses_lite` catalog field
- support serializing `reasoning.context` while preserving the existing
effort and summary path
- has not been turned on for any models yet

I've added an override to parallel tools if responses_lite is on. I've
also forced persistent reasoning when using responses_lite. It would be
ideal if we could centralize all the responses_lite plumbing, but I
think this is best for now to keep the plumbing & diffs small.

## Testing

- `cargo test -p codex-protocol
model_info_defaults_availability_nux_to_none_when_omitted`
- `RUST_MIN_STACK=8388608 cargo test -p codex-core
responses_lite_sets_all_turns_context_and_disables_parallel_tool_calls`
- `RUST_MIN_STACK=8388608 cargo test -p codex-core
configured_reasoning_summary_is_sent`
- `cargo check -p codex-core --tests`
- `RUST_MIN_STACK=8388608 cargo clippy -p codex-core --tests` (passes
with pre-existing warnings in `codex-code-mode` and
`codex-core-plugins`)
This commit is contained in:
rka-oai
2026-06-04 18:49:51 -07:00
committed by GitHub
parent ecae412740
commit e0096db6dc
17 changed files with 104 additions and 1 deletions
+8 -1
View File
@@ -44,6 +44,7 @@ use codex_api::RawMemory as ApiRawMemory;
use codex_api::RealtimeCallClient as ApiRealtimeCallClient;
use codex_api::RealtimeSessionConfig as ApiRealtimeSessionConfig;
use codex_api::Reasoning;
use codex_api::ReasoningContext;
use codex_api::RequestTelemetry;
use codex_api::ReqwestTransport;
use codex_api::ResponseCreateWsRequest;
@@ -609,6 +610,7 @@ impl ModelClient {
reasoning: effort.map(|effort| Reasoning {
effort: Some(effort),
summary: None,
context: None,
}),
};
@@ -727,6 +729,11 @@ impl ModelClient {
} else {
Some(summary)
},
// When Responses Lite is disabled, omit context so Responses uses the default,
// which is currently `current_turn`.
context: model_info
.use_responses_lite
.then_some(ReasoningContext::AllTurns),
})
} else {
None
@@ -775,7 +782,7 @@ impl ModelClient {
input,
tools,
tool_choice: "auto".to_string(),
parallel_tool_calls: prompt.parallel_tool_calls,
parallel_tool_calls: prompt.parallel_tool_calls && !model_info.use_responses_lite,
reasoning,
store: provider.is_azure_responses_endpoint(),
stream: true,