mirror of
https://github.com/pchuan98/codex.git
synced 2026-07-01 00:31:56 +08:00
Improve goal continuation based on feedback (#22045)
## Summary This PR updates the goal continuation prompt to address feedback from early adopters. There are two primary changes: 1. Goal continuation and budget-limit steering prompts now use hidden user-context messages instead of hidden developer messages. 2. The goal continuation prompt is refined to improve the model's ability to fully complete the active goal rather than stop at a smaller or merely passing subset. The user-message transition is important for two reasons. First, it eliminates an issue where older steering messages could be responded to again after a new turn. Second, it works better with compaction because user messages are treated differently from developer messages during compaction. The prompt refinements make persistence explicit, ground work in current evidence, encourage `update_plan` for multi-step progress visibility, and require stronger completion audits before calling `update_goal`. It also removes the elapsed-time reporting in the prompt; I saw evidence that this was causing the model to shortcut work as it became nervous about time. These changes were tested with evals. Chriss4123 has also been running independent evals in [#19910](https://github.com/openai/codex/issues/19910), and many of the improvements in this PR were suggested by him. ## Verification - Tested with evals. - Added and updated focused `codex-core` coverage for hidden goal user context, continuation and budget-limit request shape, prompt rendering, and objective delimiter escaping.
This commit is contained in:
@@ -7707,7 +7707,7 @@ async fn active_goal_continuation_runs_again_after_no_tool_turn() -> anyhow::Res
|
||||
.expect("goal mode should be enableable in tests");
|
||||
});
|
||||
let test = builder.build(&server).await?;
|
||||
let _responses = mount_sse_sequence(
|
||||
let responses = mount_sse_sequence(
|
||||
&server,
|
||||
vec![
|
||||
sse(vec![
|
||||
@@ -7770,6 +7770,25 @@ async fn active_goal_continuation_runs_again_after_no_tool_turn() -> anyhow::Res
|
||||
})
|
||||
.await??;
|
||||
|
||||
let continuation_request = responses
|
||||
.requests()
|
||||
.into_iter()
|
||||
.find(|request| request.body_contains_text("<goal_context>"))
|
||||
.expect("expected a goal continuation request");
|
||||
let body = continuation_request.body_json();
|
||||
let goal_context_message = body["input"]
|
||||
.as_array()
|
||||
.expect("input should be an array")
|
||||
.iter()
|
||||
.find(|item| item.to_string().contains("<goal_context>"))
|
||||
.expect("goal context message should be present");
|
||||
assert_eq!(goal_context_message["role"].as_str(), Some("user"));
|
||||
assert!(
|
||||
goal_context_message
|
||||
.to_string()
|
||||
.contains("Continue working toward the active thread goal.")
|
||||
);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
@@ -7970,10 +7989,12 @@ async fn budget_limited_accounting_steers_active_turn_without_aborting() -> anyh
|
||||
let [ResponseInputItem::Message { role, content, .. }] = pending_input.as_slice() else {
|
||||
panic!("expected one budget-limit steering message, got {pending_input:#?}");
|
||||
};
|
||||
assert_eq!("developer", role);
|
||||
assert_eq!("user", role);
|
||||
let [ContentItem::InputText { text }] = content.as_slice() else {
|
||||
panic!("expected one text span in budget-limit steering message, got {content:#?}");
|
||||
};
|
||||
assert!(text.starts_with("<goal_context>"));
|
||||
assert!(text.trim_end().ends_with("</goal_context>"));
|
||||
assert!(text.contains("budget_limited"));
|
||||
assert!(text.to_lowercase().contains("wrap up this turn soon"));
|
||||
assert!(sess.active_turn.lock().await.is_some());
|
||||
|
||||
Reference in New Issue
Block a user