Hi SWE-bench Pro team, I noticed that GPT-5.2-Codex is reported on the SWE-bench Pro leaderboard with a resolution rate of 41.04%. However, unlike several other models, I could not find the corresponding trajectories, cost information, or execution configuration in the released artifacts:
https://github.com/scaleapi/SWE-bench_Pro-os/tree/main/traj
https://docent.transluce.org/dashboard/032fb63d-4992-4bfc-911d-3b7dafcb931f/agent_run
Would it be possible to share the GPT-5.2-Codex SWE-agent trajectories or, if unavailable, the associated evaluation configuration and cost statistics? Thanks!
Hi SWE-bench Pro team, I noticed that GPT-5.2-Codex is reported on the SWE-bench Pro leaderboard with a resolution rate of 41.04%. However, unlike several other models, I could not find the corresponding trajectories, cost information, or execution configuration in the released artifacts:
https://github.com/scaleapi/SWE-bench_Pro-os/tree/main/traj
https://docent.transluce.org/dashboard/032fb63d-4992-4bfc-911d-3b7dafcb931f/agent_run
Would it be possible to share the GPT-5.2-Codex SWE-agent trajectories or, if unavailable, the associated evaluation configuration and cost statistics? Thanks!