Skip to content

Balance Windows Bazel test shards using duration estimates - #49103

Merged
copyberry[bot] merged 1 commit into
mainfrom
copyberry/codex-internal-to-codex-oss/d415ae093bfe57794ee53a2eb9d712022202dad8
Sep 28, 2026
Merged

copyberry[bot] merged 1 commit into
mainfrom
copyberry/codex-internal-to-codex-oss/d415ae093bfe57794ee53a2eb9d712022202dad8

Conversation

@copyberry

@copyberry copyberry Bot commented Sep 28, 2026 •

Copy link
Copy Markdown

Balance Windows Bazel test shards using duration estimates

Why

Hash-based sharding does not account for differences in test duration. Use estimated test time to balance the Windows Bazel workload across shards.

What changed

  • Add checked-in duration weights and a deterministic selector that assigns the longest targets first to the least-loaded shard.
  • Assign every queried target exactly once, default missing weights to one second, and ignore stale weights.
  • Run queries and selection in the foreground so failures stop the job before tests run with a partial target list.
  • Emit LF-delimited target labels for Git Bash and report estimated shard loads on stderr.

Testing

Add seven passing selector tests covering deterministic coverage and balancing, missing and stale weights, invalid inputs, duration-file overrides, and CLI output. Run them on the first Windows CI shard.

## Why

Hash-based sharding does not account for differences in test duration. Use estimated test time to balance the Windows Bazel workload across shards.

## What changed

- Add checked-in duration weights and a deterministic selector that assigns the longest targets first to the least-loaded shard.
- Assign every queried target exactly once, default missing weights to one second, and ignore stale weights.
- Run queries and selection in the foreground so failures stop the job before tests run with a partial target list.
- Emit LF-delimited target labels for Git Bash and report estimated shard loads on stderr.

## Testing

Add seven passing selector tests covering deterministic coverage and balancing, missing and stale weights, invalid inputs, duration-file overrides, and CLI output. Run them on the first Windows CI shard.

GitOrigin-RevId: d415ae093bfe57794ee53a2eb9d712022202dad8
@copyberry
copyberry Bot force-pushed the copyberry/codex-internal-to-codex-oss/d415ae093bfe57794ee53a2eb9d712022202dad8 branch from 1606e5e to 7e049b3 Compare September 28, 2026 23:52
@copyberry
copyberry Bot merged commit 7e049b3 into main Sep 28, 2026
@copyberry
copyberry Bot deleted the copyberry/codex-internal-to-codex-oss/d415ae093bfe57794ee53a2eb9d712022202dad8 branch September 28, 2026 23:52

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 1606e5ef57

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +1 to +5
# Windows Bazel duration weights; columns: target<TAB>seconds.
# Seeded from 24 successful public openai/codex Windows job logs across six main
# runs, September 21-23, 2026. Public GitHub Actions run IDs:
# 35638529858, 35665540272, 35694732932, 35912411800, 35919297566, 35924169500.
# To refresh, take the median summed inner-run time (N * average) from clean,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Split the Windows sharding change into reviewable stages

This 544-line non-mechanical CI behavior change exceeds the repository's under-500-line guidance for complex logic changes, making the target-selection algorithm, its workflow integration, and the empirical weights difficult to validate independently. Land the deterministic selector and its tests separately from the workflow wiring and checked-in timing data; the latter is a coherent second stage because this table is solely consumed by the new selector.

AGENTS.md reference: AGENTS.md:L125-L131

Useful? React with 👍 / 👎.

@github-actions github-actions Bot locked and limited conversation to collaborators Sep 28, 2026

This branch was successfully deployed

1 active deployment
issue-triage — 7e049b3e Deployed Sep 29, 2026 by hoiheejin1 via Identify potential duplicates (all issues) #49576
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants