Repository navigation
Balance Windows Bazel test shards using duration estimates - #49103
Conversation
## Why Hash-based sharding does not account for differences in test duration. Use estimated test time to balance the Windows Bazel workload across shards. ## What changed - Add checked-in duration weights and a deterministic selector that assigns the longest targets first to the least-loaded shard. - Assign every queried target exactly once, default missing weights to one second, and ignore stale weights. - Run queries and selection in the foreground so failures stop the job before tests run with a partial target list. - Emit LF-delimited target labels for Git Bash and report estimated shard loads on stderr. ## Testing Add seven passing selector tests covering deterministic coverage and balancing, missing and stale weights, invalid inputs, duration-file overrides, and CLI output. Run them on the first Windows CI shard. GitOrigin-RevId: d415ae093bfe57794ee53a2eb9d712022202dad8
1606e5e to
7e049b3
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1606e5ef57
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| # Windows Bazel duration weights; columns: target<TAB>seconds. | ||
| # Seeded from 24 successful public openai/codex Windows job logs across six main | ||
| # runs, September 21-23, 2026. Public GitHub Actions run IDs: | ||
| # 35638529858, 35665540272, 35694732932, 35912411800, 35919297566, 35924169500. | ||
| # To refresh, take the median summed inner-run time (N * average) from clean, |
There was a problem hiding this comment.
Split the Windows sharding change into reviewable stages
This 544-line non-mechanical CI behavior change exceeds the repository's under-500-line guidance for complex logic changes, making the target-selection algorithm, its workflow integration, and the empirical weights difficult to validate independently. Land the deterministic selector and its tests separately from the workflow wiring and checked-in timing data; the latter is a coherent second stage because this table is solely consumed by the new selector.
AGENTS.md reference: AGENTS.md:L125-L131
Useful? React with 👍 / 👎.
Balance Windows Bazel test shards using duration estimates
Why
Hash-based sharding does not account for differences in test duration. Use estimated test time to balance the Windows Bazel workload across shards.
What changed
Testing
Add seven passing selector tests covering deterministic coverage and balancing, missing and stale weights, invalid inputs, duration-file overrides, and CLI output. Run them on the first Windows CI shard.