fix(ci): let worker-live ride out a rolling deploy instead of failing it - #102
Conversation
The v0.4.7 deploy failed worker-live although it was correct: the manifest already read v0.4.7, and the next GET /install.sh reached an edge still serving v0.4.6, so the content check failed at once. Workers deploys roll out gradually; the re-run passed. Content checks now run as one pass retried within the existing deadline, keyed on each response's x-resq-commit. A response from another commit is rollout lag and retries; one from main's commit with wrong bytes, or with no x-resq-commit, still fails immediately. HEAD and GET of install.sh must both come from main's deploy before their lengths are compared.
Deploying with
|
| Status | Name | Latest Commit | Updated (UTC) |
|---|---|---|---|
| ✅ Deployment successful! View logs |
get-resq-software | 1aebfeb | Oct 03 2026, 11:11 PM |
|
Warning Review limit reachedYou've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Next included review available in 37 minutes. View limit detailsLimit details: You’ve used the included review currently available. Review configuration: ⚙️ Run configuration
📒 Files selected for processing (1)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
This pull request significantly improves the deployment verification workflow for the live worker. The previous implementation was susceptible to false negatives during Cloudflare Workers gradual rollouts. This change introduces a more robust polling and verification mechanism that accounts for this "eventual consistency" by checking commit hashes and retrying on mismatches until a deadline is reached. After a thorough review, I found no security vulnerabilities, logic bugs, or performance issues. The changes are well-thought-out and improve the reliability of the deployment pipeline. Therefore, I approve of these changes.
|
Summary
worker-live failed on the v0.4.7 pin bump (#101) although the deploy was correct:
On attempt 6 the manifest flipped to v0.4.7, so the poll loop exited. The very next
GET /install.shreached an edge still running v0.4.6, and the content check failed at once. Workers deploys roll out gradually; Workers Builds finished at 23:04:12 and the check gave up seconds later. A re-run passed:serves v0.4.7 @ 10a74187, 7 artifacts verified.Change
The content checks (HEAD/GET length, per-route digests, root route) now run as one pass, retried within the existing deadline. Each response's
x-resq-commitdecides what a mismatch means:x-resq-commitHEAD and GET of
/install.shmust also both come from main's deploy before their lengths are compared, because the same race could otherwise compare two versions' lengths. The manifest poll and the deadline itself are unchanged.Test plan
The workflow's own script was extracted (ORIGIN and timers overridden) and run against a local mock serving real Worker responses: main's routes for "new", and the Worker's version-locked
/v0.4.6/...routes for "old", so the old responses carry v0.4.6's real commit and bytes.x-resq-commit