Skip to content

Retry blocked Garmin calls through a browser-like client and stop one failing weigh-in from blocking the rest; 1.15.0 - #75

Merged
sturimcode merged 10 commits into
mainfrom
release/1.15.0
Oct 2, 2026
Merged

sturimcode merged 10 commits into
mainfrom
release/1.15.0

Conversation

@sturimcode

Copy link
Copy Markdown
Owner

Garmin resilience and upload retry tracking.

  • Some networks get a Cloudflare 403 on every Garmin API call after a successful login (python-garminconnect #444). A blocked call is now retried once through curl_cffi with browser impersonation and the existing token, before any relogin. Once that path works it's used for the rest of the run. The FIT upload goes through as multipart.
  • Upload results are classified: 409 counts as already uploaded only when a timed lookup finds the weigh-in on Garmin, 429 stops Garmin for the run, other 4xx are permanent, and 408, 5xx and network errors are retried. A POST is replayed only after an affirmative block, never after a gateway error.
  • A Garmin weigh-in that keeps failing no longer holds back newer ones. After 14 days or 84 attempts it stops blocking but stays queued and keeps retrying. It is given up only after 30 days of failures with newer weigh-ins uploading in that time, so an outage never drops a measurement. --status shows uploads waiting to retry.

764 tests pass locally.

…upload errors

The library sends data calls through plain requests, which Cloudflare refuses
with a 403 on some networks (upstream issue #444). A refused call now runs once
more through curl_cffi with a browser fingerprint and the same token before
the 403 counts against the session. FIT uploads go as a CurlMime multipart
body, since curl_cffi rejects files=. A Cloudflare block page that survives
the retry no longer triggers a relogin.

Upload outcomes now follow scalebridge-sync's classes: a 409 counts as
already uploaded, a 429 becomes a rate limit sync will not retry this run,
other 4xx and a refusal that outlives the relogin are permanent, and 408/5xx
stay transient for _retry.
…st of the run

A 409 on the body comp upload now counts as uploaded only when the day's
weigh-ins hold an entry within the weight window at the instant we sent,
the same match delete_weight_entry trusts. An unconfirmed 409, or a failed
lookup, raises PermanentSyncError naming the 409.

Once a browser-fingerprint retry succeeds, the client sends every later call
straight through curl_cffi instead of paying a refused plain request first.
A new GarminClient starts over on plain requests.
…h-ins

The per-target fetch cursor already re-fetches a failed upload on the next
run, so this adds no second replay path. A new upload_retries table in the
state DB records each retryable failure per user, target, and measurement.
After 84 failures (two weeks of 4-hour runs) or once the measurement is more
than 14 days old, the entry is given up with a log line, so one measurement
that never uploads stops holding back every newer one. Strava and Zwift keep
only their newest failed weight. Permanent, unsupported, and auth failures
and saved Garmin replacements are never queued, --dry-run leaves the table
untouched, and --status shows how many uploads are waiting to retry.
Hitting the age or attempt cap no longer gives an entry up on its own. A
capped measurement is still tried in order; if it fails, the run moves past
it instead of stopping the target, and marks it given up only after a newer
measurement for that target uploads. If the newer ones fail too, the target
is down: the entry stays pending and is delivered when service returns.
--dry-run shows which uploads would be retried under that rule.
A 409 now counts as uploaded only when Garmin holds an entry within 0.1 kg
whose own timestamp parses and sits within 120 seconds of the upload. A
same-weight entry with no usable timestamp no longer confirms it. The
weight-only fallback stays for delete_weight_entry, where a lone untimed
match replaces our own weight-only upload and is documented as such.

The confirming lookup runs as its own call through the same fallback,
sticky curl_cffi path, and single relogin as other reads, so its recovery
never resends the upload. A lookup that fails with a 5xx or network error
stays retryable, a 429 ends Garmin for the run like an upload 429, and a
refusal after the relogin gets the upload's reauth advice. Only a lookup
that answers without the weigh-in is permanent.

Cloudflare block detection now needs affirmative evidence: cf-mitigated, or
a challenge page on a 403, 429, or 503, or the firewall block page on a 403.
The generic "Cloudflare Ray ID" footer no longer counts, and no 5xx gateway
page is a block. The fingerprint replay of a call uses the recorded status
over the error text, so an upload is replayed only after a 403 or a
challenge, never after a 5xx or a timeout.

Tests can no longer reach a real Garmin login: conftest makes
Garmin.login raise unless a test patches the relogin.
A capped measurement that fails during a run was given up as soon as a
newer one uploaded, even when Garmin had only just come back. Now the newer
upload earns each capped failure from that run one more try through _retry
and the run's single relogin. It is given up only if that try fails again
with a retryable error. A permanent failure on the extra try (a rate limit,
a refused session) leaves the entry queued and ends Garmin for the run, and
the next run's fetch reaches back to the oldest queued Garmin failure even
when the cursor has moved past it.

The age cap now counts from an entry's first recorded failure instead of
the measurement's timestamp, so an old weigh-in reached by a backfill gets
the full two weeks before it stops holding back newer ones. The attempt cap
is unchanged.
A 409 whose confirming lookup failed with a 5xx or a network error used to
escape upload_body_composition, and sync's _retry then called the whole
method again, sending the upload a second and third time just to repeat a
read. The lookup is now retried on its own inside the client, up to three
times with a short pause, through the same fallback and single relogin. If
it still fails, the client raises RetryNextRunError: retryable, so the
measurement goes to the retry queue, but _retry passes it straight through,
so the run sends exactly one POST. The next run's POST 409s again and repeats
the lookup.
A capped upload (84 attempts, or two weeks since its first failure) no
longer blocks newer measurements and is no longer given up after one more
failure in a run where a newer one landed. It is tried once each run, in
order; a retryable failure moves past it, it stays queued, and its attempts
keep counting. A transient failure never gives it up on its own.

An entry is given up only when its first failure is more than 30 days old
and a newer measurement for the same user and target has uploaded since
then, which the queue now tracks in a last_newer_success_at column (added to
existing databases as NULL, which can only delay a give-up). During a total
outage nothing newer lands, so nothing is given up. Unsupported measurements
are still skipped at once.

The fetch still reaches back to the oldest queued Garmin failure, now no
further than 32 days. Older entries that cannot be given up stay queued, a
run logs once that they are waiting, and --backfill-days still reaches them.
A capped failure is not reported as a target error once the target has
taken a newer weigh-in. The extra in-run retry after a newer upload is gone.
@sturimcode
sturimcode merged commit 9c11c6f into main Oct 2, 2026
11 checks passed
@sturimcode
sturimcode deleted the release/1.15.0 branch October 2, 2026 15:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant