Repository navigation
feishu: card rollover never fires — long tasks silently stop updating #814
Description
Activity
Update: fixed locally in my fork, using button pagination instead of multi-card rollover.
Approach
- The card now renders
◀ prev / N/M / next ▶at the bottom; each button'svaluecarries{ga_card, page}. - Registered via
register_p2_card_action_trigger; the callback resolves the card instance from a uid registry and re-patches the same message, so the whole step history stays browsable in one card. - Pagination uses dual criteria: max 50 steps and 8000 chars per page. Measured from real logs: the original overflow fired around step 58, with single-step detail averaging 272 chars (max 9003), so a step-count-only limit is not enough.
Both underlying defects are also fixed
_rollover_locked()now clearsself.stepsand advancesturn_base, so a rollover page is actually smaller and step numbering stays contiguous._push_sync()returnslimit=Trueon the create path when it fails, so the caller can roll over and retry instead of leavingmsg_id = Noneforever.
Result: no more silent card failures, and long tasks no longer need a manual "continue".
Happy to open a PR if this fits your design — the alternatives (multi-message rollover, or a rolling window that drops old steps) both hurt reviewability.
- The card now renders
I'd like to take this one, scoped to the two defects and the minimal fix suggested in the issue body:
_rollover_locked()clearsself.stepsand carriesturn_baseso step numbering stays contiguous and the new card is actually smaller;_push_sync()returnslimit=Truewhen the create path fails, so_worker_looprolls over and retries instead of the card becoming permanently unsendable.
@cuipengcx90 you mentioned a local fork fix using button pagination — I'm deliberately not attempting the pagination rework here, keeping this PR to the rollover defects so the two approaches can coexist (yours is the nicer long-term UX; this unblocks the silent-stop failure mode with a small diff). If you'd rather own the whole thing, say the word and I'll drop this.
PR up: #820
Important finding while implementing: the code the issue references (
_rollover_locked()/_push_sync()/_worker_loop,_DETAIL_LIMIT = 4000) does not exist on current main. A rollover implementation was merged in #209 (bf002e5b) but was dropped in the later Feishu interface rework (df525560) — main today has no rollover at all, and the original freeze symptom reproduces structurally.So this PR re-lands the rollover behavior adapted to the current
_TaskCard:_push()returns(ok, limit); the create path reportslimit=Trueon failure (your defect 2 — withmsg_id is Noneevery push took the create path and never rolled over)_rollover()clearsstepsso the continuation card is actually smaller (your defect 1)turn_base+ monotonicturn_nokeepTurn Ncontiguous across cards- proactive flip at 50 steps — just under your measured ~58-step overflow point
@cuipengcx90 as agreed, scoped to the rollover defects only; your button-pagination fix is complementary and untouched here.
11 new tests (ast/exec pattern); full suite 298 passed, same 3 pre-existing Windows-env failures as before my change.
Symptom
Long Feishu tasks stop updating their card around step ~50 while the agent keeps running. The user has to send "继续" to get a working card back, and any long task reproduces it (198 occurrences of the card failure in a single session log).
Root cause (two defects)
_rollover_locked()bumpspage_noand clearsmsg_idbut notself.steps._build_locked()iterates the fullself.steps, so the "new" card is exactly as large as the overflowing one and fails the same way — rollover is a no-op._push_sync()returns(False, False)when the create path fails, while_worker_looprolls over only onlimit=True. Oncemsg_idisNone, every later push takes the create path, returnslimit=False, and rollover never runs again — the card becomes permanently unsendable.Reproduction
Any session where accumulated steps exceed the card size limit (roughly 50 steps at
_DETAIL_LIMIT = 4000). The log then repeats:while
上一张工作卡片达到飞书限制(the rollover note) never appears — direct proof that rollover never fired.Suggested fix
self.stepsin_rollover_locked()and carryturn_baseso step numbering stays contiguous.stepsreaches a threshold, instead of waiting for the API to reject the card.limit=Truefrom the create path on failure so the caller can roll over and retry.