Skip to content

fix(ai): stop giving up on a model that is still working - #75

Open
arthware-dev wants to merge 2 commits into
fix/llm-timeout-headroomfrom
fix/llm-liveness-timeouts
Open

fix(ai): stop giving up on a model that is still working#75
arthware-dev wants to merge 2 commits into
fix/llm-timeout-headroomfrom
fix/llm-liveness-timeouts

Conversation

@arthware-dev

Copy link
Copy Markdown
Contributor

Stacked on #74, which raises the ceilings. This changes what the timeout
measures, so the ceilings stop mattering.

The client waited for a whole response with nothing on the wire, so the
only question it could ask was whether too much time had passed — and a
long prefill looks exactly like a stuck server from there. It now streams
and times the silence instead: while an answer keeps arriving there is no
deadline, and a server that actually stops is caught in 90s rather than
900s.

Merge #74 first.

The AI client waited for a whole response with nothing on the wire, so
the only question it could ask was whether too much time had passed. A
long document's prefill looks exactly like a stuck server from there, so
filing a big scan on a small Mac got cancelled mid-flight and every
minute already spent on it was thrown away.

It now streams and times the silence instead. While an answer keeps
arriving there is no deadline at all, however slow the machine. A server
that actually stops is caught in 90 seconds rather than 300.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant