You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
With #215 and #218, LanguageModelSamplingMode defines seven presets along the predictability and creativity spectrum ("most-predictable" through "most-creative"), defaulting to "balanced" when omitted.
In practice, different sampling modes can have significant implementation-specific performance differences. For example, greedy sampling ("most-predictable", where topK = 1) unlocks decoding accelerations such as Multi-Token Prediction (MTP) speculative decoding that are either unavailable or less effective under higher-temperature stochastic sampling.
Currently, web developers who want the lowest latency and highest token throughput must know to select "most-predictable" as a proxy for speed. However:
"most-predictable" expresses output determinism rather than latency intent, and it may not always be the fastest mode across all user agents or model architectures.
An explicit "fastest" mode would allow implementations to select the optimal sampling preset (such as "most-predictable") while potentially applying additional runtime or decoding optimizations tailored for speed.
Proposal
Explore adding "fastest" to LanguageModelSamplingMode:
Resolution of session.samplingMode: When a session is created with { samplingMode: "fastest" }, should session.samplingMode reflect the resolved preset selected by the implementation (for example, "most-predictable"), or should it remain "fastest" (in case "fastest" enables additional decoding optimizations beyond just topK and temperature)?
Default behavior when samplingMode is omitted: Per Update spec algorithms to include samplingMode #218, omitting samplingMode defaults to "balanced". Keeping "balanced" as the default and requiring explicit { samplingMode: "fastest" } preserves predictable general-purpose quality while giving latency-sensitive applications a dedicated opt-in.
Relationship to preference: "speed" in Writing Assistance APIs: The Writing Assistance APIs (Summarizer, Writer, Rewriter) use a preference option ("speed" vs. "capability"). Should "fastest" live on LanguageModelSamplingMode, or should we consider how it relates to preference?
Background & Motivation
With #215 and #218,
LanguageModelSamplingModedefines seven presets along the predictability and creativity spectrum ("most-predictable"through"most-creative"), defaulting to"balanced"when omitted.In practice, different sampling modes can have significant implementation-specific performance differences. For example, greedy sampling (
"most-predictable", wheretopK = 1) unlocks decoding accelerations such as Multi-Token Prediction (MTP) speculative decoding that are either unavailable or less effective under higher-temperature stochastic sampling.Currently, web developers who want the lowest latency and highest token throughput must know to select
"most-predictable"as a proxy for speed. However:"most-predictable"expresses output determinism rather than latency intent, and it may not always be the fastest mode across all user agents or model architectures."fastest"mode would allow implementations to select the optimal sampling preset (such as"most-predictable") while potentially applying additional runtime or decoding optimizations tailored for speed.Proposal
Explore adding
"fastest"toLanguageModelSamplingMode:Open Questions for Discussion
session.samplingMode: When a session is created with{ samplingMode: "fastest" }, shouldsession.samplingModereflect the resolved preset selected by the implementation (for example,"most-predictable"), or should it remain"fastest"(in case"fastest"enables additional decoding optimizations beyond justtopKandtemperature)?samplingModeis omitted: Per Update spec algorithms to include samplingMode #218, omittingsamplingModedefaults to"balanced". Keeping"balanced"as the default and requiring explicit{ samplingMode: "fastest" }preserves predictable general-purpose quality while giving latency-sensitive applications a dedicated opt-in.preference: "speed"in Writing Assistance APIs: The Writing Assistance APIs (Summarizer,Writer,Rewriter) use apreferenceoption ("speed"vs."capability"). Should"fastest"live onLanguageModelSamplingMode, or should we consider how it relates topreference?cc @michaelwasserman @reillyeon @tomayac @KenjiBaheux