Skip to content

Consider adding a "fastest" LanguageModelSamplingMode value #219

Description

@isaacahouma

Background & Motivation

With #215 and #218, LanguageModelSamplingMode defines seven presets along the predictability and creativity spectrum ("most-predictable" through "most-creative"), defaulting to "balanced" when omitted.

In practice, different sampling modes can have significant implementation-specific performance differences. For example, greedy sampling ("most-predictable", where topK = 1) unlocks decoding accelerations such as Multi-Token Prediction (MTP) speculative decoding that are either unavailable or less effective under higher-temperature stochastic sampling.

Currently, web developers who want the lowest latency and highest token throughput must know to select "most-predictable" as a proxy for speed. However:

  1. "most-predictable" expresses output determinism rather than latency intent, and it may not always be the fastest mode across all user agents or model architectures.
  2. An explicit "fastest" mode would allow implementations to select the optimal sampling preset (such as "most-predictable") while potentially applying additional runtime or decoding optimizations tailored for speed.

Proposal

Explore adding "fastest" to LanguageModelSamplingMode:

enum LanguageModelSamplingMode {
  "fastest",
  "most-predictable",
  "predictable",
  "slightly-predictable",
  "balanced",
  "slightly-creative",
  "creative",
  "most-creative"
};

Open Questions for Discussion

  1. Resolution of session.samplingMode: When a session is created with { samplingMode: "fastest" }, should session.samplingMode reflect the resolved preset selected by the implementation (for example, "most-predictable"), or should it remain "fastest" (in case "fastest" enables additional decoding optimizations beyond just topK and temperature)?
  2. Default behavior when samplingMode is omitted: Per Update spec algorithms to include samplingMode #218, omitting samplingMode defaults to "balanced". Keeping "balanced" as the default and requiring explicit { samplingMode: "fastest" } preserves predictable general-purpose quality while giving latency-sensitive applications a dedicated opt-in.
  3. Relationship to preference: "speed" in Writing Assistance APIs: The Writing Assistance APIs (Summarizer, Writer, Rewriter) use a preference option ("speed" vs. "capability"). Should "fastest" live on LanguageModelSamplingMode, or should we consider how it relates to preference?

cc @michaelwasserman @reillyeon @tomayac @KenjiBaheux

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions