From 9e6354a1a29ec0f0eed3ca468b2a72c8893d1a43 Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 10:06:03 -0700 Subject: [PATCH 1/8] Add Thinking mode section to the Prompt API explainer --- README.md | 65 +++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 65 insertions(+) diff --git a/README.md b/README.md index b28f836..77b9598 100644 --- a/README.md +++ b/README.md @@ -519,6 +519,71 @@ Error-handling behavior (only applicable in contexts where legacy parameters are * If values above `maxTopK` are passed for `topK`, then `create()` will clamp to `maxTopK`. (This includes `+Infinity` and numbers above `Number.MAX_SAFE_INTEGER`.) * If fractional values are passed for `topK`, they are rounded down (using the usual [IntegerPart](https://webidl.spec.whatwg.org/#abstract-opdef-integerpart) algorithm for web specs). +### Thinking mode + +Modern language models increasingly support **reasoning** (or "thinking mode"), generating an intermediate scratchpad of reasoning tokens before producing a final response. This significantly improves accuracy on math, logic, code generation, and multi-step planning, at the cost of higher latency and compute. + +Developers can configure this reasoning process via the `thinking` option (`{ effort, includeThoughts }`), unlocking reasoning for complex prompts while dialing it back (or disabling it) on simple turns to save battery and latency. + +The allowed values for `effort` are: +* `"none"` (default): No intermediate reasoning tokens are generated before the response. +* `"low"`: A minimal reasoning budget for straightforward multi-step tasks. +* `"medium"`: A moderate reasoning budget balancing accuracy and latency. +* `"high"`: The maximum reasoning budget for complex math, logic, code generation, and planning tasks. + +Developers can check support via `LanguageModel.availability()`, set a baseline effort tier when creating a session, and override it on individual `prompt()` or `promptStreaming()` calls: + +```js +const status = await LanguageModel.availability({ + thinking: { effort: "high" } +}); + +if (status !== "unavailable") { + const session = await LanguageModel.create({ + thinking: { effort: "high" } + }); + + // Turn 1: Uses the session's default ("high" effort) for a complex task. + // By default (includeThoughts: false), prompt() returns only the final answer string. + const code = await session.prompt( + "Write a function to find the shortest path in a weighted directed graph." + ); + + // Turn 2: Override to "none" for a simple follow-up to save latency and battery. + const formatted = await session.prompt( + "Add JSDoc comments to that function.", + { thinking: { effort: "none" } } + ); +} +``` + +#### Viewing intermediate thoughts + +By default, `includeThoughts` is `false` so existing applications only receive the final text response. When set to `true`, `promptStreaming()` and `prompt()` emit structured `{ type, value }` dictionaries (following the same pattern as [Tool use](#tool-use)) so applications can render a collapsible `"Thinking..."` UI: + +```js +const session = await LanguageModel.create({ + thinking: { + effort: "medium", + includeThoughts: true + } +}); + +const stream = session.promptStreaming("Plan a 3-day itinerary for Tokyo."); + +for await (const chunk of stream) { + if (chunk.type === "thought") { + thinkingContainer.textContent += chunk.value; + } else if (chunk.type === "text") { + responseContainer.textContent += chunk.value; + } +} +``` + +Rather than exposing model-specific token counts, user agents map each `effort` tier to an appropriate reasoning budget for the underlying model. This budget acts as an upper bound: models stop thinking early once they reach a conclusion, or transition to the final answer if the ceiling is reached. Following standard reasoning-model behavior, intermediate thoughts are stripped from the conversation history after each turn completes and do not permanently accumulate in `session.contextUsage`. + +_Other possibilities under investigation include an `"auto"` effort level for dynamic prompt routing, as well as how `"thought"` blocks interleave with [Tool use](#tool-use) and [`samplingMode`](#configuration-of-sampling-modes)._ + ### Session persistence and cloning Each language model session consists of a persistent series of interactions with the model: From aea5ece1ae3f77ec627648418d2e206e1bfad9c2 Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 11:44:02 -0700 Subject: [PATCH 2/8] Update README.md Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 77b9598..b71421a 100644 --- a/README.md +++ b/README.md @@ -573,7 +573,7 @@ const stream = session.promptStreaming("Plan a 3-day itinerary for Tokyo."); for await (const chunk of stream) { if (chunk.type === "thought") { - thinkingContainer.textContent += chunk.value; + thinkingContainer.append(chunk.value); } else if (chunk.type === "text") { responseContainer.textContent += chunk.value; } From 8befd5633add02fb317f205b807665505abca89a Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 11:44:12 -0700 Subject: [PATCH 3/8] Update README.md Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index b71421a..2ecc74f 100644 --- a/README.md +++ b/README.md @@ -565,7 +565,7 @@ By default, `includeThoughts` is `false` so existing applications only receive t const session = await LanguageModel.create({ thinking: { effort: "medium", - includeThoughts: true + includeThoughts: true, } }); From 60b06b776a0d72cef4957efa0a6dbb4f6e653c33 Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 11:44:18 -0700 Subject: [PATCH 4/8] Update README.md Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 2ecc74f..e64d857 100644 --- a/README.md +++ b/README.md @@ -575,7 +575,7 @@ for await (const chunk of stream) { if (chunk.type === "thought") { thinkingContainer.append(chunk.value); } else if (chunk.type === "text") { - responseContainer.textContent += chunk.value; + responseContainer.append(chunk.value); } } ``` From e2e0193cdb236cfe2fb6ba82165317bc07f587e4 Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 12:02:01 -0700 Subject: [PATCH 5/8] Clarify that thought and text chunks may contain Markdown --- README.md | 1 + 1 file changed, 1 insertion(+) diff --git a/README.md b/README.md index e64d857..c7026fd 100644 --- a/README.md +++ b/README.md @@ -572,6 +572,7 @@ const session = await LanguageModel.create({ const stream = session.promptStreaming("Plan a 3-day itinerary for Tokyo."); for await (const chunk of stream) { + // Both "thought" and "text" chunks may contain Markdown; we append raw text here for simplicity. if (chunk.type === "thought") { thinkingContainer.append(chunk.value); } else if (chunk.type === "text") { From 41ececfa02014487591fae29cb756f2cb75781c4 Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 16:17:35 -0700 Subject: [PATCH 6/8] Update README.md Co-authored-by: Mike Wasserman --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index c7026fd..98e1c25 100644 --- a/README.md +++ b/README.md @@ -583,7 +583,7 @@ for await (const chunk of stream) { Rather than exposing model-specific token counts, user agents map each `effort` tier to an appropriate reasoning budget for the underlying model. This budget acts as an upper bound: models stop thinking early once they reach a conclusion, or transition to the final answer if the ceiling is reached. Following standard reasoning-model behavior, intermediate thoughts are stripped from the conversation history after each turn completes and do not permanently accumulate in `session.contextUsage`. -_Other possibilities under investigation include an `"auto"` effort level for dynamic prompt routing, as well as how `"thought"` blocks interleave with [Tool use](#tool-use) and [`samplingMode`](#configuration-of-sampling-modes)._ +_Open questions include: offering an `"auto"` effort level, how `"thought"` blocks interleave with [Tool use](#tool-use), and whether [`samplingMode`](#configuration-of-sampling-modes) options would apply during reasoning._ ### Session persistence and cloning From 17ced3e52eb029bfdb7a5d70a50afaf886504689 Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 16:22:04 -0700 Subject: [PATCH 7/8] Update README.md Co-authored-by: Reilly Grant --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 98e1c25..2d407aa 100644 --- a/README.md +++ b/README.md @@ -559,7 +559,7 @@ if (status !== "unavailable") { #### Viewing intermediate thoughts -By default, `includeThoughts` is `false` so existing applications only receive the final text response. When set to `true`, `promptStreaming()` and `prompt()` emit structured `{ type, value }` dictionaries (following the same pattern as [Tool use](#tool-use)) so applications can render a collapsible `"Thinking..."` UI: +By default, `includeThoughts` is `false` so applications only receive the final text response, matching the output with thinking disabled. When set to `true`, `promptStreaming()` and `prompt()` emit structured `{ type, value }` dictionaries (following the same pattern as [Tool use](#tool-use)) so applications can render a collapsible `"Thinking..."` UI: ```js const session = await LanguageModel.create({ From 9262bc4a4df09279039b57d81b452d4b50e1b202 Mon Sep 17 00:00:00 2001 From: Isaac Ahouma Date: Mon, 21 Sep 2026 16:25:37 -0700 Subject: [PATCH 8/8] Note emitThoughts naming and non-text thought representation in open questions --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 2d407aa..84c088a 100644 --- a/README.md +++ b/README.md @@ -583,7 +583,7 @@ for await (const chunk of stream) { Rather than exposing model-specific token counts, user agents map each `effort` tier to an appropriate reasoning budget for the underlying model. This budget acts as an upper bound: models stop thinking early once they reach a conclusion, or transition to the final answer if the ceiling is reached. Following standard reasoning-model behavior, intermediate thoughts are stripped from the conversation history after each turn completes and do not permanently accumulate in `session.contextUsage`. -_Open questions include: offering an `"auto"` effort level, how `"thought"` blocks interleave with [Tool use](#tool-use), and whether [`samplingMode`](#configuration-of-sampling-modes) options would apply during reasoning._ +_Open questions include: offering an `"auto"` effort level, naming (`emitThoughts` vs `includeThoughts`), designating thoughts via a separate field rather than `type: "thought"` to support non-text thoughts (e.g., images, audio, or [Tool use](#tool-use)), and whether [`samplingMode`](#configuration-of-sampling-modes) options would apply during reasoning._ ### Session persistence and cloning