Skip to content

tool use without auto execution - #162

Open
jingyun19 wants to merge 32 commits into
webmachinelearning:mainfrom
jingyun19:patch-1
Open

jingyun19 wants to merge 32 commits into
webmachinelearning:mainfrom
jingyun19:patch-1

Conversation

@jingyun19

@jingyun19 jingyun19 commented Nov 19, 2025 •

Copy link
Copy Markdown

Update explainer and spec to support tool use functionalities without automatic execution.

Explainer: added an example and explained how to make tool calls
Spec: reflect IDL changes in https://chromium-review.googlesource.com/c/chromium/src/+/7092943


Preview | Diff

Added detailed explanations for tool use modes, including examples for open loop and closed loop execution.
Updated README to clarify tool-call and tool-result usage.
Added new types and enums for tool calls and responses.
@jingyun19

Copy link
Copy Markdown
Author

@tomayac tomayac left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tried to make the code samples more readable and correct. Maybe consider running them all through a tool like prettier, which catches typos like missing commas or parentheses.

As general feedback, could the explainer outline why developers would choose closed vs. open?

Comment thread README.md Outdated
Comment thread README.md
Comment thread README.md Outdated
Comment thread README.md Outdated
Comment thread README.md Outdated
Comment thread README.md Outdated
Comment thread README.md Outdated
Comment thread README.md Outdated
Comment thread README.md Outdated
jingyun19 and others added 5 commits November 26, 2025 09:08
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Added explanation about automatic execution and constraints in planner loop.
@jingyun19

jingyun19 commented Dec 1, 2025 •

Copy link
Copy Markdown
Author

I added a new section to describe use cases where open loop is preferred. cc @tomayac

@reillyeon

Copy link
Copy Markdown
Collaborator

I hadn't previously considered the context compression use case. That's interesting and motivating to enable developers to manipulate the conversation at this low level.

@nico-martin nico-martin left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this implementation has a few weaknesses when it comes to distinguishing between a Message and a Message.Content element:
Message: Can have a specific role (whether it comes from the user, the assistant, or a tool call); it essentially describes the sender.
Message.Content: There can be multiple instances per Message; it describes the type of content.
I tried to make this concrete with a couple of comments.

Comment thread index.bs Outdated
Comment thread index.bs
Comment thread README.md Outdated
Comment thread README.md Outdated
Comment thread README.md Outdated

#### Do I need auto execution?

In general, automatic execution is suitable for use cases where the model quality is good enough via prompt tuning. That can either mean you are tolerable for certain mistakes that the model makes when making tool calls, or the task is simple enough for the model to handle (e.g, just a few distinct tools, short and clean tool output, short context window, etc)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not sure if I get this right. For the user, both, the closed and the open loop, are executed automatically. The only difference is that in an open loop, the developer has to execute the tools and start the next generation, while in the closed loop the loop will run without any extra steps.
Also if I dont want to have "automatic execution" as a developer, I could always intercept in the execute function. I would even argue for the wohle LLM conversation it is better to intercept a tool execution inside the execute function. Because then it allows you to return a reason why the tool was not executed intead of letting the model generate the tool call and then it does not know why it was not executed.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From the API client's perspective, there's no automatic execution in open loop. The API client need to read the tool name and arguments and invoke the functino themselves.
I also updated those sections in the explainer, PTAL and let me know if they make more sense now!

@anssiko

anssiko commented Aug 14, 2026

Copy link
Copy Markdown
Member

@jingyun19 does the current PR reflect the Chromium implementation? The initial CL was merged. Remaining work?

@jingyun19

Copy link
Copy Markdown
Author

@jingyun19 does the current PR reflect the Chromium implementation? The initial CL was merged. Remaining work?

Yes it reflects the Chromium implementation. I believe the only remaining work is to support ToolCall type in input.

However, for us to do a dev trial, we also have the remaining work to fully support parsing and formatting tool types in the inference engine infrastructure.

jingyun19 and others added 10 commits September 28, 2026 14:43
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
@jingyun19

Copy link
Copy Markdown
Author

The updated spec is ready for review.

@reillyeon @michaelwasserman

Comment thread index.bs
Comment thread index.bs
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread index.bs
1. If |content|["{{LanguageModelMessageContent/value}}"] is not a {{LanguageModelToolResponse}}, then throw a {{TypeError}}.

1. If |content|["{{LanguageModelMessageContent/value}}"] is a {{LanguageModelToolSuccess}}, then [=list/for each=] |resultItem| of |content|["{{LanguageModelMessageContent/value}}"]'s [=LanguageModelToolSuccess/result=]:
1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/image}}" or "{{LanguageModelToolResultType/audio}}" and the user agent does not support multimodal tool result content, then throw a "{{NotSupportedError}}" {{DOMException}}.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel like this deserves a note that implementations should only expect to receive tool call response objects that they themselves could generate, so this case should never be hit in practice..

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Because LanguageModelToolSuccess is constructed by API client, a developer may pass { type: "image", value: ... } or { type: "audio", value: ... }. If the expectedInputs only supports "text" and "object" tool results (or when the session wasn't configured with "image" / "audio" in expectedInputs), this NotSupportedError check can be hit in practice

Comment thread index.bs
Comment thread README.md Outdated
Comment thread README.md Outdated
Comment thread index.bs

@michaelwasserman michaelwasserman left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for picking up this complex and important spec work again!

Comment thread index.bs Outdated
Comment thread index.bs
Comment thread index.bs
Comment thread index.bs
Comment thread index.bs Outdated
Comment thread README.md Outdated
Comment thread README.md Outdated
console.log(result);
```

Note that a `"tool-response"` message should immediately follow the `"tool-call"` generated by the model.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should the spec yield an error if that isn't the case?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Because "measureContextUsage(input)" runs validate and canonicalize a prompt on input on just the input, and clients may want to measure the token cost of a tool-response message before or independent of the current session state. If validate and canonicalize a prompt required a preceding tool-call, measureContextUsage() on a tool-response would throw.

Also for the model, the sequence will actually get prefilled to the model, so nothing throws an error in the stack, it just produces a non-standard token sequence in the model context, which is why I put it as a note (suggestion) rather than a normative error step.

Comment thread index.bs Outdated
Comment thread index.bs Outdated
Comment thread README.md Outdated
@michaelwasserman

Copy link
Copy Markdown
Collaborator

FYI @FrankLi-MSFT and @sushraja-msft

Comment thread README.md Outdated
Comment thread README.md Outdated

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants