Conversation
Added detailed explanations for tool use modes, including examples for open loop and closed loop execution.
Updated README to clarify tool-call and tool-result usage.
Added new types and enums for tool calls and responses.
tomayac
left a comment
There was a problem hiding this comment.
Tried to make the code samples more readable and correct. Maybe consider running them all through a tool like prettier, which catches typos like missing commas or parentheses.
As general feedback, could the explainer outline why developers would choose closed vs. open?
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Added explanation about automatic execution and constraints in planner loop.
|
I added a new section to describe use cases where open loop is preferred. cc @tomayac |
|
I hadn't previously considered the context compression use case. That's interesting and motivating to enable developers to manipulate the conversation at this low level. |
nico-martin
left a comment
There was a problem hiding this comment.
I think this implementation has a few weaknesses when it comes to distinguishing between a Message and a Message.Content element:
Message: Can have a specific role (whether it comes from the user, the assistant, or a tool call); it essentially describes the sender.
Message.Content: There can be multiple instances per Message; it describes the type of content.
I tried to make this concrete with a couple of comments.
|
|
||
| #### Do I need auto execution? | ||
|
|
||
| In general, automatic execution is suitable for use cases where the model quality is good enough via prompt tuning. That can either mean you are tolerable for certain mistakes that the model makes when making tool calls, or the task is simple enough for the model to handle (e.g, just a few distinct tools, short and clean tool output, short context window, etc) |
There was a problem hiding this comment.
Not sure if I get this right. For the user, both, the closed and the open loop, are executed automatically. The only difference is that in an open loop, the developer has to execute the tools and start the next generation, while in the closed loop the loop will run without any extra steps.
Also if I dont want to have "automatic execution" as a developer, I could always intercept in the execute function. I would even argue for the wohle LLM conversation it is better to intercept a tool execution inside the execute function. Because then it allows you to return a reason why the tool was not executed intead of letting the model generate the tool call and then it does not know why it was not executed.
There was a problem hiding this comment.
From the API client's perspective, there's no automatic execution in open loop. The API client need to read the tool name and arguments and invoke the functino themselves.
I also updated those sections in the explainer, PTAL and let me know if they make more sense now!
|
@jingyun19 does the current PR reflect the Chromium implementation? The initial CL was merged. Remaining work? |
Yes it reflects the Chromium implementation. I believe the only remaining work is to support ToolCall type in input. However, for us to do a dev trial, we also have the remaining work to fully support parsing and formatting tool types in the inference engine infrastructure. |
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
Co-authored-by: Thomas Steiner <tomac@google.com>
|
The updated spec is ready for review. |
| 1. If |content|["{{LanguageModelMessageContent/value}}"] is not a {{LanguageModelToolResponse}}, then throw a {{TypeError}}. | ||
|
|
||
| 1. If |content|["{{LanguageModelMessageContent/value}}"] is a {{LanguageModelToolSuccess}}, then [=list/for each=] |resultItem| of |content|["{{LanguageModelMessageContent/value}}"]'s [=LanguageModelToolSuccess/result=]: | ||
| 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/image}}" or "{{LanguageModelToolResultType/audio}}" and the user agent does not support multimodal tool result content, then throw a "{{NotSupportedError}}" {{DOMException}}. |
There was a problem hiding this comment.
I feel like this deserves a note that implementations should only expect to receive tool call response objects that they themselves could generate, so this case should never be hit in practice..
There was a problem hiding this comment.
Because LanguageModelToolSuccess is constructed by API client, a developer may pass { type: "image", value: ... } or { type: "audio", value: ... }. If the expectedInputs only supports "text" and "object" tool results (or when the session wasn't configured with "image" / "audio" in expectedInputs), this NotSupportedError check can be hit in practice
michaelwasserman
left a comment
There was a problem hiding this comment.
Thanks for picking up this complex and important spec work again!
| console.log(result); | ||
| ``` | ||
|
|
||
| Note that a `"tool-response"` message should immediately follow the `"tool-call"` generated by the model. |
There was a problem hiding this comment.
Should the spec yield an error if that isn't the case?
There was a problem hiding this comment.
Because "measureContextUsage(input)" runs validate and canonicalize a prompt on input on just the input, and clients may want to measure the token cost of a tool-response message before or independent of the current session state. If validate and canonicalize a prompt required a preceding tool-call, measureContextUsage() on a tool-response would throw.
Also for the model, the sequence will actually get prefilled to the model, so nothing throws an error in the stack, it just produces a non-standard token sequence in the model context, which is why I put it as a note (suggestion) rather than a normative error step.
|
FYI @FrankLi-MSFT and @sushraja-msft |
Clarified expectedInputs and LanguageModelToolSuccess.result structure in README.
Clarify behavior of session.prompt() with expectedOutputs.
Update explainer and spec to support tool use functionalities without automatic execution.
Explainer: added an example and explained how to make tool calls
Spec: reflect IDL changes in https://chromium-review.googlesource.com/c/chromium/src/+/7092943
Preview | Diff