Skip to content
@ToolSlack

ToolSlack

ToolSlack: Exploiting Tool Execution Windows for Efficient LLM Agent Serving

ToolSlack

ToolSlack schedules agent memory and exact-prefix KV preparation during real tool execution. The anonymous artifact includes the dense Qwen3-8B / LangGraph / LangMem implementation, CPU correctness checks, and a native GPU benchmark workflow.

Code and reproduction instructions

Quick start

git clone https://github.com/ToolSlack/code.git
cd code
./run.sh cpu

The CPU command bootstraps the pinned environment and runs correctness checks on Linux or macOS. Python 3.9 or newer is required to start it; no API key, model weights, CUDA installation, or GPU allocation is needed.

For the native GPU workflow, first read the hardware and reproduction prerequisites. Use Linux x86_64 with one independently reserved, empty GPU, and replace 0 with its physical nvidia-smi index:

./run.sh smoke --gpu-indices 0 --exclusive-gpus
./run.sh full --gpu-indices 0 --exclusive-gpus

CPU validation does not establish GPU performance. Fresh native-engine installation, GPU smoke, native DMA correctness, and paper performance reproduction remain unverified; see the validation status and scope.

Popular repositories Loading

  1. code code Public

    ToolSlack: Exploiting Tool Execution Windows for Efficient LLM Agent Serving

    Python

  2. .github .github Public

    ToolSlack organization profile

Repositories

Showing 2 of 2 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…