❤️
LLM infers, human laughs
Interested in Desktop LLM inference (vllm-metal & DGX-Spark) and speculative decoding.
Highlights
- Pro
Pinned Loading
-
vllm-project/vllm-metal
vllm-project/vllm-metal PublicCommunity maintained hardware plugin for vLLM on Apple Silicon
-
SiliconBench
SiliconBench PublicThis repository includes code and materials for the paper "SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory Desktops"
-
vllm-project/speculators
vllm-project/speculators PublicA unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
-
vllm-project/vllm
vllm-project/vllm PublicA high-throughput and memory-efficient inference and serving engine for LLMs
-
whack-a-mole-local-llm
whack-a-mole-local-llm PublicPlay whack-a-mole using local llm (vllm-metal or vllm dgx-spark)
JavaScript
-
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.





