Actions: google/gemma.cpp
Actions
Showing runs from all workflows
2,500+ workflow runs
2,500+ workflow runs
layout.total_bytes per attention subtask and kv_out_mem per token step by storing worker_workspaces and kv_out_mem on AttentionActivations / AttentionActivationsPtrs. This eliminates ~20M minor page faults and improves decode throughput by +18% to +26% on AMD Turin.
build
#8119:
Pull request #1025
opened
by
copybara-service
Bot