Skip to content

[Blog] Exploring Speculative Decoding in vLLM on AMD GPUs - #310

Open
junkang1991 wants to merge 1 commit into
vllm-project:mainfrom
junkang1991:spec-dec-study
Open

[Blog] Exploring Speculative Decoding in vLLM on AMD GPUs#310
junkang1991 wants to merge 1 commit into
vllm-project:mainfrom
junkang1991:spec-dec-study

Conversation

@junkang1991

Copy link
Copy Markdown
Contributor

Speculative decoding allows vLLM to verify multiple drafted tokens in a single target-model pass. In our experiments, its effect on output-token throughput varied across drafting methods and proposal lengths, and also depended on the model family, draft checkpoint, workload, and acceptance behavior.

Signed-off-by: Jun Kang Chow <junkangchow@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant