Your Question
Summary
This RFC proposes enabling DSpark speculative decoding support on Ascend NPU within the Vime framework. The change adapts the NPU batch-reordering logic to accommodate DSpark's speculative-decoding layout requirements, allowing DSpark to be launched alongside standard attention groups on Ascend hardware.
https://github.com/momo609/vime/tree/ascend
Design
#377
Usage
DSpark is enabled through the standard vLLM speculative-config interface. The configuration is passed as a JSON string via --vllm-speculative-config, and the launch script is invoked with the usual RL-training arguments plus a set of vLLM-specific flags. The two essential fields for DSpark are method (set to "dspark") and model (pointing to the DSpark draft checkpoint). The optional fields control the number of speculative tokens, tensor-parallel size of the draft model, and where the updated draft weights are saved after training.
example:
SPEC_CONFIG=$(printf '%s' \
'{"method":"dspark",' \
'"model":"{your path}/dspark_qwen3_4b_block7",' \
'"num_speculative_tokens":7,' \
'"draft_tensor_parallel_size":1}')
VLLM_ARGS=(
--rollout-num-gpus-per-engine 4
--vllm-weight-sync-mode native
--vllm-enable-sleep-mode
--vllm-gpu-memory-utilization 0.6
--vllm-max-model-len 4096
--vllm-speculative-config "$SPEC_CONFIG"
--draft-save-hf {your path}/dspark_qwen3_4b_block7_after
)
TEST
we have test Qwen3-4B with dspark on Ascend:
test1.bmp
test2.bmp
test3.bmp
What I've Tried
test
Environment (if relevant)
Additional Context
No response
Pre-submission Checklist
Your Question
Summary
This RFC proposes enabling DSpark speculative decoding support on Ascend NPU within the Vime framework. The change adapts the NPU batch-reordering logic to accommodate DSpark's speculative-decoding layout requirements, allowing DSpark to be launched alongside standard attention groups on Ascend hardware.
https://github.com/momo609/vime/tree/ascend
Design
#377
Usage
DSpark is enabled through the standard vLLM speculative-config interface. The configuration is passed as a JSON string via --vllm-speculative-config, and the launch script is invoked with the usual RL-training arguments plus a set of vLLM-specific flags. The two essential fields for DSpark are method (set to "dspark") and model (pointing to the DSpark draft checkpoint). The optional fields control the number of speculative tokens, tensor-parallel size of the draft model, and where the updated draft weights are saved after training.
example:
TEST
we have test Qwen3-4B with dspark on Ascend:
test1.bmp
test2.bmp
test3.bmp
What I've Tried
test
Environment (if relevant)
Additional Context
No response
Pre-submission Checklist