转推:用 Qwen3.8-27B+MTP 时升级 DFlash 可提速,附 llama.cpp 命令
RT @ggerganov:如果你正在使用 Qwen3.8-27B + MTP,请务必升级到 DFlash 以获得额外提速:
llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-dflash --spec-draft-n-max 7
需要最新的 llama.cpp v0.6.0
llama serve -hf ggml-org/Qwen3.8-27B-GGUF --spec-type draft-dflash --spec-draft-n-max 7
需要最新的 llama.cpp v0.6.0