How to use the mtp gguf file in llama -server

#7
by nps798 - opened

How to use the mtp gguf file in llama -server ? any comment on how to properly set up the mtp part
Upload Step3.7-flash-mtp-BF16.gguf
Step3.7-flash-mtp-Q8_0.gguf??

StepFun org

You can try the following command:

./llama-server \
    -m Step-3.7-Q4_K_S.gguf \
    --spec-type draft-mtp \
    --spec-draft-model Step3.7-flash-mtp-Q8_0.gguf \
    -c 35000 \
    -np 1 \
    -b 2048 \
    -ub 1024 \
    --temp 0 \
    --spec-draft-n-max 3 \
    --spec-draft-p-min 0.06 \
    --host 127.0.0.1 \
    --port 8080

I’m not getting any increase in speed on my Mac with MTP. Is this something you’re seeing on your end as well?

Also, really looking forward to next model from this team

StepFun org

@MilestoneCap Yeah, MTP may not give much of a speedup on mac right now. It’s partly due to how MoE models work and partly the backend implementation. llama.cpp is still working on optimizations here.

And for the next model — shouldn’t be too long :)

@MilestoneCap Yeah, MTP may not give much of a speedup on mac right now. It’s partly due to how MoE models work and partly the backend implementation. llama.cpp is still working on optimizations here.

And for the next model — shouldn’t be too long :)

Awesome, hope it's still 128GB friendly. 3.7 was my favorite model to run locally until a few weeks ago

Sign up or log in to comment