--n-cpu-moe36--spec-type draft-mtp --spec-draft-n-max3 does seem to speed up token generation for me
Can’t use llama-bench for MTP. In a basic tests it seems to improve from about 26 to 30 tokens per second output. But it seems to hurt my input speed from about 1300 pp down to 800.
--n-cpu-moe 36 --spec-type draft-mtp --spec-draft-n-max 3does seem to speed up token generation for meCan’t use llama-bench for MTP. In a basic tests it seems to improve from about 26 to 30 tokens per second output. But it seems to hurt my input speed from about 1300 pp down to 800.