Hey there,
I followed you over from huggingface :-) interesting project!
I was curious about the following:
DFlash + hybrid (mlx_mtp.dflash, mlx_mtp.hybrid)
Block-diffusion external drafter (e.g. z-lab/Qwen3.6-27B-DFlash) with a tunable block size, plus a unified loop that picks MTP vs DFlash per context by tokens-per-second. Requires DFlash drafter weights — optional, "where available."
Would you be able to add to the readme any references and/or knowledge that was used for designing the logic behind "picks MTP vs DFlash per context by tokens-per-second"?
Hey there,
I followed you over from huggingface :-) interesting project!
I was curious about the following:
Would you be able to add to the readme any references and/or knowledge that was used for designing the logic behind "picks MTP vs DFlash per context by tokens-per-second"?