Qwen 3.5

#1
by NIK2703 - opened

What about Qwen 3.5? I tried a quantized version from Yoursmiling, but it’s crashing when loading on my mobile device https://huggingface.co/Yoursmiling/Qwen3.5-4B-LiteRT/discussions/1#69d10722f5041eaf894f2315

LiteRT Community (FKA TFLite) org

Qwen 3.5 is a great model that it uses much less KV cache memory(and memory bandwidth) than previous ones. It requires a new cache contract that LiteRT and LiteRT-LM do not support yet. We are working on adding the support.

How about bonsai 1bit and ternary models? Is there a plan to support them in the future as well?
69e10ee7705d14e8165aa4f5_p-vs-s

Sign up or log in to comment