Logan Chu
|
5cdd3d1dad
|
model : fix MTP context kv cache allocation for deepseek2, glm4moe, … (#28630)
* model : fix MTP context kv cache allocation for deepseek2, glm4moe, cohere2moe architectures (#28626)
* model: add inverse architecture gating and comprehensive architecture testing for mtp layer filtering
* model : slim NextN filter comment, drop test-llama-archs changes
|
2026-09-11 12:02:31 +03:00 |
|