#3492·mlc-llm

针对 AGX Orin 优化的 qwen3.5 / qwen3.6 支持 POC

作者: alansrobotlab2创建于 2026年5月1日更新于 2026年7月27日
标签new-models

qwen3.6-35b-a3b benchmark results

tg MLC q4f16_1 v2+FI LLaMA.cpp Q4_K_S¹ LLaMA.cpp Q4_K_XL² ratio (vs Q4_K_XL)
512 54.46 29.19 28.26 1.927×
1024 54.30 29.30 28.21 1.925×
2048 54.07 29.31 28.13 1.922×
4096 53.69 29.04 28.07 1.913×
8192 53.00 28.48 27.86 1.902×
Δ tg512→tg8192 −2.7 % −2.4 % −1.4 % flat

qwen3.5-0.8b benchmark results

tg LLaMA.cpp Q4_K_XL (pure tg) MLC q4f16_g16e + FI ratio
512 100.3 134.82 1.345×
1024 100.1 134.29 1.341×
2048 99.7 133.54 1.340×
4096 98.0 132.17 1.349×
8192 96.5 129.59 1.343×

Claude's summary

Starting point (~10 tps). The benchmark harness was reporting blended pp+tg numbers, batch_decode was being routed through the wrong path, and the dlight-default GEMM grid was tuned for Hopper.