针对 AGX Orin 优化的 qwen3.5 / qwen3.6 支持 POC
作者: alansrobotlab2创建于 2026年5月1日更新于 2026年7月27日
标签new-models
qwen3.6-35b-a3b benchmark results
| tg | MLC q4f16_1 v2+FI | LLaMA.cpp Q4_K_S¹ | LLaMA.cpp Q4_K_XL² | ratio (vs Q4_K_XL) |
|---|---|---|---|---|
| 512 | 54.46 | 29.19 | 28.26 | 1.927× |
| 1024 | 54.30 | 29.30 | 28.21 | 1.925× |
| 2048 | 54.07 | 29.31 | 28.13 | 1.922× |
| 4096 | 53.69 | 29.04 | 28.07 | 1.913× |
| 8192 | 53.00 | 28.48 | 27.86 | 1.902× |
| Δ tg512→tg8192 | −2.7 % | −2.4 % | −1.4 % | flat |
qwen3.5-0.8b benchmark results
| tg | LLaMA.cpp Q4_K_XL (pure tg) | MLC q4f16_g16e + FI | ratio |
|---|---|---|---|
| 512 | 100.3 | 134.82 | 1.345× |
| 1024 | 100.1 | 134.29 | 1.341× |
| 2048 | 99.7 | 133.54 | 1.340× |
| 4096 | 98.0 | 132.17 | 1.349× |
| 8192 | 96.5 | 129.59 | 1.343× |
Claude's summary
Starting point (~10 tps). The benchmark harness was reporting blended pp+tg numbers, batch_decode was being routed through the wrong path, and the dlight-default GEMM grid was tuned for Hopper.
内容来源: mlc-ai/mlc-llm