Qwen 上的 TensorRT 图像编辑

作者: ohnlily创建于 2026年3月12日更新于 2026年4月24日

We are currently using Alibaba's Qwen image edit model, and we have compiled it using TensorRT to speed up inference. However, the current AOT speedup is only 16% (this evaluation specifically targets the Transformer at fp16 accuracy), with two 2K input images. This speedup is far from our expectations. We guess you may have conducted similar compiler acceleration experiments as well. Therefore, we would like to seek your suggestions to improve our speedup performance. Below are some details of our compilation process.

内容来源: QwenLM/Qwen-Image