Could you give me some guidance on how I can adapt this to vllm like llava?

Author: Yingshu-LiCreated Sep 28, 2024Updated May 16, 2025

Hello,

I’m currently exploring how to visualize the heatmap on LLAVA or other kinds of multimodal large language model to understand the model’s focus during text generation. I am familiar with using Grad-CAM for single-target classification tasks. However, with LLAVA generating complete sentences, I’m unsure how to obtain heatmaps for individual words. Could you provide any guidance or advice on how to approach this?

Source: jacobgil/pytorch-grad-cam