_generate_planning_step should support image inputs
Author: iamxtsoCreated May 22, 2025Updated Sep 17, 2026
Labelsenhancement
Currently, _generate_planning_step only processes textual messages. In multimodal workflows, it’s often necessary to include images (e.g., screenshots or visual context) as part of the reasoning input.
Source: huggingface/smolagents