In the example provided in the code, the image generation function only supports text input. Can the image generation support multimodal input? How should image input be handled?
Author: missyncCreated Jul 12, 2025Updated Jul 12, 2025
In the example provided in the code, the image generation function only supports text input. Can the image generation support multimodal input? How should image input be handled?
Source: deepseek-ai/Janus