[Question] How fast is inference supposed to be?
Author: SouljaVRCreated Mar 10, 2025Updated Feb 5, 2026
Hi, I am using an RTX 4090 for inference with CUDA PyTorch 12.6. To inference this text takes me around 25-30 seconds with any length audio input for voice cloning:
With the direction of the Sun, the orientation of the solar panels, and extreme cold temperatures in the crater, Intuitive Machines does not expect Athena to recharge, the company said in a statement. The mission has concluded, and teams are continuing to assess the data collected throughout the mission.
So 25 second generation time for around 15s of audio - Is this speed of inference correct? Or is mine working slower than expected? I would like to see some benchmark data or something if possible, thank you.
Source: SparkAudio/Spark-TTS