#228·MeloTTS

MPS device limitation error when running TTS on Apple Silicon Mac

Author: kiwamizamuraiCreated Jan 13, 2025Updated Feb 26, 2026

Environment

  • OS: macOS 14.7
  • Hardware: Apple M3 Pro
  • Python: 3.11.11
  • MeloTTS: 0.1.2

Dependencies

dependencies = [ "transformers==4.27.4", "torch>=2.0.0", "sentencepiece>=0.1.99", "click>=8.1.7", "rich>=13.7.0", "pydantic>=2.6.0", "melotts @ git+https://github.com/myshell-ai/MeloTTS.git", "unidic>=1.1.0", "sounddevice>=0.5.1", "nltk>=3.8.1", ]

Issue Description

When trying to run TTS on an Apple Silicon Mac using the MPS (Metal Performance Shaders) device, the following error occurs:

Error: Output channels > 65536 not supported at the MPS device. As a temporary fix, you can set the environment variable `PYTORCH_ENABLE_MPS_FALLBACK=1` to use the CPU as a fallback for this op. WARNING: this will be slower than running natively on MPS.

This error occurs during the speech synthesis process when calling tts_to_file() method.

Steps to Reproduce

  1. Initialize TTS with device='auto' or device='mps'
  2. Call tts_to_file() with any English text
  3. Error occurs during model inference

Current Workaround

Currently, we have two workarounds:

  1. Set environment variable: PYTORCH_ENABLE_MPS_FALLBACK=1
  2. Force CPU usage by initializing TTS with device='cpu'

Both workarounds result in slower performance compared to potential MPS acceleration.

Additional Context

This seems to be related to a limitation in PyTorch's MPS backend regarding the maximum number of output channels. It would be beneficial if the model architecture could be adjusted to work within MPS device limitations, or if there's a way to optimize the operations to stay under the 65536 channel limit.

Code Example

python
from melo.api import TTS

# This fails on MPS
engine = TTS(language='EN', device='auto')
audio = engine.tts_to_file("Test text", speaker_id=0, output_path=None)

# Current workaround
engine = TTS(language='EN', device='cpu')  # Force CPU usage
audio = engine.tts_to_file("Test text", speaker_id=0, output_path=None)