#835·kotaemon

Feature Request: Add audio/video document support via FunASR transcription

Author: LauraGPTCreated May 31, 2026Updated Jul 14, 2026

[!NOTE] License and capability clarification (2026-07-14): FunASR is a toolkit, not a single checkpoint. The FunASR and SenseVoice repository source code is MIT; model weights follow each model card. SenseVoiceSmall supports Chinese, Cantonese, English, Japanese, and Korean, and its weights use the linked FunASR Model Open Source License Agreement. Fun-ASR-Nano-2512 is Apache-2.0. Language coverage, punctuation, and performance depend on the selected model and runtime configuration.

Feature Request

kotaemon is excellent for document-based RAG QA. Adding audio/video document support via FunASR would extend the knowledge base to include audio content.

Use case: Upload meeting recordings, podcasts, lectures → FunASR transcribes → index and search like any document.

Why FunASR?

  • SenseVoice: 50+ languages, 5-10x faster than Whisper
  • Speaker diarization: Identifies speakers — improves retrieval for meetings
  • Timestamps: Enables citation back to audio position
  • Self-hosted: Aligns with kotaemon's local-first design
  • OpenAI-compatible API: Easy integration

Quick start:

bash
funasr-server --device cuda
# /v1/audio/transcriptions