百科.dev
全部条目趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
C

csm

> 编程语言
开源

对话式语音生成模型

14.7K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

对话式语音生成模型

中央管理系统 2025/05/20 - CSM在Hugging Face 变形器的4.52.1版本中是本地有用的,我们模式2025/03/13中有更多的信息,我们正在发布1B CSM的变体。 检查站设在Hugging Face上。 --- CSM(Conversational Speech Model)是来自芝麻的语音生成模型,从文本和音频输入生成RVQ音频代码. 该模型架构使用LLaMA主干线和更小的音频解码器来生成米米音频代码. CSM的一个微调变体赋予互动语音演示功能, 还有一个主机Hugging Face空间可供测试音频生成. A 兼容CUDA的GPU 该代码已在CUDA 12.4和 12.6上测试过,但也可能在其他版本上工作类似,Python 3.10被推荐,但较新的版本可能很好 对于一些音频操作,ffmpeg可能被要求访问后继.

核心特点

  • •A CUDA-compatible GPU
  • •The code has been tested on CUDA 12.4 and 12.6, but it may also work on other versions
  • •Similarly, Python 3.10 is recommended, but newer versions may be fine
  • •For some audio operations, ffmpeg may be required
  • •Access to the following Hugging Face models:
  • •Llama-3.2-1B
  • •Impersonation or Fraud: Do not use this model to generate speech that mimics real individuals without their explicit consent.
  • •Misinformation or Deception: Do not use this model to create deceptive or misleading content, such as fake news or fraudulent calls.
  • •Illegal or Harmful Activities: Do not use this model for any illegal, harmful, or malicious purposes.

> 标签

Python

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月9日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言