espnet

语音识别、合成、翻译一把抓,一个工具箱全搞定

端到端语音处理工具包,覆盖语音识别、语音合成、语音翻译、说话人识别等多项任务。提供统一训练框架、丰富预训练模型与评测流程,是学术界与工业界语音技术研究的常用基础设施。

开源 开源免费 音频音乐

项目数据

分类音频音乐
开发团队espnet
所属国家
定价模式开源免费
价格区间
是否开源
开源协议Apache-2.0
主要语言Python
技术栈/模型chainer,deep-learning,end-to-end,kaldi,machine-translation,pytorch,singing-voice-synthesis,speaker-diarization,speech-enhancement,speech-recognition
GitHub 星标★ 9916
30天Star增速
HF 下载量
上线时间2017-12-13 00:00:00
最近更新2026-08-07 00:00:00
维护状态活跃
中文支持
访问方式
移动端支持
综合评分
收录时间2026-08-09
浏览次数0

核心亮点

    不足之处

      适用场景

        替代项目

        项目介绍

        端到端语音处理工具包,覆盖语音识别、语音合成、语音翻译、说话人识别等多项任务。提供统一训练框架、丰富预训练模型与评测流程,是学术界与工业界语音技术研究的常用基础设施。。主要使用 Python 开发,GitHub 星标 9916。

        上一篇:piper

        下一篇:EmotiVoice

        同类项目推荐

        StyleTTS2 开源

        StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversari···

        ★ 6328 2026-08-09
        tuneflow-py 开源

        + Build your music algorithms and AI models with the next-gen DAW

        ★ 890 2026-08-09
        open-webui-tools 开源

        Open‑WebUI Tools is a modular toolkit designed to extend and enrich your Open WebUI···

        ★ 785 2026-08-09
        audio-diffusion 开源

        Apply diffusion models using the new Hugging Face diffusers package to synthesize mu···

        ★ 793 2026-08-09