跳到正文
Joeplover
学习笔记·2026-04-13·约 4 分钟阅读

edge-tts:微软 Edge 在线语音合成

Python 模块,通过 Microsoft Edge TTS 服务实现文本转语音,支持流式播放和字幕生成。

语音合成

项目背景

edge-tts 是一个 Python 库,它封装了微软 Edge 浏览器的在线 TTS(Text-to-Speech)服务。为什么需要这个项目?因为传统的 TTS 方案要么价格昂贵(如各大云厂商的 TTS API),要么需要下载数 GB 的模型文件(如 Coqui TTS、Tortoise TTS),而 Edge 浏览器内置的 TTS 服务完全免费、音质优秀、支持超过 400 种语音,却一直没有官方的 API。

这个项目通过逆向 Edge 浏览器的 WebSocket 通信协议,以纯 Python 实现的方式调用 Edge TTS 服务,让开发者可以在任何 Python 项目中使用高质量的语音合成能力。

技术选型

组件选择理由
网络请求aiohttp异步 HTTP/WebSocket 客户端
SSL 证书certifi可靠的 CA 证书管理
表格输出tabulateCLI 模式下格式化展示语音列表
字幕生成内置 SubMakerSRT 格式字幕生成
Python 版本3.8+广泛兼容
包管理setuptools标准 Python 打包

架构设计

edge-tts 的核心架构非常简洁,分为三层:

  1. CLI 层(__main__.py / util.py):提供命令行接口,支持 -t(文本)、-f(文件)、-v(语音)、--write-media(输出音频)、--write-subtitles(输出字幕)等参数
  2. 通信层(communicate.py):核心模块,实现与 Edge TTS 服务的 WebSocket 通信协议
  3. 工具层(submaker.py、voices.py 等):字幕生成、语音列表查询等辅助功能

核心实现

与 Edge TTS 服务的 WebSocket 通信

核心类是 Communicate,它封装了与 Edge TTS WebSocket 端点的完整握手和流式数据传输:

class Communicate:
    def __init__(self, text: str, voice: str = DEFAULT_VOICE,
                 rate: str = "+0%", volume: str = "+0%",
                 pitch: str = "+0Hz", proxy: Optional[str] = None):
        self.text = text
        self.voice = voice
        self.rate = rate
        self.volume = volume
        self.pitch = pitch
        self.proxy = proxy

    async def stream(self) -> AsyncGenerator[dict, None]:
        async with aiohttp.ClientSession() as session:
            async with session.ws_connect(
                WSS_URL,
                headers=WSS_HEADERS,
                proxy=self.proxy,
                ssl=_SSL_CTX,
            ) as ws:
                # 1. 发送配置消息
                await ws.send_json({
                    "context": {
                        "synthesis": {
                            "audio": {
                                "metadataoptions": {
                                    "sentenceBoundaryEnabled": "true",
                                    "wordBoundaryEnabled": "true"
                                },
                                "outputFormat": "audio-24khz-96kbitrate-mono-mp3"
                            }
                        }
                    }
                })
                # 等待配置确认
                msg = await ws.receive()
                # 2. 发送 SSML 格式的合成请求
                ssml = self._build_ssml()
                await ws.send_str(ssml)
                # 3. 流式接收音频数据
                async for chunk in self._receive_audio(ws):
                    yield chunk
    async def _receive_audio(self, ws):
        while True:
            msg = await ws.receive()
            if msg.type == aiohttp.WSMsgType.BINARY:
                # 解析二进制帧
                # Turn data → headers + audio data
                # ...
                yield {"type": "audio", "data": audio_bytes}
            elif msg.type == aiohttp.WSMsgType.TEXT:
                data = json.loads(msg.data)
                if data.get("type") == "WordBoundary":
                    yield {
                        "type": "WordBoundary",
                        "offset": data.get("offset"),
                        "duration": data.get("duration"),
                        "text": data.get("text", {}).get("Text", ""),
                    }
            elif msg.type == aiohttp.WSMsgType.CLOSED:
                break

CLI 接口

async def _run_tts(args: UtilArgs) -> None:
    communicate = Communicate(args.text, args.voice,
                               rate=args.rate, volume=args.volume,
                               pitch=args.pitch, proxy=args.proxy)
    submaker = SubMaker()
    audio_file = open(args.write_media, "wb") if args.write_media else sys.stdout.buffer
    sub_file = open(args.write_subtitles, "w", encoding="utf-8") if args.write_subtitles else None

    async for chunk in communicate.stream():
        if chunk["type"] == "audio":
            audio_file.write(chunk["data"])
        elif chunk["type"] in ("WordBoundary", "SentenceBoundary"):
            submaker.feed(chunk)

    if sub_file:
        sub_file.write(submaker.get_srt())

声音列表查询

async def list_voices(proxy: Optional[str] = None) -> list[dict]:
    """获取所有可用的语音列表"""
    async with aiohttp.ClientSession() as session:
        async with session.get(
            "https://speech.platform.bing.com/consumer/speech/synthesize/"
            "readaloud/voices/list?trustedclienttoken=6A5DAA5BB23972AE",
            proxy=proxy,
        ) as resp:
            voices = await resp.json()
            return voices

踩坑记录

1. SSL 证书验证
在部分 Linux 发行版上,系统 CA 证书不完整,导致 WebSocket 连接握手失败。解决方案:使用 certifi 提供的最新 CA 证书包,显式创建 SSL 上下文。

2. 特殊字符导致服务报错
OCR 识别出的文本可能包含垂直制表符(VT, 0x0B)等控制字符,Edge TTS 服务无法处理这些字符。解决方案:实现 remove_incompatible_characters 函数,过滤掉 0-8、11-12、14-31 范围内的控制字符。

3. 音频格式兼容性
Edge 返回的是 24kHz 96kbps 的 MP3 格式,部分老旧系统或需要特定格式的场景可能需要转码。解决方案:文档中建议配合 ffmpeg 使用进行格式转换,不作为内置依赖以减少包体积。

4. 代理环境支持
在企业网络环境中,需要通过 HTTP 代理访问外部服务。解决方案:所有网络请求均支持 proxy 参数,透传给 aiohttp 的代理配置。

总结

edge-tts 给我的最大启发是:有时候最好的 API 不是付费的云服务,而是"发现"已有的能力。Edge 浏览器每天为成千上万的用户提供无障碍语音服务,它的 TTS 引擎效果堪比商业方案。通过 Python 封装这个能力,我们可以在视频配音、语音助手、辅助阅读、自动化播报等场景中低成本地使用高质量的语音合成。

这个项目的代码量不大(约 658 行核心代码),但对网络协议的理解要求很高。WebSocket 的帧结构、SSML 语音合成标记语言、音频流的处理,都值得深入学习。目前社区中 edge-tts 已经在 GitHub 上获得超过 10k Stars,被广泛应用于各种 AI 项目中。