Files
BaiLongma/待办.txt
2026-08-23 13:07:58 +08:00

60 lines
3.7 KiB
Plaintext
Raw Permalink Blame History

This file contains invisible Unicode characters
This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
待办事项(2026-08-09 更新)
============================
【1】✅ CosyVoice2-0.5B 本地语音合成(已完成,替代 VoxCPM2)
------------------------------------------------------------
背景:VoxCPM2 验证后 RTF≈20(3060 8G 上一句话 2 分钟),性能不达标已弃用。
CosyVoice2-0.5B:fp16 下显存 3.9G/8G,RTF≈2.7,流式首包 ~7s,女声克隆(MeloTTS 参考音频)。
已完成:
✅ 模型:src/voice/models/CosyVoice2-0.5B/(3.8G:llm.pt 2G + flow.pt + hift.pt + speech_tokenizer_v2.onnx + campplus.onnx + CosyVoice-BlankEN/)
✅ 源码:src/voice/cosyvoice-src/(cosyvoice 包 + third_party/Matcha-TTS + asset/)
✅ 依赖:hyperpyyaml、onnxruntime、x-transformers、transformers 4.51.3、wetext、pypinyin、conformer、lightning、matcha(源码)、pyworld 0.3.5、openai-whisper、modelscope、torchcodec 0.9.1 等
✅ torch cu126 2.9.1 已装(注意:别装 torchvision,会把 torch 覆盖回 CPU 版!)
✅ 常驻服务:src/voice/tts_cosyvoice.py(stdin JSON 行 → stdout [4B长度+WAV],READY 协议)
✅ provider:tts-providers.js 新增 cosyvoice(单例进程 + 串行队列,崩溃自动重建)
✅ 验证:cosyvoice-test.py / cosyvoice-opt-test.py / test-cosyvoice-provider.mjs(端到端 47s 含加载)
✅ 女声参考固化:src/voice/models/melo-tts/ref_female.wav(MeloTTS 合成)
待办/注意:
⏳ 打包策略未定:CosyVoice2 模型 3.8G,安装包会从 391MB 涨到 4G+。
建议:打包排除模型(files 里加 !src/voice/models/CosyVoice2-0.5B/**),
首次使用 TTS 时提示手动放置/下载。
⏳ 应用内验证:设置 → TTS 服务商选 CosyVoice2 → 对话测试出声
⏳ jit 实测更慢(RTF 5.9 vs 2.7),已弃用;flow.decoder.estimator onnx(onnxruntime-gpu)未试,是后续提速方向
【2】⏳ 安装新版安装包,验证应用内 TTS
------------------------------------
背景:安装包 TTS 失败根因 = streamMelo 脚本路径没做 asar→unpacked 转换(已修复)。
dist/Bailongma-Setup-2.1.300.exe(03:19,391MB)已修复该 bug,但用户尚未重装验证。
待办:退出旧版 → 覆盖安装 03:19 的包 → 设置里 TTS 选 MeloTTS 本地 → 测试出声。
安装版配置已改好(%APPDATA%\Bailongma\config.json:ttsProvider=melo)。
【3】⏳ 代码提交(语音集成修复 + CosyVoice2 集成)
--------------------------------------------------
未提交的改动:
- src/voice/tts-providers.js(streamMelo asar 修复 + cosyvoice provider)
- src/voice/tts_cosyvoice.py(CosyVoice2 常驻服务)
- src/voice/cosyvoice-src/(CosyVoice2 源码,~10MB,含 Matcha-TTS)
- src/voice/cosyvoice-test.py / cosyvoice-opt-test.py / test-cosyvoice-provider.mjs
- src/voice/voxcpm-test.py / voxcpm-clone-test.py(可保留或删除)
- 待办.txt 本身
提交后 push gitea(远端:gitea,token 在 .git/config)。
注意:CosyVoice2 模型 3.8G 不入库(src/voice/models/CosyVoice2-0.5B 需加入 .gitignore)
【4】⚠️ 安全事项:config.json 含明文 API 密钥
---------------------------------------------
config.json(豆包/火山密钥)被 git 跟踪过且 Gitea 仓库是 public —— 密钥已公开。
建议:1) 到控制台轮换密钥;2) git rm --cached config.json + gitignore。
历史提交里的密钥无法删除,轮换是根本解。
【5】💡 可选优化(做完 1-3 后再考虑)
------------------------------------
- 语音唤醒词("小克小克",备份 F:\optimized-backup\voice-system 里有现成脚本)
- CosyVoice2 提速:onnxruntime-gpu + flow.decoder.estimator.onnx(未试)
- 安装包默认引擎已设为本地(voice-panel.js 默认 'local')