กลับไปหน้า Tools

GetNotes Tools

jamiepine/voicebox

Tool นี้คืออะไร

Voicebox คือสตูดิโอเสียง AI แบบโอเพนซอร์สที่ทำงานบนเครื่องของคุณ ช่วยให้คุณโคลนเสียง สร้างคำพูดในหลายภาษาและหลายเอนจิน รวมถึงใช้การป้อนเสียงและเอาต์พุตเสียงสำหรับแอปและเอเจนต์ AI ได้อย่างเป็นส่วนตัวและมีประสิทธิภาพ

ข้อมูลโปรเจกต์

ดาว

43.8K

Forks

5.3K

License

MIT

อัปเดต GitHub ล่าสุด

13 ก.ค. 2569

เพิ่มใน GetNotes

21 ก.ค. 2569

Repository

jamiepine/voicebox

เหมาะกับอาชีพ

Ecosystem

TypeScript

แปลและเรียบเรียงโดย AI

เนื้อหาฉบับภาษาไทย

ใช้อ่านเพื่อทำความเข้าใจเบื้องต้น โปรดตรวจสอบรายละเอียดสำคัญกับเอกสารต้นฉบับด้านล่าง

DownloadsReleaseStarsLicenseAsk DeepWiki
jamiepine%2Fvoicebox | Trendshift

Voicebox คืออะไร?

Voicebox คือ สตูดิโอเสียง AI ที่ทำงานบนเครื่องของคุณเป็นหลัก — เป็นทางเลือกโอเพนซอร์สฟรีสำหรับ ElevenLabs และ WisprFlow ในแอปเดียว โคลนเสียงจากไฟล์เสียงเพียงไม่กี่วินาที สร้างคำพูดใน 23 ภาษาผ่านเอนจิน TTS 7 ตัว ป้อนข้อความลงในช่องข้อความใดก็ได้ด้วยปุ่มลัดสากล และให้เอเจนต์ AI ที่รองรับ MCP มีเสียงที่คุณเลือก

ผู้ให้บริการคลาวด์สองรายที่ครองตลาดอยู่ในปัจจุบันนั้นอยู่คนละฝั่งของวงจร Voice I/O — ElevenLabs อยู่ที่เอาต์พุต และ WisprFlow อยู่ที่อินพุต Voicebox ทำได้ทั้งสองอย่าง เชื่อมโยงเข้าด้วยกันด้วย LLM ในเครื่องที่มาพร้อมกับแอปสำหรับการปรับแต่งและบุคลิกเฉพาะโปรไฟล์ และรันทั้งหมดนี้บนเครื่องของคุณ

  • ความเป็นส่วนตัวสมบูรณ์ — โมเดล ข้อมูลเสียง และการบันทึกจะไม่มีวันออกจากเครื่องของคุณ
  • เอนจิน TTS 7 ตัว — Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, HumeAI TADA และ Kokoro
  • การโคลนเสียงและเสียงที่ตั้งไว้ล่วงหน้า — การโคลนแบบ zero-shot จากตัวอย่างอ้างอิง หรือเสียงที่ตั้งไว้ล่วงหน้ากว่า 50 เสียงที่คัดสรรมาอย่างดีผ่าน Kokoro และ Qwen CustomVoice
  • 23 ภาษา — ตั้งแต่ภาษาอังกฤษไปจนถึงภาษาอาหรับ ญี่ปุ่น ฮินดี สวาฮีลี และอื่นๆ
  • เอฟเฟกต์หลังการประมวลผล — การปรับระดับเสียง (pitch shift), รีเวิร์บ (reverb), ดีเลย์ (delay), คอรัส (chorus), คอมเพรสชัน (compression) และฟิลเตอร์ (filters)
  • การพูดที่แสดงอารมณ์ — แท็กทางภาษาศาสตร์ เช่น [laugh], [sigh], [gasp] ผ่าน Chatterbox Turbo; การควบคุมการส่งเสียงด้วยภาษาธรรมชาติผ่าน Qwen CustomVoice
  • ความยาวไม่จำกัด — การแบ่งส่วนอัตโนมัติพร้อมการครอสเฟดสำหรับสคริปต์ บทความ และบทต่างๆ
  • โปรแกรมแก้ไขเรื่องราว — ไทม์ไลน์แบบหลายแทร็กสำหรับการสนทนา พอดแคสต์ และเรื่องเล่า
  • การป้อนเสียง — ปุ่มลัดการป้อนข้อความสากลพร้อมโหมดกดเพื่อพูดและสลับ, การวางอัตโนมัติที่ได้รับการยืนยันการเข้าถึงบน macOS, ไมโครโฟนในแอปบนทุกช่องข้อความ, STT ที่ใช้ Whisper
  • เอาต์พุตเสียงของเอเจนต์ — การเรียกใช้เครื่องมือเพียงครั้งเดียว (voicebox.speak) และเอเจนต์ที่รองรับ MCP (Claude Code, Cursor, Cline) จะพูดกับคุณด้วยเสียงที่คุณโคลนไว้
  • บุคลิกเสียง — แนบบุคลิกแบบอิสระเข้ากับโปรไฟล์เสียงใดก็ได้ จากนั้น Compose, Rewrite หรือ Respond ผ่าน LLM ในเครื่องที่มาพร้อมกับแอป — เอเจนต์สามารถเรียกใช้โหมดเดียวกันผ่าน MCP ได้
  • API-first — REST API พร้อมเซิร์ฟเวอร์ MCP ในตัวสำหรับการรวม Voice I/O เข้ากับแอปและเอเจนต์ของคุณเอง
  • ประสิทธิภาพแบบเนทีฟ — สร้างด้วย Tauri (Rust) ไม่ใช่ Electron
  • ทำงานได้ทุกที่ — macOS (MLX/Metal), Windows (CUDA), Linux, AMD ROCm, Intel Arc, Docker

ดาวน์โหลด

แพลตฟอร์มดาวน์โหลด
macOS (Apple Silicon)ดาวน์โหลด DMG
macOS (Intel)ดาวน์โหลด DMG
Windowsดาวน์โหลด MSI
Dockerdocker compose up

Linux — ไบนารีที่สร้างไว้ล่วงหน้ายังไม่พร้อมใช้งาน ดู voicebox.sh/linux-install สำหรับคำแนะนำในการสร้างจากซอร์สโค้ด

มีปัญหาใช่ไหม? ดู คู่มือการแก้ไขปัญหา สำหรับปัญหาการติดตั้ง การสร้าง โมเดลดาวน์โหลด และ GPU ที่พบบ่อย


คุณสมบัติ

การโคลนเสียงแบบหลายเอนจิน

เอนจิน TTS เจ็ดตัวที่มีจุดแข็งแตกต่างกัน สามารถสลับได้ต่อการสร้างแต่ละครั้ง:

เอนจินภาษาจุดแข็ง
Qwen3-TTS (0.6B / 1.7B)10การโคลนหลายภาษาคุณภาพสูง, คำสั่งการส่งเสียง ("พูดช้าๆ", "กระซิบ")
Qwen CustomVoice10เสียงที่ตั้งไว้ล่วงหน้า 9 เสียงที่คัดสรรมาอย่างดีพร้อมการควบคุมการส่งเสียงด้วยภาษาธรรมชาติ — ไม่ต้องใช้ไฟล์เสียงอ้างอิง
LuxTTSอังกฤษน้ำหนักเบา (~1GB VRAM), เอาต์พุต 48kHz, เร็วกว่าเรียลไทม์ 150 เท่าบน CPU
Chatterbox Multilingual23ครอบคลุมภาษาได้กว้างที่สุด — อาหรับ, เดนมาร์ก, ฟินแลนด์, กรีก, ฮิบรู, ฮินดี, มาเลย์, นอร์เวย์, โปแลนด์, สวาฮีลี, สวีเดน, ตุรกี และอื่นๆ
Chatterbox Turboอังกฤษโมเดล 350M ที่รวดเร็วพร้อมแท็กอารมณ์/เสียงทางภาษาศาสตร์
TADA (1B / 3B)10โมเดลภาษาพูดของ HumeAI — เสียงที่สอดคล้องกันยาวนานกว่า 700 วินาที, การจัดเรียงข้อความ-เสียงแบบคู่
Kokoro8เสียงที่ตั้งไว้ล่วงหน้า 50 เสียงที่คัดสรรมาอย่างดี, โมเดล 82M ขนาดเล็ก, การอนุมานบน CPU ที่รวดเร็ว

อารมณ์และแท็กทางภาษาศาสตร์

เฉพาะ Chatterbox Turbo เท่านั้นที่ตีความแท็กทางภาษาศาสตร์ เช่น [laugh] และ [sigh] Qwen3-TTS, LuxTTS, Chatterbox Multilingual และ HumeAI TADA จะอ่านแท็กเหล่านี้ ตามตัวอักษรเป็นข้อความ

เมื่อเลือก Chatterbox Turbo ให้พิมพ์ / ในช่องป้อนข้อความเพื่อเปิดตัวแทรกแท็ก และเพิ่มแท็กที่แสดงอารมณ์ในบรรทัดเดียวกับคำพูด:

[laugh] [chuckle] [gasp] [cough] [sigh] [groan] [sniff] [shush] [clear throat]

เอฟเฟกต์หลังการประมวลผล

เอฟเฟกต์เสียง 8 แบบที่ขับเคลื่อนโดยไลบรารี pedalboard ของ Spotify ใช้หลังจากสร้างเสียง, ดูตัวอย่างแบบเรียลไทม์, สร้างพรีเซ็ตที่นำกลับมาใช้ใหม่ได้

เอฟเฟกต์คำอธิบาย
Pitch Shiftปรับขึ้นหรือลงได้สูงสุด 12 เซมิโทน
Reverbปรับขนาดห้อง, การหน่วง, อัตราส่วน wet/dry ได้
Delayเสียงสะท้อนพร้อมปรับเวลา, ฟีดแบ็ก และมิกซ์ได้
Chorus / Flangerดีเลย์แบบโมดูเลตสำหรับเสียงโลหะหรือเสียงที่นุ่มนวล
Compressorการบีบอัดช่วงไดนามิก
Gainการปรับระดับเสียง (-40 ถึง +40 dB)
High-Pass Filterลบความถี่ต่ำ
Low-Pass Filterลบความถี่สูง

มาพร้อมกับพรีเซ็ตในตัว 4 แบบ (Robotic, Radio, Echo Chamber, Deep Voice) และรองรับพรีเซ็ตที่กำหนดเอง เอฟเฟกต์สามารถกำหนดเป็นค่าเริ่มต้นต่อโปรไฟล์ได้

ความยาวการสร้างไม่จำกัด

ข้อความจะถูกแบ่งโดยอัตโนมัติที่ขอบเขตประโยค และแต่ละส่วนจะถูกสร้างขึ้นอย่างอิสระ จากนั้นจึงนำมาครอสเฟดเข้าด้วยกัน ใช้งานได้กับทุกเอนจิน

  • กำหนดขีดจำกัดการแบ่งส่วนอัตโนมัติ (100–5,000 ตัวอักษร)
  • แถบเลื่อน Crossfade (0–200ms) สำหรับการเปลี่ยนผ่านที่ราบรื่น
  • ความยาวข้อความสูงสุด: 50,000 ตัวอักษร
  • การแบ่งส่วนอัจฉริยะที่คำนึงถึงตัวย่อ, เครื่องหมายวรรคตอน CJK และ [tags]

เวอร์ชันการสร้าง

การสร้างทุกครั้งรองรับหลายเวอร์ชันพร้อมการติดตามแหล่งที่มา:

  • ต้นฉบับ (Original) — ผลลัพธ์ TTS ที่สะอาดตา เก็บรักษาไว้เสมอ
  • เวอร์ชันเอฟเฟกต์ (Effects versions) — ใช้ชุดเอฟเฟกต์ที่แตกต่างกันจากเวอร์ชันแหล่งที่มาใดก็ได้
  • เทค (Takes) — สร้างใหม่ด้วย seed ใหม่เพื่อความหลากหลาย
  • การติดตามแหล่งที่มา (Source tracking) — แต่ละเวอร์ชันบันทึกสายเลือดของตน
  • รายการโปรด (Favorites) — ติดดาวการสร้างเพื่อการเข้าถึงที่รวดเร็ว

คิวการสร้างแบบอะซิงโครนัส

การสร้างไม่บล็อก ส่งคำสั่งแล้วเริ่มพิมพ์รายการถัดไปได้ทันที

  • คิวการทำงานแบบอนุกรมป้องกันการแย่งชิง GPU
  • การสตรีมสถานะ SSE แบบเรียลไทม์
  • การสร้างที่ล้มเหลวสามารถลองใหม่ได้
  • การสร้างที่ค้างจากการขัดข้องจะกู้คืนอัตโนมัติเมื่อเริ่มต้น

การจัดการโปรไฟล์เสียง

  • สร้างโปรไฟล์จากไฟล์เสียงหรือบันทึกโดยตรงในแอป
  • นำเข้า/ส่งออกโปรไฟล์เพื่อแชร์หรือสำรองข้อมูล
  • รองรับหลายตัวอย่างเพื่อการโคลนคุณภาพสูงขึ้น
  • ชุดเอฟเฟกต์เริ่มต้นต่อโปรไฟล์
  • จัดระเบียบด้วยคำอธิบายและแท็กภาษา

โปรแกรมแก้ไขเรื่องราว (Stories Editor)

โปรแกรมแก้ไขไทม์ไลน์หลายเสียงสำหรับบทสนทนา พอดแคสต์ และเรื่องเล่า

  • การจัดองค์ประกอบหลายแทร็กด้วยการลากและวาง
  • การตัดและแบ่งเสียงแบบอินไลน์
  • การเล่นอัตโนมัติพร้อม playhead ที่ซิงโครไนซ์
  • การปักหมุดเวอร์ชันต่อคลิปแทร็ก

การป้อนตามคำบอกและการป้อนเสียงทั่วโลก

อีกครึ่งหนึ่งของวงจร I/O เสียง กดปุ่มลัดค้างไว้ที่ใดก็ได้ในระบบของคุณ พูด แล้วปล่อย — บน macOS ข้อความที่ถอดเสียงจะถูกวางลงในช่องข้อความที่โฟกัสโดยตรง หรือกดไมโครโฟนบนช่องป้อนข้อความ Voicebox ใดก็ได้แล้วป้อนตามคำบอกโดยตรงในแอป

  • การผูกคอร์ดที่กำหนดค่าได้ (Configurable chord bindings) — คอร์ดกดค้างเพื่อพูดและแตะเพื่อสลับ ซึ่งแต่ละคอร์ดสามารถผูกใหม่ได้ในตัวเลือกคอร์ดในแอป การกดปุ่ม push-to-talk ค้างไว้แล้วแตะ Space ระหว่างที่กดค้างอยู่จะอัปเกรดเป็นการสลับเซสชันโดยไม่มีช่องว่างในเสียง
  • การวางที่รับรู้เป้าหมาย (macOS) (Target-aware paste (macOS)) — การแทรกที่ได้รับการยืนยันการเข้าถึงลงในช่องข้อความที่โฟกัส พร้อมการบันทึก/กู้คืนคลิปบอร์ดแบบอะตอมิก เพื่อไม่ให้คลิปบอร์ดของคุณถูกเขียนทับ
  • UX การอนุญาตครั้งแรก (First-run permissions UX) — ประตูในแอปจะนำคุณผ่านการอนุญาต macOS Accessibility และ Input Monitoring พร้อมลิงก์โดยตรงไปยัง System Settings
  • ปุ่มไมโครโฟนในแอป (In-app mic button) บนช่องข้อความ Voicebox ทุกช่อง — ฟอร์มการสร้าง, คำอธิบายโปรไฟล์, ชื่อเรื่องราว, ทุกที่ที่คุณจะพิมพ์
  • การปรับปรุง LLM (LLM refinement) — การล้างคำว่า "อืม", การติดอ่าง และการเริ่มต้นผิดพลาดก่อนการวาง (ไม่บังคับ)
  • ปุ่มบนหน้าจอ (On-screen pill) — โอเวอร์เลย์ลอยตัวที่แสดงสถานะ recording, transcribing, refining และ speaking ปุ่มเดียวกันนี้ที่เอเจนต์ใช้เมื่อพูดกับคุณ ดังนั้นจึงมีโมเดลความคิดเดียวสำหรับทั้งสองทิศทางของวงจร

การแปลงเสียงเป็นข้อความ (Speech-to-Text)

Voicebox รัน OpenAI Whisper สำหรับการถอดเสียง — โมเดลเดียวกับที่รองรับการป้อนตามคำบอก, แท็บ Captures และ API /transcribe ทำงานบน MLX (Apple Silicon) หรือ PyTorch (CUDA / ROCm / DirectML / CPU) ขึ้นอยู่กับแพลตฟอร์มของคุณ

ขนาดหมายเหตุ
Base / Small / Medium / Largeระดับคุณภาพ Whisper มาตรฐาน
Turboเร็วกว่า Whisper Large ประมาณ 8 เท่า, คุณภาพลดลงน้อยที่สุด

มีแผนจะเพิ่มเอนจินอื่น ๆ (Parakeet v3, Qwen3-ASR) — ดู Roadmap

การบันทึก (Captures)

การป้อนตามคำบอกทุกครั้ง, การบันทึกในแอป และไฟล์เสียงที่อัปโหลดจะไปอยู่ในแท็บ Captures — เสียงต้นฉบับจับคู่กับข้อความที่ถอดเสียง เก็บรักษาไว้เสมอ

  • เล่นซ้ำ, ถอดเสียงใหม่, ปรับปรุง (Replay, re-transcribe, refine) — รัน STT ใหม่ด้วยขนาด Whisper ใดก็ได้ หรือรันข้อความที่ถอดเสียงดิบผ่าน LLM ในเครื่องด้วยแฟล็กที่แตกต่างกัน (การล้างคำฟุ่มเฟือย, การลบการแก้ไขตัวเอง, การเก็บรักษาคำศัพท์ทางเทคนิค)
  • แก้ไขแบบอินไลน์ (Edit inline) — ปรับแต่งข้อความที่ถอดเสียงและบันทึกเมื่อออกจากโฟกัส
  • เล่นเป็นโปรไฟล์เสียง (Play as voice profile) — เปลี่ยนการบันทึกใด ๆ ให้เป็นเสียงด้วยเสียงที่โคลนมา เพียงคลิกเดียว
  • เลื่อนขั้นเป็นตัวอย่างเสียง (Promote to voice sample) — ใช้เสียง + ข้อความที่ถอดเสียงจากการบันทึกเป็นตัวอย่างอ้างอิงในโปรไฟล์เสียงใดก็ได้
  • การจัดเก็บการบันทึกในเครื่อง (Local capture storage) — เสียงต้นฉบับและข้อความที่ถอดเสียงจะอยู่ในไดเรกทอรีข้อมูล Voicebox ของคุณ พร้อมทางลัดโฟลเดอร์ใน Settings

เอาต์พุตเสียงของเอเจนต์

เอเจนต์ทุกตัวมีเสียงของตัวเอง การเรียกใช้เครื่องมือเพียงครั้งเดียวและเอเจนต์ที่รองรับ MCP ใด ๆ ก็สามารถพูดกับคุณด้วยเสียงที่คุณโคลนมาได้ — การทำงานที่เสร็จสมบูรณ์, คำถาม, การแจ้งเตือน ปุ่มเดียวกันที่ปรากฏขึ้นระหว่างการป้อนตามคำบอกจะปรากฏขึ้นระหว่างการพูดของเอเจนต์ ดังนั้นคุณจะเห็นเสมอว่ามีอะไรออกมาจากเครื่องของคุณ

ts
// ในเอเจนต์ที่รองรับ MCP ใด ๆ:
await voicebox.speak({
  text: "Deploy complete.",
  profile: "Morgan",
});

ยังเปิดเผยเป็น POST /speak สำหรับสิ่งใดก็ตามที่ไม่พูด MCP — ACP, A2A, สคริปต์เชลล์, ฮาร์เนสแบบกำหนดเอง

  • ปุ่มสองทิศทาง (Bidirectional pill)recording, transcribing, refining และ speaking ล้วนเป็นสถานะของโอเวอร์เลย์ระดับ OS เดียวกัน ดังนั้นการป้อนตามคำบอกและการพูดของเอเจนต์จึงใช้พื้นผิวเดียวกัน
  • การผูกเสียงต่อเอเจนต์ (Per-agent voice binding) — ใน Settings → MCP ปักหมุด Claude Code ให้ Morgan และ Cursor ให้ Scarlett เพื่อให้คุณสามารถบอกได้ว่าเอเจนต์ตัวไหนกำลังพูดโดยไม่ต้องมอง การประทับเวลา last_seen_at ของไคลเอนต์แต่ละรายยืนยันว่าการติดตั้งสำเร็จจริง
  • มองเห็นได้เสมอ (Always visible) — ไม่มี TTS พื้นหลังแบบเงียบ; การพูดที่เริ่มต้นโดยเอเจนต์ทุกครั้งจะแสดงปุ่มพร้อมชื่อโปรไฟล์เสียงตลอดระยะเวลา
  • การขนส่ง HTTP + stdio (HTTP + stdio transports) — ติดตั้งเป็น URL ใน Claude Code / Cursor / Windsurf / VS Code MCP หรือชี้ไคลเอนต์ที่รองรับ stdio เท่านั้นไปยังไบนารี voicebox-mcp ที่มาพร้อมเครื่อง

บุคลิกเสียง (Voice Personalities)

แนบบุคลิกอิสระเข้ากับโปรไฟล์เสียงใดก็ได้ — ว่าเสียงนี้คือใคร, พูดอย่างไร, สนใจอะไร การกระทำสองอย่างจะปรากฏบนกล่องสร้างเมื่อมีการตั้งค่าบุคลิก ขับเคลื่อนโดย Qwen3 LLM ที่มาพร้อมเครื่องซึ่งทำงานในเครื่องทั้งหมด

  • แต่ง (Compose) — ปุ่มสุ่มที่วางข้อความใหม่ที่ตรงกับบุคลิกในพื้นที่ข้อความ; แก้ไขและพูด หรือคลิกอีกครั้งเพื่อลองแบบอื่น
  • พูดตามบุคลิก (Speak in character) — ปุ่มสลับที่ส่งข้อความที่คุณป้อนผ่าน LLM บุคลิกเพื่อเขียนใหม่ในเสียงของพวกเขา ก่อนที่จะแปลงเป็น TTS

เอเจนต์สามารถเข้าถึงเส้นทางการเขียนใหม่เดียวกันผ่าน MCP โดยส่ง personality: true ไปยัง voicebox.speak ซึ่งจะเปลี่ยนเครื่องมือให้เป็นไปป์ไลน์ text-in → personality-LLM → TTS LLM เดียวกันนี้รองรับขั้นตอนการปรับปรุงของการป้อนตามคำบอก — LLM หนึ่งตัวในแอป, แคชโมเดลหนึ่งตัว, รอยเท้าหน่วยความจำ GPU หนึ่งตัว

ตัวเลือก LLM ในเครื่อง: Qwen3 0.6B / 1.7B / 4B, ใช้รันไทม์ TTS ร่วมกัน (MLX บน Apple Silicon, PyTorch ที่อื่น)

กรณีการใช้งาน: วงจรการพัฒนาเอเจนต์ (ป้อนคำถาม, ฟังคำตอบด้วยเสียงที่โคลนมา), ตัวละครแบบโต้ตอบสำหรับเกมและเครื่องมือเล่าเรื่อง, การช่วยเหลือการพูดสำหรับผู้ที่ไม่สามารถพูดด้วยเสียงต้นฉบับของตนเองได้

การจัดการโมเดล

  • การยกเลิกการโหลดต่อโมเดลเพื่อเพิ่มหน่วยความจำ GPU โดยไม่ต้องลบไฟล์ที่ดาวน์โหลด
  • ไดเรกทอรีโมเดลแบบกำหนดเองผ่าน VOICEBOX_MODELS_DIR
  • การย้ายโฟลเดอร์โมเดลพร้อมการติดตามความคืบหน้า
  • UI สำหรับยกเลิก/ล้างการดาวน์โหลด

การรองรับ GPU

แพลตฟอร์มแบ็กเอนด์หมายเหตุ
macOS (Apple Silicon)MLX (Metal)เร็วขึ้น 4-5 เท่าผ่าน Neural Engine
Windows / Linux (NVIDIA)PyTorch (CUDA)ดาวน์โหลดไบนารี CUDA อัตโนมัติจากภายในแอป
Linux (AMD)PyTorch (ROCm)กำหนดค่า HSA_OVERRIDE_GFX_VERSION อัตโนมัติ
Windows (GPU ใดก็ได้)DirectMLรองรับ GPU Windows สากล
Intel ArcIPEX/XPUการเร่งความเร็ว GPU แยกของ Intel
ใด ๆCPUใช้งานได้ทุกที่ แต่ช้ากว่า

API

Voicebox เปิดเผย REST API สำหรับการรวม I/O เสียงเข้ากับแอปและเอเจนต์ของคุณเอง

bash
# สร้างเสียงพูด
curl -X POST http://127.0.0.1:17493/generate \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello world", "profile_id": "abc123", "language": "en"}'

# เอาต์พุตเสียงของเอเจนต์ — แอปหรือสคริปต์ใด ๆ สามารถพูดด้วยเสียงที่โคลนมาได้
curl -X POST http://127.0.0.1:17493/speak \
  -H "Content-Type: application/json" \
  -H "X-Voicebox-Client-Id: my-script" \
  -d '{"text": "Deploy complete.", "profile": "Morgan"}'

# ถอดเสียงไฟล์เสียง
curl -X POST http://127.0.0.1:17493/transcribe \
  -F "audio=@recording.wav" \
  -F "model=whisper-turbo"

# แสดงรายการโปรไฟล์เสียง
curl http://127.0.0.1:17493/profiles

POST /speak ยอมรับ profile เป็นชื่อ (ไม่คำนึงถึงตัวพิมพ์เล็กใหญ่) หรือ ID และแก้ไขตามลำดับความสำคัญเดียวกับเครื่องมือ MCP: อาร์กิวเมนต์ที่ระบุ → การผูกต่อไคลเอนต์ → capture_settings.default_playback_voice_id

เซิร์ฟเวอร์ MCP

Voicebox มาพร้อมกับเซิร์ฟเวอร์ Model Context Protocol ในตัว เพื่อให้เอเจนต์ที่รองรับ MCP ใด ๆ (Claude Code, Cursor, Windsurf, Cline, ส่วนขยาย VS Code MCP) สามารถพูด, ถอดเสียง และเรียกดูการบันทึกและโปรไฟล์ได้

คำสั่งเดียวสำหรับ Claude Code:

code
claude mcp add voicebox \
  --transport http \
  --url http://127.0.0.1:17493/mcp \
  --header "X-Voicebox-Client-Id: claude-code"

ไคลเอนต์ HTTP MCP ใด ๆ (Cursor, Windsurf, VS Code, ฯลฯ):

json
{
  "mcpServers": {
    "voicebox": {
      "url": "http://127.0.0.1:17493/mcp",
      "headers": { "X-Voicebox-Client-Id": "cursor" }
    }
  }
}

**Stdio fallback** สำหรับไคลเอนต์ที่ไม่รองรับ HTTP MCP — ชี้ไปที่ไบนารี `voicebox-mcp` ที่มาพร้อมกับแอป:

```json
{
  "mcpServers": {
    "voicebox": {
      "command": "/Applications/Voicebox.app/Contents/MacOS/voicebox-mcp",
      "env": { "VOICEBOX_CLIENT_ID": "claude-desktop" }
    }
  }
}

มีเครื่องมือสี่อย่างที่มาพร้อมกับแอป: voicebox.speak, voicebox.transcribe, voicebox.list_captures, voicebox.list_profiles การผูกเสียงต่อไคลเอนต์จะถูกจัดการใน Voicebox → Settings → MCP ดู คู่มือ MCP ฉบับเต็ม สำหรับลายเซ็นเครื่องมือ, ลำดับความสำคัญในการแก้ไข, สัญญาการพูด, และข้อควรระวังด้านความปลอดภัย

ts
// ในเอเจนต์ที่รองรับ MCP:
await voicebox.speak({
  text: "Tests passing. Ready to merge.",
  profile: "Morgan",      // ไม่บังคับ — จะใช้การผูกต่อไคลเอนต์เป็นค่าเริ่มต้น
  personality: true,      // ไม่บังคับ — จะเขียนข้อความใหม่ผ่าน LLM บุคลิกภาพของโปรไฟล์ก่อน
});

กรณีการใช้งาน: วงจรการพัฒนาเอเจนต์ (เสียงเข้า, เสียงออก), บทสนทนาในเกม, การผลิตพอดแคสต์, เครื่องมือช่วยการเข้าถึง, ผู้ช่วยเสียง, การทำงานอัตโนมัติของเนื้อหา

เอกสาร API ฉบับเต็มมีอยู่ที่ http://127.0.0.1:17493/docs


Tech Stack

เลเยอร์เทคโนโลยี
แอปเดสก์ท็อปTauri (Rust)
ส่วนหน้าReact, TypeScript, Tailwind CSS
สถานะZustand, React Query
ส่วนหลังFastAPI (Python)
เอนจิน TTSQwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox, Chatterbox Turbo, TADA, Kokoro
STTWhisper / Whisper Turbo (PyTorch or MLX)
LLM ในเครื่องQwen3 (0.6B / 1.7B / 4B), รันไทม์ร่วมกับ TTS / STT
เซิร์ฟเวอร์ MCPFastMCP ติดตั้งที่ /mcp (Streamable HTTP) + ไบนารี stdio shim ที่มาพร้อม
Native ShimRust (ภายใน Tauri) สำหรับ global hotkey, paste injection, focus introspection
เอฟเฟกต์Pedalboard (Spotify)
การอนุมานMLX (Apple Silicon) / PyTorch (CUDA/ROCm/XPU/CPU)
ฐานข้อมูลSQLite
เสียงWaveSurfer.js, librosa

แผนงาน

ฟีเจอร์คำอธิบาย
การวางอัตโนมัติบน Windows / Linuxความเท่าเทียมในการวางคำบอก — SendInput บน Windows, uinput / AT-SPI บน Linux
การขยายเอนจิน STTParakeet v3 และ Qwen3-ASR เข้าร่วม Whisper — 50+ ภาษา, คุณภาพภาษาที่ไม่ใช่ภาษาอังกฤษดีขึ้น
การกำหนดเส้นทางไปป์ไลน์กำหนดค่าสายโซ่ source → transform → sink ได้พร้อม webhook + MCP sinks และตัวแก้ไขพรีเซ็ต
การถอดเสียงแบบสตรีมมิ่งWebSocket /transcribe/stream สำหรับการถอดเสียงบางส่วนขณะที่คุณพูด
LLM เสียงแบบ End-to-endMoshi, GLM-4-Voice, Qwen2.5 Omni — เสียงเป็นเสียงจริง, ไม่มีข้อความคั่น
การออกแบบเสียงสร้างเสียงใหม่จากคำอธิบายข้อความ
การบันทึกแบบยาวเครื่องบันทึกสองสตรีม (ไมโครโฟน + เสียงระบบ) พร้อมการแปลง LLM สรุป
Platform sinksApple Notes, Obsidian และการผสานรวมอื่นๆ ที่เลือกใช้ได้
สถาปัตยกรรมปลั๊กอินขยายด้วยโมเดล, การแปลง และ sinks ที่กำหนดเอง
แอปคู่หูบนมือถือควบคุม Voicebox จากโทรศัพท์ของคุณ

สำหรับ สถานะทางวิศวกรรมฉบับเต็ม, การจัดลำดับความสำคัญของปัญหาที่เปิดอยู่, และคิวงานที่จัดลำดับความสำคัญแล้ว โปรดดูที่ docs/PROJECT_STATUS.md — เอกสารที่มีการอัปเดตอยู่เสมอที่ติดตามสิ่งที่ได้จัดส่งไปแล้ว, สิ่งที่กำลังดำเนินการอยู่, เอนจิน TTS ที่กำลังประเมิน, และเหตุผลที่เรายอมรับหรือเลื่อนการผสานรวมบางอย่างออกไป


การพัฒนา

ดู CONTRIBUTING.md สำหรับรายละเอียดการตั้งค่าและแนวทางการมีส่วนร่วม

เริ่มต้นอย่างรวดเร็ว

bash
git clone https://github.com/jamiepine/voicebox.git
cd voicebox

just setup   # สร้าง Python venv, ติดตั้ง dependencies ทั้งหมด
just dev     # เริ่มต้น backend + แอปเดสก์ท็อป

ติดตั้ง just: brew install just หรือ cargo install just รัน just --list เพื่อดูคำสั่งทั้งหมด

ข้อกำหนดเบื้องต้น: Bun, Rust, Python 3.11+, Tauri Prerequisites และ Xcode บน macOS

repo นี้มาพร้อมกับไฟล์ .mcp.json ที่กำหนดค่าไว้ล่วงหน้าใน root — การรัน Claude Code ภายใน checkout นี้จะดึงเครื่องมือ Voicebox MCP ขึ้นมาโดยอัตโนมัติเมื่อแอป dev ทำงานอยู่

การสร้างในเครื่อง

bash
just build          # สร้างไบนารีเซิร์ฟเวอร์ CPU + แอป Tauri
just build-local    # (Windows) สร้างไบนารีเซิร์ฟเวอร์ CPU + CUDA + แอป Tauri

การเพิ่มโมเดลเสียงใหม่

สถาปัตยกรรมแบบหลายเอนจินทำให้การเพิ่มเอนจิน TTS ใหม่เป็นเรื่องง่าย คู่มือทีละขั้นตอน ครอบคลุมกระบวนการทั้งหมด: การวิจัย dependencies, การนำโปรโตคอล backend ไปใช้, การเชื่อมต่อ frontend และการรวม PyInstaller

คู่มือนี้ได้รับการปรับให้เหมาะสมสำหรับ AI coding agents ทักษะเอเจนต์ สามารถรับชื่อโมเดลและจัดการการผสานรวมทั้งหมดได้โดยอัตโนมัติ — คุณเพียงแค่ทดสอบการสร้างในเครื่อง

โครงสร้างโปรเจกต์

code
voicebox/
├── app/              # ส่วนหน้า React ที่ใช้ร่วมกัน
├── tauri/            # แอปเดสก์ท็อป (Tauri + Rust)
├── web/              # การปรับใช้บนเว็บ
├── backend/          # เซิร์ฟเวอร์ Python FastAPI
├── landing/          # เว็บไซต์การตลาด
└── scripts/          # สคริปต์การสร้างและเผยแพร่

การมีส่วนร่วม

ยินดีรับการมีส่วนร่วม! ดู CONTRIBUTING.md สำหรับแนวทาง

  1. 1Fork repo
  2. 2สร้าง feature branch
  3. 3ทำการเปลี่ยนแปลงของคุณ
  4. 4ส่ง PR

ความปลอดภัย

พบช่องโหว่ด้านความปลอดภัยหรือไม่? โปรดรายงานอย่างรับผิดชอบ ดู SECURITY.md สำหรับรายละเอียด


ใบอนุญาต

MIT License — ดู LICENSE สำหรับรายละเอียด


เอกสารโปรเจกต์

อ่านเอกสารต้นฉบับ

README วิธีติดตั้ง วิธีใช้งาน และข้อกำหนดจาก repository ต้นฉบับ

ดูไฟล์บน GitHub
Voicebox
DownloadsReleaseStarsLicenseAsk DeepWiki
jamiepine%2Fvoicebox | Trendshift
Voicebox App Screenshot
Voicebox Screenshot 2
Voicebox Screenshot 3

What is Voicebox?

Voicebox is a local-first AI voice studio — a free and open-source alternative to ElevenLabs and WisprFlow in one app. Clone voices from a few seconds of audio, generate speech in 23 languages across 7 TTS engines, dictate into any text field with a global hotkey, and give any MCP-aware AI agent a voice of your choosing.

The two cloud incumbents sit on opposite halves of the voice I/O loop — ElevenLabs on output, WisprFlow on input. Voicebox does both, bridges them with a bundled local LLM for refinement and per-profile personas, and runs the whole thing on your machine.

  • Complete privacy — models, voice data, and captures never leave your machine
  • 7 TTS engines — Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, HumeAI TADA, and Kokoro
  • Voice cloning and preset voices — zero-shot cloning from a reference sample, or 50+ curated preset voices via Kokoro and Qwen CustomVoice
  • 23 languages — from English to Arabic, Japanese, Hindi, Swahili, and more
  • Post-processing effects — pitch shift, reverb, delay, chorus, compression, and filters
  • Expressive speech — paralinguistic tags like [laugh], [sigh], [gasp] via Chatterbox Turbo; natural-language delivery control via Qwen CustomVoice
  • Unlimited length — auto-chunking with crossfade for scripts, articles, and chapters
  • Stories editor — multi-track timeline for conversations, podcasts, and narratives
  • Voice input — global dictation hotkey with push-to-talk and toggle modes, accessibility-verified auto-paste on macOS, in-app mic on every text field, Whisper-based STT
  • Agent voice output — one tool call (voicebox.speak) and any MCP-aware agent (Claude Code, Cursor, Cline) speaks to you in a voice you've cloned
  • Voice personalities — attach a free-form persona to any voice profile, then Compose, Rewrite, or Respond via a bundled local LLM — agents can invoke the same modes over MCP
  • API-first — REST API plus a built-in MCP server for integrating voice I/O into your own apps and agents
  • Native performance — built with Tauri (Rust), not Electron
  • Runs everywhere — macOS (MLX/Metal), Windows (CUDA), Linux, AMD ROCm, Intel Arc, Docker

Download

PlatformDownload
macOS (Apple Silicon)Download DMG
macOS (Intel)Download DMG
WindowsDownload MSI
Dockerdocker compose up

Linux — Pre-built binaries are not yet available. See voicebox.sh/linux-install for build-from-source instructions.

Having trouble? See the Troubleshooting Guide for common install, generation, model-download, and GPU issues.


Features

Multi-Engine Voice Cloning

Seven TTS engines with different strengths, switchable per-generation:

EngineLanguagesStrengths
Qwen3-TTS (0.6B / 1.7B)10High-quality multilingual cloning, delivery instructions ("speak slowly", "whisper")
Qwen CustomVoice109 curated preset voices with natural-language delivery control — no reference audio required
LuxTTSEnglishLightweight (~1GB VRAM), 48kHz output, 150x realtime on CPU
Chatterbox Multilingual23Broadest language coverage — Arabic, Danish, Finnish, Greek, Hebrew, Hindi, Malay, Norwegian, Polish, Swahili, Swedish, Turkish and more
Chatterbox TurboEnglishFast 350M model with paralinguistic emotion/sound tags
TADA (1B / 3B)10HumeAI speech-language model — 700s+ coherent audio, text-acoustic dual alignment
Kokoro850 curated preset voices, tiny 82M model, fast CPU inference

Emotions & Paralinguistic Tags

Only Chatterbox Turbo interprets paralinguistic tags like [laugh] and [sigh]. Qwen3-TTS, LuxTTS, Chatterbox Multilingual, and HumeAI TADA read them literally as text.

With Chatterbox Turbo selected, type / in the text input to open the tag inserter and add expressive tags inline with speech:

[laugh] [chuckle] [gasp] [cough] [sigh] [groan] [sniff] [shush] [clear throat]

Post-Processing Effects

8 audio effects powered by Spotify's pedalboard library. Apply after generation, preview in real time, build reusable presets.

EffectDescription
Pitch ShiftUp or down by up to 12 semitones
ReverbConfigurable room size, damping, wet/dry mix
DelayEcho with adjustable time, feedback, and mix
Chorus / FlangerModulated delay for metallic or lush textures
CompressorDynamic range compression
GainVolume adjustment (-40 to +40 dB)
High-Pass FilterRemove low frequencies
Low-Pass FilterRemove high frequencies

Ships with 4 built-in presets (Robotic, Radio, Echo Chamber, Deep Voice) and supports custom presets. Effects can be assigned per-profile as defaults.

Unlimited Generation Length

Text is automatically split at sentence boundaries and each chunk is generated independently, then crossfaded together. Works with all engines.

  • Configurable auto-chunking limit (100–5,000 chars)
  • Crossfade slider (0–200ms) for smooth transitions
  • Max text length: 50,000 characters
  • Smart splitting respects abbreviations, CJK punctuation, and [tags]

Generation Versions

Every generation supports multiple versions with provenance tracking:

  • Original — clean TTS output, always preserved
  • Effects versions — apply different effects chains from any source version
  • Takes — regenerate with a new seed for variation
  • Source tracking — each version records its lineage
  • Favorites — star generations for quick access

Async Generation Queue

Generation is non-blocking. Submit and immediately start typing the next one.

  • Serial execution queue prevents GPU contention
  • Real-time SSE status streaming
  • Failed generations can be retried
  • Stale generations from crashes auto-recover on startup

Voice Profile Management

  • Create profiles from audio files or record directly in-app
  • Import/export profiles to share or back up
  • Multi-sample support for higher quality cloning
  • Per-profile default effects chains
  • Organize with descriptions and language tags

Stories Editor

Multi-voice timeline editor for conversations, podcasts, and narratives.

  • Multi-track composition with drag-and-drop
  • Inline audio trimming and splitting
  • Auto-playback with synchronized playhead
  • Version pinning per track clip

Global Dictation & Voice Input

The other half of the voice I/O loop. Hold a hotkey anywhere on your system, speak, release — on macOS the transcript pastes straight into the focused text field. Or hit the mic on any Voicebox text input and dictate directly into the app.

  • Configurable chord bindings — hold-to-speak and tap-to-toggle chords, each rebindable in the in-app chord picker. Holding push-to-talk and tapping Space mid-hold upgrades into a toggle session without a gap in audio
  • Target-aware paste (macOS) — accessibility-verified injection into the focused text field, with atomic clipboard save/restore so your clipboard isn't clobbered
  • First-run permissions UX — in-app gates walk you through the macOS Accessibility and Input Monitoring grants with deep-links to System Settings
  • In-app mic button on every Voicebox text field — generation form, profile descriptions, story titles, anywhere you'd type
  • LLM refinement — optional cleanup of ums, stutters, and false starts before paste
  • On-screen pill — floating overlay surfacing recording, transcribing, refining, and speaking states. Same pill agents use when they speak to you, so there's one mental model for both directions of the loop

Speech-to-Text

Voicebox runs OpenAI Whisper for transcription — the same model that backs dictation, the Captures tab, and the /transcribe API. Running on MLX (Apple Silicon) or PyTorch (CUDA / ROCm / DirectML / CPU) depending on your platform.

SizeNotes
Base / Small / Medium / LargeStandard Whisper quality ladder
Turbo~8x faster than Whisper Large, minimal quality loss

More engines (Parakeet v3, Qwen3-ASR) are planned — see Roadmap.

Captures

Every dictation, in-app recording, and uploaded audio file lands in the Captures tab — original audio paired with transcript, always preserved.

  • Replay, re-transcribe, refine — rerun STT with any Whisper size, or re-run the raw transcript through the local LLM with different flags (filler cleanup, self-correction removal, technical-term preservation)
  • Edit inline — tweak the transcript and save on blur
  • Play as voice profile — turn any capture into speech with a cloned voice, one click
  • Promote to voice sample — use a capture's audio + transcript as a reference sample on any voice profile
  • Local capture storage — original audio and transcript stay in your Voicebox data directory, with a folder shortcut in Settings

Agent Voice Output

Every agent gets a voice. One tool call and any MCP-aware agent can speak to you in a voice you've cloned — task completions, questions, notifications. The same pill that surfaces during dictation surfaces during agent speech, so you always see what's coming out of your machine.

ts
// In any MCP-aware agent:
await voicebox.speak({
  text: "Deploy complete.",
  profile: "Morgan",
});

Also exposed as POST /speak for anything that doesn't speak MCP — ACP, A2A, shell scripts, custom harnesses.

  • Bidirectional pillrecording, transcribing, refining, and speaking are all states of the same OS-level overlay, so dictation and agent speech share one surface
  • Per-agent voice binding — in Settings → MCP, pin Claude Code to Morgan and Cursor to Scarlett so you can tell which agent is talking without looking. Each client's last_seen_at timestamp confirms the install actually took
  • Always visible — no silent background TTS; every agent-initiated speak surfaces the pill with the voice profile name for the full duration
  • HTTP + stdio transports — install as a URL in Claude Code / Cursor / Windsurf / VS Code MCP, or point stdio-only clients at the bundled voicebox-mcp binary

Voice Personalities

Attach a free-form personality to any voice profile — who this voice is, how they speak, what they care about. Two actions appear on the generate box when a personality is set, powered by a bundled Qwen3 LLM running entirely locally.

  • Compose — a shuffle button that drops a fresh in-character line into the textarea; edit and speak, or click again for a different take
  • Speak in character — a toggle that routes your input text through the personality LLM to be rewritten in their voice before TTS

Agents can reach the same rewrite path over MCP by passing personality: true to voicebox.speak, turning the tool into a text-in → personality-LLM → TTS pipeline. The same LLM backs dictation's refinement step — one LLM in the app, one model cache, one GPU-memory footprint.

Local LLM options: Qwen3 0.6B / 1.7B / 4B, sharing the TTS runtime (MLX on Apple Silicon, PyTorch elsewhere).

Use cases: agent dev loops (dictate a question, hear the answer in a cloned voice), interactive characters for games and narrative tools, speech assistance for people who can't speak in their original voice.

Model Management

  • Per-model unload to free GPU memory without deleting downloads
  • Custom models directory via VOICEBOX_MODELS_DIR
  • Model folder migration with progress tracking
  • Download cancel/clear UI

GPU Support

PlatformBackendNotes
macOS (Apple Silicon)MLX (Metal)4-5x faster via Neural Engine
Windows (NVIDIA)PyTorch (CUDA)Auto-downloads CUDA binary from within the app
Linux (NVIDIA)PyTorch (CUDA)Use a local/remote Python backend with CUDA PyTorch
Linux (AMD)PyTorch (ROCm)Auto-configures HSA_OVERRIDE_GFX_VERSION
Windows (any GPU)DirectMLUniversal Windows GPU support
Intel ArcIPEX/XPUIntel discrete GPU acceleration
AnyCPUWorks everywhere, just slower

API

Voicebox exposes a REST API for integrating voice I/O into your own apps and agents.

bash
# Generate speech
curl -X POST http://127.0.0.1:17493/generate \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello world", "profile_id": "abc123", "language": "en"}'

# Agent voice output — any app or script can speak in a cloned voice
curl -X POST http://127.0.0.1:17493/speak \
  -H "Content-Type: application/json" \
  -H "X-Voicebox-Client-Id: my-script" \
  -d '{"text": "Deploy complete.", "profile": "Morgan"}'

# Transcribe an audio file
curl -X POST http://127.0.0.1:17493/transcribe \
  -F "audio=@recording.wav" \
  -F "model=whisper-turbo"

# List voice profiles
curl http://127.0.0.1:17493/profiles

POST /speak accepts profile as a name (case-insensitive) or id, and resolves via the same precedence as the MCP tool: explicit arg → per-client binding → capture_settings.default_playback_voice_id.

MCP server

Voicebox ships a built-in Model Context Protocol server so any MCP-aware agent (Claude Code, Cursor, Windsurf, Cline, VS Code MCP extensions) can speak, transcribe, and browse captures and profiles.

Claude Code one-liner:

code
claude mcp add voicebox \
  --transport http \
  --url http://127.0.0.1:17493/mcp \
  --header "X-Voicebox-Client-Id: claude-code"

Any HTTP MCP client (Cursor, Windsurf, VS Code, etc.):

json
{
  "mcpServers": {
    "voicebox": {
      "url": "http://127.0.0.1:17493/mcp",
      "headers": { "X-Voicebox-Client-Id": "cursor" }
    }
  }
}

Stdio fallback for clients that don't speak HTTP MCP — point at the bundled voicebox-mcp binary inside the app:

json
{
  "mcpServers": {
    "voicebox": {
      "command": "/Applications/Voicebox.app/Contents/MacOS/voicebox-mcp",
      "env": { "VOICEBOX_CLIENT_ID": "claude-desktop" }
    }
  }
}

Four tools ship: voicebox.speak, voicebox.transcribe, voicebox.list_captures, voicebox.list_profiles. Per-client voice bindings are managed in Voicebox → Settings → MCP. See the full MCP guide for tool signatures, resolution precedence, the speaking-pill contract, and security notes.

ts
// In any MCP-aware agent:
await voicebox.speak({
  text: "Tests passing. Ready to merge.",
  profile: "Morgan",      // optional — falls back to the per-client binding
  personality: true,      // optional — rewrites text through the profile's personality LLM first
});

Use cases: agent dev loops (voice in, voice out), game dialogue, podcast production, accessibility tools, voice assistants, content automation.

Full API documentation available at http://127.0.0.1:17493/docs.


Tech Stack

LayerTechnology
Desktop AppTauri (Rust)
FrontendReact, TypeScript, Tailwind CSS
StateZustand, React Query
BackendFastAPI (Python)
TTS EnginesQwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox, Chatterbox Turbo, TADA, Kokoro
STTWhisper / Whisper Turbo (PyTorch or MLX)
Local LLMQwen3 (0.6B / 1.7B / 4B), shared runtime with TTS / STT
MCP ServerFastMCP mounted at /mcp (Streamable HTTP) + bundled stdio shim binary
Native ShimRust (inside Tauri) for global hotkey, paste injection, focus introspection
EffectsPedalboard (Spotify)
InferenceMLX (Apple Silicon) / PyTorch (CUDA/ROCm/XPU/CPU)
DatabaseSQLite
AudioWaveSurfer.js, librosa

Roadmap

FeatureDescription
Windows / Linux auto-pasteDictation paste parity — SendInput on Windows, uinput / AT-SPI on Linux
STT engine expansionParakeet v3 and Qwen3-ASR joining Whisper — 50+ languages, better non-English quality
Pipeline routingConfigurable source → transform → sink chains with webhook + MCP sinks and a preset editor
Streaming transcriptionWebSocket /transcribe/stream for partial transcripts as you speak
End-to-end speech LLMsMoshi, GLM-4-Voice, Qwen2.5 Omni — real voice-to-voice, no text between
Voice DesignCreate new voices from text descriptions
Long-form captureDual-stream recorder (mic + system audio) with summary LLM transform
Platform sinksApple Notes, Obsidian, and other opt-in integrations
Plugin architectureExtend with custom models, transforms, and sinks
Mobile companionControl Voicebox from your phone

For the full engineering status, open-issue triage, and prioritized work queue, see docs/PROJECT_STATUS.md — a living document that tracks what's shipped, what's in-flight, candidate TTS engines under evaluation, and why we've accepted or backlogged specific integrations.


Development

See CONTRIBUTING.md for detailed setup and contribution guidelines.

Quick Start

bash
git clone https://github.com/jamiepine/voicebox.git
cd voicebox

just setup   # creates Python venv, installs all deps
just dev     # starts backend + desktop app

Install just: brew install just or cargo install just. Run just --list to see all commands.

Prerequisites: Bun, Rust, Python 3.11+, Tauri Prerequisites, and Xcode on macOS.

The repo ships a pre-wired .mcp.json at the root — running Claude Code inside this checkout picks up the Voicebox MCP tools automatically once the dev app is running.

Building Locally

bash
just build          # Build CPU server binary + Tauri app
just build-local    # (Windows) Build CPU + CUDA server binaries + Tauri app

Adding New Voice Models

The multi-engine architecture makes adding new TTS engines straightforward. A step-by-step guide covers the full process: dependency research, backend protocol implementation, frontend wiring, and PyInstaller bundling.

The guide is optimized for AI coding agents. An agent skill can pick up a model name and handle the entire integration autonomously — you just test the build locally.

Project Structure

code
voicebox/
├── app/              # Shared React frontend
├── tauri/            # Desktop app (Tauri + Rust)
├── web/              # Web deployment
├── backend/          # Python FastAPI server
├── landing/          # Marketing website
└── scripts/          # Build & release scripts

Contributing

Contributions welcome! See CONTRIBUTING.md for guidelines.

  1. 1Fork the repo
  2. 2Create a feature branch
  3. 3Make your changes
  4. 4Submit a PR

Security

Found a security vulnerability? Please report it responsibly. See SECURITY.md for details.


License

MIT License — see LICENSE for details.


#ai#cuda#mlx#qwen3-tts#qwen3-tts-ui#voice-ai#voice-clone#whisper