กลับไปหน้า Tools

GetNotes Tools

Anil-matcha/Open-Generative-AI

Tool นี้คืออะไร

Open Generative AI เป็นทางเลือกโอเพนซอร์สฟรีสำหรับแพลตฟอร์มวิดีโอ AI ช่วยให้สร้างภาพและวิดีโอ AI โดยใช้โมเดลล้ำสมัยกว่า 400 แบบจาก 14 สตูดิโอ โดยไม่มีข้อจำกัดด้านเนื้อหา ไม่มีระบบปิด และไม่มีค่าสมัครสมาชิก

ข้อมูลโปรเจกต์

ดาว

26.6K

Forks

4.7K

License

MIT

อัปเดต GitHub ล่าสุด

17 ส.ค. 2569

เพิ่มใน GetNotes

17 ส.ค. 2569

Repository

Anil-matcha/Open-Generative-AI

รูปแบบ

Self-hosted

เหมาะกับงาน

AI และ Agents

เหมาะกับอาชีพ

Ecosystem

JavaScript

แปลและเรียบเรียงโดย AI

เนื้อหาฉบับภาษาไทย

ใช้อ่านเพื่อทำความเข้าใจเบื้องต้น โปรดตรวจสอบรายละเอียดสำคัญกับเอกสารต้นฉบับด้านล่าง

Open Generative AI — ทางเลือกโอเพนซอร์สที่ไม่จำกัดสำหรับแพลตฟอร์มวิดีโอ AI

Powered by MuAPI

ทางเลือกโอเพนซอร์สฟรีสำหรับแพลตฟอร์มวิดีโอ AI สร้างภาพและวิดีโอ AI โดยใช้โมเดลล้ำสมัยกว่า 400 แบบจาก 14 สตูดิโอ — ไม่มีตัวกรองเนื้อหา ไม่มีระบบปิด ไม่มีค่าสมัครสมาชิก

ชุมชน: เข้าร่วม Discord สำหรับการพูดคุยและสนับสนุน

Awesome Generative AI Apps

💰 เปลี่ยนสิ่งนี้ให้เป็นผลิตภัณฑ์ของคุณเอง — White Label และ Resell

ต้องการเปิดตัวสิ่งนี้ในฐานะ สตูดิโอ AI แบรนด์ของคุณเอง และเรียกเก็บเงินจากลูกค้าของคุณหรือไม่? MuAPI White Label ช่วยให้คุณสร้างแอปเวอร์ชัน White Label ได้อย่างสมบูรณ์ — โลโก้ของคุณ สีของคุณ โดเมนที่กำหนดเองของคุณ ราคาของคุณเอง — โดยไม่ต้องจัดการโครงสร้างพื้นฐานใดๆ คุณจะได้รับส่วนต่างกำไรจากการสร้างทุกครั้ง; MuAPI จะจัดการโมเดล คิว และระบบการเรียกเก็บเงินที่อยู่เบื้องหลัง

  • แบรนด์ของคุณ — โลโก้ ธีมสี และโดเมนที่กำหนดเอง (เช่น studio.yourbrand.com)
  • ราคาของคุณ — กำหนดราคาเครดิต/การสมัครสมาชิกของคุณเองสำหรับผู้ใช้ปลายทาง เก็บส่วนต่างกำไร
  • ไม่มีโครงสร้างพื้นฐาน — ไม่ต้องรันเซิร์ฟเวอร์, worker, หรือโฮสต์โมเดลด้วยตัวเอง
  • รวมทุกสตูดิโอ — รูปภาพ, วิดีโอ, เสียง, Lip Sync, Cinema, Workflows และอื่นๆ ขึ้นอยู่กับแผน

แผนเริ่มต้นที่ $49/เดือน เริ่มต้นใช้งาน White Label →

สตูดิโอ AI ที่คล้ายกันเรียกเก็บเงินจากผู้ใช้เท่าไร

แพลตฟอร์ม AI สำหรับรูปภาพ/วิดีโอสำหรับผู้บริโภคเกือบทั้งหมดทำงานบนการสมัครสมาชิกรายเดือนแบบชำระเงิน — นี่คือกลยุทธ์เดียวกันกับที่คุณจะใช้ภายใต้แบรนด์ของคุณเอง:

แพลตฟอร์มช่วงราคาการสมัครสมาชิกทั่วไป
Midjourney~$10–$120/เดือน (Basic → Mega)
Runway~$12

🌐 ลองใช้งานออนไลน์ — ไม่ต้องติดตั้ง

เวอร์ชันโฮสต์: https://muapi.ai/open-generative-ai?utm_source=github&utm_medium=readme&utm_campaign=open-generative-ai

ใช้สตูดิโอทั้งหมด (Image, Video, Audio, AI Clipping, Vibe Motion, Lip Sync, Cinema, Marketing, Workflows, Agents, Design Agent, Apps, MCP & CLI) ได้โดยตรงในเบราว์เซอร์ของคุณ — ไม่ต้องใช้ Node.js ไม่ต้องตั้งค่า สมัครบัญชีฟรีเพื่อเริ่มสร้างสรรค์ผลงาน เวอร์ชันโฮสต์จะอัปเดตด้วยโมเดลล่าสุดอยู่เสมอ

ติดตาม ผู้สร้าง เพื่อรับการอัปเดต


⬇️ ดาวน์โหลดแอปเดสก์ท็อป

ตัวติดตั้งแบบคลิกเดียว — ไม่ต้องใช้ Node.js หรือเทอร์มินัล

แพลตฟอร์มดาวน์โหลด
macOS Apple Silicon (M1/M2/M3/M4)Open Generative AI-1.0.9-arm64.dmg
macOS Intel (x64)Open Generative AI-1.0.9.dmg
Windows (x64)Open Generative AI Setup 1.0.9.exe
Linux (Ubuntu x64)v1.0.9 release (.AppImage / .deb), หรือสร้างเองในเครื่องด้วย npm run electron:build:linux

ทุกรีลีส: github.com/Anil-matcha/Open-Generative-AI/releases

คู่มือการติดตั้ง macOS

เนื่องจากแอปไม่ได้ถูกรับรองโดย Apple, macOS Gatekeeper จะบล็อกการเปิดใช้งานครั้งแรก ทำตามขั้นตอนเหล่านี้:

ขั้นตอนที่ 1 — เมานต์ DMG และลากแอปไปที่ /Applications

ขั้นตอนที่ 2 — เปิด Terminal และรัน:

bash
xattr -cr "/Applications/Open Generative AI.app"

ขั้นตอนที่ 3 — คลิกขวาที่แอปใน /Applications → คลิก Open → คลิก Open อีกครั้งในกล่องโต้ตอบ

คุณต้องทำสิ่งนี้เพียงครั้งเดียว หลังจากนั้น แอปจะเปิดตามปกติ

ทางเลือกอื่น (ไม่ต้องใช้ Terminal):

  1. 1ลองเปิดแอป — macOS จะบล็อก
  2. 2ไปที่ System Settings → Privacy & Security
  3. 3เลื่อนลงเพื่อค้นหา "Open Generative AI was blocked"
  4. 4คลิก Open AnywayOpen

การติดตั้ง Windows — แก้ไขคำเตือน SmartScreen

Windows SmartScreen อาจแสดงคำเตือนเนื่องจากตัวติดตั้งไม่ได้ลงนามด้วยโค้ด:

  1. 1คลิก More info ในกล่องโต้ตอบ SmartScreen
  2. 2คลิก Run anyway

แอปจะติดตั้งอย่างเงียบ ๆ ไปยัง %LocalAppData% พร้อมทางลัดใน Start Menu

การติดตั้ง Ubuntu / Linux

ไฟล์ Linux พร้อมใช้งานเมื่อสร้างด้วย Electron Builder:

bash
# Build Linux installers (AppImage + .deb)
npm run electron:build:linux

ไฟล์ที่สร้างขึ้นจะถูกเขียนไปยังโฟลเดอร์ release/:

  • AppImage — พกพาได้, รันได้โดยตรงหลังจากทำให้เป็น executable:chmod +x "release/Open Generative AI-*.AppImage" ./release/Open\ Generative\ AI-*.AppImage
  • .deb — ติดตั้งบน Debian/Ubuntu:sudo apt install ./release/open-generative-ai_*_amd64.deb

หาก AppImage ไม่สามารถเริ่มทำงานบนระบบเก่าได้ ให้ติดตั้ง libfuse2:

bash
sudo apt install libfuse2

Ubuntu 24.04+ / ข้อจำกัด AppArmor sandbox

Ubuntu 24.04 และเวอร์ชันที่ใหม่กว่าจะเปิดใช้งานนโยบายความปลอดภัยของเคอร์เนล (apparmor_restrict_unprivileged_userns) ที่บล็อก user-namespace sandbox ของ Chromium หากแอปไม่สามารถเริ่มทำงานอย่างเงียบ ๆ หรือหยุดทำงานทันที คุณมีสองทางเลือก:

ตัวเลือก A — แนะนำ: ติดตั้ง .deb แทน แพ็กเกจ .deb มาพร้อมกับโปรไฟล์ AppArmor ที่ให้สิทธิ์ที่จำเป็นโดยอัตโนมัติเมื่อติดตั้งโดยไม่มีการเปลี่ยนแปลงทั่วทั้งระบบ

ตัวเลือก B — การแก้ไขระบบชั่วคราว (ผู้ใช้ AppImage):

bash
sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0

สิ่งนี้จะคงอยู่จนกว่าจะรีบูตครั้งถัดไป หากต้องการให้ถาวร:

bash
echo 'kernel.apparmor_restrict_unprivileged_userns=0' | sudo tee /etc/sysctl.d/99-userns.conf

Open Generative AI เป็นสตูดิโอ AI สร้างภาพ วิดีโอ ภาพยนตร์ และลิปซิงค์แบบโอเพนซอร์สฟรี ที่นำเวิร์กโฟลว์สร้างสรรค์มาสู่ทุกคน ไม่มีตัวกรองเนื้อหา ไม่มีการปฏิเสธพรอมต์ ไม่มีข้อจำกัด — เพียงแค่มีอิสระในการสร้างสรรค์อย่างเต็มที่ ขับเคลื่อนโดย Muapi.ai รองรับการสร้างข้อความเป็นภาพ, ภาพเป็นภาพ, ข้อความเป็นวิดีโอ, ภาพเป็นวิดีโอ และการลิปซิงค์ที่ขับเคลื่อนด้วยเสียงผ่านโมเดลต่างๆ เช่น Flux, Nano Banana, Midjourney, Kling, Sora, Veo, Seedream, Infinite Talk, LTX Lipsync, Wan 2.2 และอื่นๆ อีกมากมาย — ทั้งหมดนี้จากอินเทอร์เฟซที่ทันสมัยและใช้งานง่ายที่คุณสามารถโฮสต์เองและปรับแต่งได้

ทำไมต้อง Open Generative AI แทนแพลตฟอร์ม AI Video อื่นๆ?

  • ไม่มีตัวกรอง — ไม่มีตัวกรองเนื้อหา ไม่มีข้อจำกัด ไม่มีพรอมต์ถูกปฏิเสธ
  • ฟรี & โอเพนซอร์ส — ไม่มีค่าสมัคร ไม่มีข้อผูกมัดกับผู้ขาย
  • โฮสต์เอง — ข้อมูลของคุณยังคงอยู่ในเครื่องของคุณ ควบคุมการสร้างสรรค์ได้อย่างเต็มที่
  • 200+ โมเดล — ข้อความเป็นภาพ, ภาพเป็นภาพ, ข้อความเป็นวิดีโอ, ภาพเป็นวิดีโอ, ลิปซิงค์
  • อินพุตหลายภาพ — ป้อนภาพอ้างอิงได้สูงสุด 14 ภาพไปยังโมเดลที่เข้ากันได้
  • Lip Sync Studio — ทำให้ภาพบุคคลเคลื่อนไหวหรือซิงค์ริมฝีปากกับเสียงใดๆ ด้วย 9 โมเดลเฉพาะ
  • ขยายได้ — เพิ่มโมเดลของคุณเอง, ปรับเปลี่ยน UI, สร้างต่อยอดจากมัน

สำหรับข้อมูลเชิงลึกเกี่ยวกับสถาปัตยกรรมทางเทคนิคและปรัชญาเบื้องหลังเวิร์กโฟลว์ภาพยนตร์ "Infinite Budget" โปรดดู คู่มือและแผนงานฉบับสมบูรณ์ ของเรา

⚡ การอนุมานโมเดลในเครื่อง (เฉพาะแอปเดสก์ท็อป)

แอปเดสก์ท็อปรองรับ สองเอนจินในเครื่องที่เป็นอิสระต่อกัน เลือกเอนจินที่เหมาะกับเครื่องที่คุณใช้งานจริง:

เอนจินคืออะไรเหมาะที่สุดสำหรับ
sd.cpp (รวมมาให้)เอนจิน C++ จาก stable-diffusion.cpp รันบนเครื่องเดียวกับแอป Metal GPU บน Apple Silicon, CUDA/Vulkan/ROCm บน Linux/Windowsโมเดลเฉพาะภาพ ทำงานบน Mac M-series
Wan2GP (เซิร์ฟเวอร์ BYO)ไคลเอนต์ HTTP ไปยังเซิร์ฟเวอร์ Wan2GP ที่ผู้ใช้รันเอง เซิร์ฟเวอร์รัน Python + PyTorch บน CUDA/ROCm GPU; แอปเดสก์ท็อปเพียงแค่ส่งพรอมต์และรับผลลัพธ์โมเดลวิดีโอ (Wan 2.2, Hunyuan, LTX) และโมเดลภาพขนาดใหญ่ (Flux, Qwen-Image) ต้องใช้ NVIDIA/AMD GPU บน เซิร์ฟเวอร์; ตัวแอปเดสก์ท็อปสามารถรันบน Mac ได้

ทั้งสองเอนจินใช้ UI เดียวกัน: เปิด Settings → Local Models เพื่อกำหนดค่าแต่ละเอนจิน

เอนจิน 1 — sd.cpp (รวมมาให้)

โมเดลประเภทขนาดหมายเหตุ
Z-Image TurboDiffusion Transformer2.5 GB + 2.7 GB aux8-step turbo ใช้หน่วยความจำมาก
Z-Image BaseDiffusion Transformer3.5 GB + 2.7 GB aux50-step คุณภาพสูง ใช้หน่วยความจำมาก
Dreamshaper 8SD 1.52.1 GB20-step อเนกประสงค์ ตัวเลือกที่เบาที่สุดที่ทดสอบบน Mac
Realistic Vision v5.1SD 1.52.1 GB25-step ภาพสมจริง
Anything v5SD 1.52.1 GB20-step อะนิเมะ/ภาพประกอบ
SDXL Base 1.0SDXL6.9 GB30-step ความละเอียดสูง

โมเดล Z-Image ต้องใช้ไฟล์เสริมที่ใช้ร่วมกันสองไฟล์ (ดาวน์โหลดครั้งเดียว, ใช้ร่วมกันทั้งสองโมเดล):

  • Qwen3-4B Text Encoder — 2.4 GB
  • FLUX VAE — 335 MB

วิธีใช้งาน:

  1. 1เปิด Settings → Local Models ในแอปเดสก์ท็อป
  2. 2ติดตั้ง sd.cpp inference engine (คลิกเดียว — ดาวน์โหลดอัตโนมัติ)
  3. 3ดาวน์โหลดโมเดลที่คุณเลือก (และไฟล์เสริมสำหรับ Z-Image)
  4. 4ใน Image Studio คลิกสวิตช์ ⚡ Local ถัดจากตัวเลือกโมเดล
  5. 5เลือกโมเดลในเครื่องของคุณแล้วสร้าง — ไม่ต้องใช้ API key

การดาวน์โหลดทั้งหมดเกิดขึ้นภายในแอป ไม่มีอะไรถูกติดตั้งทั่วทั้งระบบ

โดยค่าเริ่มต้น sd.cpp จะเก็บเอนจิน, น้ำหนักโมเดล และการดาวน์โหลดชั่วคราวไว้ในไดเรกทอรีข้อมูลแอปของ Electron เส้นทางทั่วไปคือ:

  • macOS: ~/Library/Application Support/open-generative-ai/local-ai
  • Windows: %APPDATA%\open-generative-ai\local-ai
  • Linux: ~/.config/open-generative-ai/local-ai

หากต้องการเก็บน้ำหนักโมเดลหลาย GB ไว้ในไดรฟ์อื่น ให้ตั้งค่า OPEN_GENERATIVE_AI_LOCAL_AI_DIR ก่อนเปิดแอปเดสก์ท็อป แอปจะสร้าง bin/, models/ และ tmp/ ภายในไดเรกทอรีนั้น และ Settings -> Local Models จะแสดงโฟลเดอร์โมเดลที่แก้ไขแล้ว เอาต์พุตเอนจินในเครื่องและข้อผิดพลาดในการดาวน์โหลดจะถูกเขียนไปยังคอนโซลกระบวนการของแอป ดังนั้นให้เปิด จาก Terminal หรือ PowerShell เมื่อคุณต้องการบันทึกการแก้ไขปัญหา

เอนจิน 2 — Wan2GP (เซิร์ฟเวอร์ Gradio ระยะไกล)

แอป ไม่ ได้รวม Python หรือน้ำหนักโมเดลสำหรับ Wan2GP คุณต้องรัน Wan2GP ด้วยตัวเองบนเครื่องที่มี CUDA หรือ ROCm GPU และชี้แอปเดสก์ท็อปไปยัง URL ของมัน

bash
# บนเครื่อง GPU ของคุณ
git clone https://github.com/deepbeepmeep/Wan2GP
cd Wan2GP
./install.sh                          # หรือ install.bat บน Windows
python wgp.py --listen --server-name 0.0.0.0   # ผูกกับทุกอินเทอร์เฟซ

จากนั้นในแอปเดสก์ท็อป: Settings → Local Models → Wan2GP server วาง URL (เช่น http://192.168.1.42:7860) คลิก Test จากนั้น Save โมเดล Wan2GP จะพร้อมใช้งาน — โมเดลภาพใน Image Studio โมเดลวิดีโอสามารถเข้าถึงได้ผ่าน API การสร้างเดียวกัน (Image Studio ปฏิเสธเอาต์พุตวิดีโออย่างชัดเจน; การเชื่อมต่อ Video Studio เต็มรูปแบบอยู่ในแผนงาน)

โมเดลประเภทหมายเหตุ
Flux.1 Devรูปภาพ1024px, 28 ขั้นตอน
Qwen Imageรูปภาพ1024px, 30 ขั้นตอน
Wan 2.2 (T2V / I2V)วิดีโอช้าบน GPU ทั่วไป
Hunyuan VideoวิดีโอT2V คุณภาพสูง
LTX Videoวิดีโอตัวเลือกวิดีโอที่เร็วที่สุด

ทำไมต้องมีเซิร์ฟเวอร์แยกต่างหาก? รันไทม์ของ Wan2GP (Sage attention, flash-attn, AWQ/GGUF kernels) เป็นแบบ CUDA-only — ไม่มีเส้นทางสำหรับ MPS / Apple Silicon การใช้งานในฐานะเซิร์ฟเวอร์ระยะไกลช่วยให้ผู้ใช้ Mac เท่านั้นสามารถเก็บแอปเดสก์ท็อปไว้ได้ ในขณะที่ส่งการประมวลผล inference ไปยังเครื่อง Linux/Windows GPU, พีซีสำหรับเล่นเกมใน LAN หรืออินสแตนซ์ RunPod/vast.ai ที่เช่ามา

การประมวลผล inference แบบโลคอลมีให้ใช้งานเฉพาะในแอปเดสก์ท็อปเท่านั้น เวอร์ชันเว็บที่โฮสต์ไว้จะใช้ Cloud API เสมอ

หมายเหตุเกี่ยวกับฮาร์ดแวร์

  • sd.cpp ทำงานบน CPU (ทุกแพลตฟอร์ม) และ Metal GPU บน Apple Silicon (M1/M2/M3/M4); CUDA/Vulkan/ROCm บน Linux/Windows
  • การเร่งความเร็ว Metal GPU ถูกสร้างมาในไบนารีเดสก์ท็อปของ macOS ซึ่งเร็วกว่า CPU-only อย่างมาก
  • แนะนำสำหรับ sd.cpp Z-Image: RAM 16 GB (น้ำหนักโมเดล 7.4 GB + บัฟเฟอร์คำนวณ 2.4 GB) บน Mac M-series พื้นฐานที่มี RAM 8 GB, Z-Image เป็นที่ทราบกันว่าทำให้ระบบค้างได้ — ให้ใช้ SD 1.5 แทน
  • สำหรับ SD 1.5 บน M2: คาดว่าจะได้ประมาณ 1–2 วินาที/ขั้นตอน เมื่อ Metal dylib ทำงานอยู่ หากคุณเห็นประมาณ 10 วินาที/ขั้นตอนแทน ไบนารีอาจกลับไปใช้ CPU — ดูการตรวจสอบด้านล่าง

การตรวจสอบเส้นทาง SD 1.5 (การทดสอบความถูกต้องที่เร็วที่สุดบน Mac)

หากคุณต้องการยืนยันว่า sd.cpp ติดตั้งอย่างถูกต้องโดยไม่ต้องผ่าน UI คุณสามารถเรียกใช้ sd-cli ได้โดยตรง นี่คือไบนารีเดียวกันกับที่แอปใช้

bash
# 1. App data layout (created on first app launch)
APP_DATA="${OPEN_GENERATIVE_AI_LOCAL_AI_DIR:-$HOME/Library/Application Support/open-generative-ai/local-ai}"
ls "$APP_DATA/bin"     # sd-cli, libstable-diffusion.dylib
ls "$APP_DATA/models"  # whatever you've downloaded

# 2. Grab a small SD 1.5 model directly (Dreamshaper 8, ~2 GB)
curl -L --fail --progress-bar \
  -o "$APP_DATA/models/DreamShaper_8_pruned.safetensors" \
  "https://huggingface.co/Lykon/DreamShaper/resolve/main/DreamShaper_8_pruned.safetensors"

# 3. Run a single 512x512 / 12-step inference
DYLD_LIBRARY_PATH="$APP_DATA/bin" "$APP_DATA/bin/sd-cli" \
  -m "$APP_DATA/models/DreamShaper_8_pruned.safetensors" \
  -p "a serene mountain lake at sunrise, oil painting" \
  -o /tmp/sd15-test.png \
  --steps 12 -H 512 -W 512 --cfg-scale 7.5 --seed 42 \
  --sampling-method euler_a

การรันที่สมบูรณ์บน Apple Silicon จะพิมพ์ total params memory size = 1969.78MB (VRAM 1969.78MB, RAM 0.00MB) (รองรับ Metal) และสร้างไฟล์ PNG ขนาด 512×512 ที่สอดคล้องกัน หาก VRAM เป็น 0.00MB แทน แสดงว่า dylib เป็นแบบ CPU-only — ตรวจสอบ otool -L "$APP_DATA/bin/libstable-diffusion.dylib" | grep -i metal และติดตั้งเอนจินใหม่จาก Settings → Local Models หาก Metal หายไป


✨ คุณสมบัติ

  • Image Studio — สร้างรูปภาพจากข้อความแจ้ง (โมเดล text-to-image กว่า 50+ รายการ) หรือแปลงรูปภาพที่มีอยู่ (โมเดล image-to-image กว่า 55+ รายการ) สลับชุดโมเดลโดยอัตโนมัติตามว่ามีรูปภาพอ้างอิงให้หรือไม่ การควบคุมคุณภาพและความละเอียดจะแสดงสำหรับโมเดลที่รองรับ
  • Local Inference — สองเอนจิน: sd.cpp (รวมมาให้, ทำงานบน Mac/Win/Linux ด้วย Metal/CUDA/Vulkan/ROCm) สำหรับ SD 1.5, SDXL และ Z-Image; และ Wan2GP (เซิร์ฟเวอร์ Gradio ของคุณเอง) สำหรับ Flux, Qwen-Image และโมเดลวิดีโอ (Wan 2.2, Hunyuan, LTX) กำหนดค่าทั้งสองได้ใน Settings → Local Models
  • Multi-Image Input — อัปโหลดรูปภาพอ้างอิงได้สูงสุด 14 รูปสำหรับโมเดลแก้ไขที่เข้ากันได้ (Nano Banana 2 Edit, Flux Kontext Dev, GPT-4o Edit และอื่นๆ) ตัวเลือกแบบหลายรายการพร้อมป้ายลำดับ, การอัปโหลดแบบกลุ่ม และขั้นตอนการยืนยัน "Use Selected"
  • Video Studio — สร้างวิดีโอจากข้อความแจ้ง (โมเดล text-to-video กว่า 40+ รายการ) หรือทำให้รูปภาพเฟรมเริ่มต้นเคลื่อนไหว (โมเดล image-to-video กว่า 60+ รายการ) การสลับโหมดอัจฉริยะแบบเดียวกับ Image Studio
  • Audio Studio — สร้างและแก้ไขเสียง/เพลง AI จากข้อความแจ้ง
  • AI Clipping — ตัดต่อและดึงไฮไลท์จากเนื้อหาวิดีโอที่ยาวขึ้นโดยอัตโนมัติ
  • Vibe Motion Studio — สตูดิโอสร้างการเคลื่อนไหว/แอนิเมชันสำหรับเอฟเฟกต์วิดีโอที่มีสไตล์
  • Lip Sync Studio — ทำให้รูปภาพบุคคลเคลื่อนไหวหรือซิงค์ริมฝีปากในวิดีโอที่มีอยู่โดยใช้เสียง โมเดลเฉพาะ 9 รายการในสองโหมด: รูปภาพบุคคล + เสียง → วิดีโอพูดคุย และวิดีโอ + เสียง → วิดีโอลิปซิงค์
  • Body Swap (Recast) Studio — สลับ/เปลี่ยนร่างกายหรือรูปลักษณ์ของตัวแบบในรูปภาพหรือวิดีโอ
  • Cinema Studio — อินเทอร์เฟซสำหรับการถ่ายภาพยนตร์ที่สมจริงด้วยการควบคุมกล้องระดับมืออาชีพ (Lens, Focal Length, Aperture)
  • Marketing Studio — สร้างโฆษณาและรูปแบบสร้างสรรค์พร้อมสำหรับการตลาดจากอินพุตเดียว
  • Workflow Studio — สร้างและรันไปป์ไลน์ AI แบบหลายขั้นตอนด้วยภาพ เชื่อมโยงโมเดลรูปภาพ วิดีโอ และเสียงเข้าด้วยกันเป็นเวิร์กโฟลว์อัตโนมัติ เรียกดูเทมเพลตจากชุมชน สร้างของคุณเองด้วยตัวแก้ไขแบบโหนด และรันผ่านสนามเด็กเล่นแบบโต้ตอบ
  • Agent Studio — เอเจนต์สร้างสรรค์แบบหลายเทิร์นที่วางแผนและดำเนินการงานสร้างสรรค์แบบสนทนา
  • Design Agent Studio — เอเจนต์ออกแบบอัตโนมัติบน Canvas สำหรับงานภาพแบบวนซ้ำ
  • Explore Apps — ไดเรกทอรีของเทมเพลตแอปและกรณีการใช้งานที่สร้างขึ้นจากแคตตาล็อกโมเดลเดียวกัน
  • AI Influencer Studio — เครื่องมือสำหรับการสร้างและจัดการเนื้อหา AI Persona/Influencer ที่สอดคล้องกัน
  • Upload History — รูปภาพอ้างอิงจะถูกอัปโหลดเพียงครั้งเดียวและจัดเก็บไว้ในเครื่อง แผงตัวเลือกช่วยให้คุณสามารถนำรูปภาพที่อัปโหลดก่อนหน้านี้กลับมาใช้ใหม่ได้ในทุกเซสชัน — ไม่ต้องอัปโหลดซ้ำ
  • Smart Controls — ตัวเลือกอัตราส่วนภาพ ความละเอียด/คุณภาพ และระยะเวลาแบบไดนามิกที่ปรับให้เข้ากับความสามารถของแต่ละโมเดล (รวมถึงโมเดล t2i ที่มีตัวเลือกความละเอียดหรือคุณภาพ)
  • Generation History — เรียกดู เยี่ยมชมซ้ำ และดาวน์โหลดผลลัพธ์ที่สร้างขึ้นทั้งหมด (เก็บไว้ในที่จัดเก็บของเบราว์เซอร์)
  • Image & Video Download — ดาวน์โหลดผลลัพธ์ที่สร้างขึ้นด้วยความละเอียดเต็มรูปแบบได้ในคลิกเดียว
  • API Key Management — การจัดเก็บ API key อย่างปลอดภัยใน browser localStorage (ไม่เคยส่งไปยังเซิร์ฟเวอร์ใดๆ ยกเว้น Muapi)
  • Responsive Design — ทำงานได้อย่างราบรื่นบนเดสก์ท็อปและมือถือด้วย UI แบบ dark glassmorphism

🖼️ Image Studio — โหมดคู่

Image Studio จะสลับระหว่างชุดโมเดลสองชุดโดยอัตโนมัติ:

โหมดตัวกระตุ้นโมเดลข้อความแจ้ง
Text-to-Imageค่าเริ่มต้น (ไม่มีรูปภาพ)โมเดล t2i กว่า 50+ รายการ (Flux, Nano Banana 2, Seedream 5.0, Ideogram, GPT-4o, Midjourney…)จำเป็น
Image-to-Imageอัปโหลดรูปภาพอ้างอิงโมเดล i2i กว่า 55+ รายการ (Kontext, Nano Banana 2 Edit, Seedream 5.0 Edit, Seededit, Upscaler…)ไม่จำเป็น

โมเดลที่เพิ่มใหม่

โมเดลประเภทคุณสมบัติหลัก
Nano Banana 2Text-to-ImageGoogle Gemini 3.1 Flash Image · ความละเอียด 1K/2K/4K · การปรับปรุง Google Search · อัตราส่วนภาพ auto
Nano Banana 2 EditImage-to-Imageรูปภาพอ้างอิงสูงสุด 14 รูป · ความละเอียด 1K/2K/4K · การปรับปรุง Google Search
Seedream 5.0Text-to-ImageByteDance · คุณภาพพื้นฐาน/สูง · 8 อัตราส่วนภาพ · สูงสุด 4K
Seedream 5.0 EditImage-to-ImageByteDance · การถ่ายโอนสไตล์ภาษาธรรมชาติ · คุณภาพพื้นฐาน/สูง
MiniMax Image 01Text-to-ImageMiniMax · 8 อัตราส่วนภาพ · สูงสุด 4 รูปภาพต่อคำขอ · ข้อความแจ้ง 1500 ตัวอักษร

อินพุตหลายรูปภาพ

โมเดลที่รับรูปภาพอ้างอิงหลายรูปภาพจะแสดงตัวเลือกแบบหลายรายการเมื่อเปิดใช้งาน:

โมเดลรูปภาพสูงสุด
Nano Banana 2 Edit14
Nano Banana Edit10
Flux Kontext Dev I2I10
Kling O1 Edit Image10
GPT-4o Edit / GPT Image 1.5 Edit10
Bytedance Seedream Edit v4 / v4.510
Vidu Q2 Reference to Image7
Flux 2 Flex/Pro Edit8
Nano Banana Pro Edit8
Flux Kontext Pro/Max I2I2
Wan 2.5/2.6 Image Edit2–3
Qwen Image Edit Plus / 25113
GPT-4o Image to Image5
Flux 2 Klein 4b/9b Edit4

เมื่อเลือกโมเดลหลายรูปภาพ ตัวกระตุ้นการอัปโหลดจะเปลี่ยนเป็นโหมดเลือกหลายรายการ:

  • ช่องทำเครื่องหมายพร้อมหมายเลขลำดับ — รูปภาพจะถูกส่งไปยังโมเดลตามลำดับที่คุณเลือก
  • การอัปโหลดแบบกลุ่ม — เลือกไฟล์หลายไฟล์พร้อมกันจากกล่องโต้ตอบไฟล์ของคุณ
  • ป้ายนับจำนวน บนตัวกระตุ้นจะแสดงจำนวนรูปภาพที่ใช้งานอยู่; ป้าย + จะปรากฏขึ้นเมื่อมีช่องว่างเพิ่มเติม
  • ปุ่ม "Use Selected" ยืนยันและปิดตัวเลือก

🎬 Video Studio — โหมดคู่

Video Studio เป็นไปตามรูปแบบเดียวกัน:

โหมดตัวกระตุ้นโมเดลข้อความแจ้ง
Text-to-Videoค่าเริ่มต้น (ไม่มีรูปภาพ)โมเดล t2v กว่า 40+ รายการ (Kling, Sora, Veo, Wan, Seedance 2.0, Hailuo, Runway…)จำเป็น
Image-to-Videoอัปโหลดเฟรมเริ่มต้นโมเดล i2v กว่า 60+ รายการ (Kling I2V, Veo3 I2V, Runway I2V, Wan I2V, Seedance 2.0 I2V, Midjourney I2V…)ไม่จำเป็น

โมเดลที่เพิ่มใหม่

โมเดลประเภทคุณสมบัติหลัก
Seedance 2.0Text-to-VideoByteDance · อัตราส่วนภาพ 16:9 / 9:16 / 4:3 / 3:4 · ระยะเวลา 5 / 10 / 15 วินาที · คุณภาพพื้นฐาน/สูง
Seedance 2.0 I2VImage-to-VideoByteDance · ทำให้รูปภาพเคลื่อนไหวเป็นวิดีโอ · รูปภาพอ้างอิงสูงสุด 9 รูป · อัตราส่วนภาพ 16:9 / 9:16 / 4:3 / 3:4 · ระยะเวลา 5 / 10 / 15 วินาที · คุณภาพพื้นฐาน/สูง
Seedance 2.0 Extendส่วนขยายวิดีโอByteDance · ดำเนินการสร้าง Seedance 2.0 ต่อไปได้อย่างราบรื่น · รักษาสไตล์ การเคลื่อนไหว และเสียง · ข้อความแจ้งการดำเนินการต่อเสริม · ระยะเวลา 5 / 10 / 15 วินาที · คุณภาพพื้นฐาน/สูง
Grok Imagine T2VText-to-VideoxAI · ระยะเวลา 6 / 10 / 15 วินาที · โหมด: สนุก / ปกติ / เผ็ดร้อน · อัตราส่วนภาพ 9:16 / 16:9 / 2:3 / 3:2 / 1:1
Grok Imagine I2VImage-to-VideoxAI · ระยะเวลา 6 / 10 / 15 วินาที · โหมด: สนุก / ปกติ / เผ็ดร้อน · การเคลื่อนไหวแบบภาพยนตร์จากภาพนิ่ง
MiniMax Hailuo 02 / 2.3 Standard & ProText-to-Video / Image-to-VideoMiniMax · วิดีโอ Full HD · อัตราส่วนภาพหลายแบบ · มีรุ่นที่รวดเร็วรวมอยู่ด้วย

🎙️ Lip Sync Studio

Lip Sync Studio สร้างวิดีโอพูดที่ขับเคลื่อนด้วยเสียงโดยใช้ 9 โมเดลในสองโหมดอินพุต:

โหมดทริกเกอร์คำอธิบาย
Portrait Imageค่าเริ่มต้นอัปโหลดรูปภาพบุคคล + ไฟล์เสียง → วิดีโอพูดแบบเคลื่อนไหว
Videoสลับไปที่โหมด Videoอัปโหลดวิดีโอที่มีอยู่ + ไฟล์เสียง → วิดีโอลิปซิงค์

โมเดลที่ใช้รูปภาพ (Portrait Image + Audio → Video)

โมเดลEndpointResolutionsPrompt
Infinite Talkinfinitetalk-image-to-video480p, 720pไม่บังคับ
Wan 2.2 Speech to Videowan2.2-speech-to-video480p, 720pไม่บังคับ
LTX 2.3 Lipsyncltx-2.3-lipsync480p, 720p, 1080pไม่บังคับ
LTX 2 19B Lipsyncltx-2-19b-lipsync480p, 720p, 1080pไม่บังคับ

โมเดลที่ใช้วิดีโอ (Video + Audio → Lipsync Video)

โมเดลEndpointResolutionsPrompt
Sync Lipsyncsync-lipsync
LatentSynclatentsync-video
Creatify Lipsynccreatify-lipsync
Veed Lipsyncveed-lipsync
Infinite Talk V2Vinfinitetalk-video-to-video480p, 720pไม่บังคับ

วิธีการทำงาน:

  1. 1เลือกโหมด Portrait Image หรือ Video โดยใช้ปุ่มสลับ
  2. 2อัปโหลดรูปภาพบุคคล (หรือวิดีโอ) ของคุณโดยใช้ปุ่มอัปโหลดรูปภาพ/วิดีโอ
  3. 3อัปโหลดไฟล์เสียงของคุณโดยใช้ปุ่มอัปโหลดเสียง
  4. 4เลือกใส่ prompt เพื่อแนะนำสไตล์การเคลื่อนไหว
  5. 5เลือกรุ่นและความละเอียด (หากรองรับ) จากนั้นคลิก Generate

ประวัติการสร้างจะถูกบันทึกแยกต่างหากใน lipsync_history และงานที่รอดำเนินการจะกลับมาทำงานต่อโดยอัตโนมัติเมื่อโหลดหน้าซ้ำ

🔀 Workflow Studio

Workflow Studio ช่วยให้คุณสร้างและรันไปป์ไลน์ AI แบบหลายขั้นตอนได้โดยไม่ต้องเขียนโค้ด

ความสามารถหลัก:

  • Templates — เริ่มต้นจากเวิร์กโฟลว์ที่สร้างไว้ล่วงหน้า (image chains, video pipelines และอื่นๆ)
  • My Workflows — บันทึกและจัดการไปป์ไลน์ที่คุณกำหนดเอง
  • Community — เรียกดูและรันเวิร์กโฟลว์ที่เผยแพร่โดยผู้ใช้รายอื่น
  • Node-based Builder — ตัวแก้ไขภาพแบบลากและวางเพื่อเชื่อมต่อโมเดลและกำหนดเส้นทางเอาต์พุตระหว่างขั้นตอน
  • Playground — รันเวิร์กโฟลว์ใดๆ แบบโต้ตอบด้วย UI แบบฟอร์ม; ผลลัพธ์จะแสดงผลแบบอินไลน์
  • API execution — ทุกเวิร์กโฟลว์สามารถเรียกใช้ผ่าน Muapi API ได้

💡 ต้องการเพิ่มเวิร์กโฟลว์ลงในแอปของคุณเองหรือไม่? ลองดู Vibe Workflow — เอ็นจิ้นเวิร์กโฟลว์โอเพนซอร์สที่ขับเคลื่อนคุณสมบัตินี้ นำไปใช้ในโปรเจกต์ใดก็ได้

🎥 Cinema Studio Controls

Cinema Studio ให้การควบคุมที่แม่นยำเหนือกล้องเสมือนจริง โดยแปลงตัวเลือกของคุณให้เป็นตัวปรับแต่ง prompt ที่เหมาะสมที่สุด:

หมวดหมู่ตัวเลือกที่มี
CamerasModular 8K Digital, Full-Frame Cine Digital, Grand Format 70mm Film, Studio Digital S35, Classic 16mm Film, Premium Large Format Digital
LensesCreative Tilt, Compact Anamorphic, Extreme Macro, 70s Cinema Prime, Classic Anamorphic, Premium Modern Prime, Warm Cinema Prime, Swirl Bokeh Portrait, Vintage Prime, Halation Diffusion, Clinical Sharp Prime
Focal Lengths8mm (Ultra-Wide), 14mm, 24mm, 35mm (Human Eye), 50mm (Portrait), 85mm (Tight Portrait)
Aperturesf/1.4 (Shallow DoF), f/4 (Balanced), f/11 (Deep Focus)

📁 Upload History & Picker

รูปภาพทุกรูปที่คุณอัปโหลดจะถูกบันทึกไว้ในเครื่อง (URL + รูปย่อ) เพื่อให้คุณไม่ต้องอัปโหลดไฟล์เดิมซ้ำสอง:

  • คลิกปุ่มอัปโหลดเพื่อเปิด reference image picker
  • รูปภาพที่เคยอัปโหลดจะปรากฏในตาราง 3 คอลัมน์พร้อมรูปย่อ
  • Single-image models — คลิกรูปย่อเพื่อเลือกและปิดทันที
  • Multi-image models — สลับรูปย่อหลายรูป (แสดงพร้อมลำดับที่)

เอกสารโปรเจกต์

อ่านเอกสารต้นฉบับ

README วิธีติดตั้ง วิธีใช้งาน และข้อกำหนดจาก repository ต้นฉบับ

ดูไฟล์บน GitHub

Open Generative AI — Unrestricted Open-Source Alternative to AI Video Platforms

Powered by MuAPI

The free, open-source alternative to AI Video Platforms. Generate AI images and videos using 400+ state-of-the-art models across 14 studios — no content filters, no closed ecosystem, no subscription fees.

Community: Join Discord for discussions and support

Awesome Generative AI Apps

💰 Turn This Into Your Own Product — White Label & Resell

Want to launch this as your own branded AI studio and charge your own customers for it? MuAPI White Label lets you spin up a fully white-labeled version of this app — your logo, your colors, your custom domain, your own pricing — with zero infra to manage. You keep the markup on every generation; MuAPI handles the models, the queue, and the billing plumbing underneath.

  • Your branding — logo, color theme, and a custom domain (e.g. studio.yourbrand.com)
  • Your pricing — set your own credit/subscription prices for end users, keep the margin
  • No infra — no servers, workers, or model hosting to run yourself
  • All studios included — Image, Video, Audio, Lip Sync, Cinema, Workflows, and more, depending on plan

Plans start at $49/mo. Get started with White Label →

What similar AI studios charge their users

Consumer AI image/video platforms almost all run on paid monthly subscriptions — this is the same playbook you'd run under your own brand:

PlatformTypical subscription range
Midjourney~$10–$120/mo (Basic → Mega)
Runway~$12–$76/mo (Standard → Unlimited), custom Enterprise
Kling AI~$10–$92/mo across Standard → Premier tiers
Luma Dream Machine~$10–$100+/mo
Pika~$8–$58/mo

(Figures are approximate, general-market ranges and change over time — check each platform's current pricing page before quoting them.) With MuAPI White Label, you set these numbers yourself for your own end users — the subscription revenue is yours.


Related Projects

🎵 MiniMax Music 3.0 prompts & guide: awesome-minimax-music-3-prompts — curated song prompts, lyrics-formatting rules, and a Python client for MiniMax Music 3.0 text-to-music generation.

🎚️ MiniMax Music 3.0 ComfyUI nodes: minimax-music-3-comfyui — native ComfyUI custom nodes for generating full songs or instrumentals with MiniMax Music 3.0.

🎞️ MiniMax H3 API Python SDK: MiniMax-H3-API — Python SDK for MiniMax H3 text-to-video, image-to-video, and first/last-frame video workflows through Muapi.

📝 MiniMax H3 prompt gallery: awesome-minimax-h3-prompts — runnable MuAPI examples and creator-ready prompt references for MiniMax H3 video generation.

🌊 Wan 3.0 API Python SDK: Wan-3.0-API — Python SDK and MCP server for Wan 3.0-compatible text-to-video, image-to-video, multimodal references, uploads, and asynchronous generation jobs.

🆕 FLUX 3 Python SDK: Flux-3-Dev-API — Python wrapper for Black Forest Labs' newly announced FLUX 3 — text-to-image, image-to-image, text-to-video, and image-to-video through one client, including the fast/low-cost Dev variant.

🖼️ Grok Imagine Image 2.0 API: Grok-Imagine-Image-2-API — Python SDK and MCP server for xAI image generation, image editing, and multi-reference workflows through MuAPI.

🎼 MiniMax Music 3.0 Python SDK: minimax-music-3-api — standalone Python SDK for MiniMax Music 3.0 text-to-music generation through MuAPI.

🎬 FLUX 3 video, specifically: flux-3-video-api — a focused Python wrapper for just the FLUX 3 Text-to-Video and Image-to-Video endpoints, with native synchronized audio.

🎓 Learn to monetize generative AI: ai-creator-academy — a free, open-source curriculum teaching creators, freelancers, and agencies how to make money with generative AI image/video/audio tools, covering monetization, pricing, and client delivery — not just how the models work.

🤖 Automate media generations with AI coding agents: Generative-Media-Skills — a library of skills that let agents like Claude Code, Codex, and other coding assistants drive 200+ image/video models end-to-end (prompt → generate → edit → stitch) directly from your terminal. Perfect for building automated media pipelines without touching a UI.

🎬 Seedance 2.5 prompts & API guide: awesome-seedance-2.5-api-prompts — Curated prompt templates, camera control vocabulary, MuAPI reference, and cinematic examples for Seedance 2.5 video generation.

🎥 Seedance 2.5 Python SDK: Seedance-2.5-API — Python wrapper for ByteDance's Seedance 2.5 API — text-to-video, image-to-video, realistic human faces, character consistency.

🧩 Seedance 2.5 ComfyUI pack: seedance2.5-comfyui — native ComfyUI nodes and example workflows for the same MuAPI video model.

🤖 Seedance MCP servers: seedance-2-mcp and seedance-2.5-mcp — focused MCP tools for driving Seedance 2 and Seedance 2.5 from Claude, Cursor, and other AI assistants.

🍌 Claude Fable 5 use cases + 20% off on MuAPI: awesome-claude-fable-5 — 60 curated real-world use cases, prompts, and benchmarks for Claude Fable 5, with 20% off Fable 5 access via MuAPI.

🌐 Try it Online — No Install Required

Hosted version: https://muapi.ai/open-generative-ai?utm_source=github&utm_medium=readme&utm_campaign=open-generative-ai

Use all studios (Image, Video, Audio, AI Clipping, Vibe Motion, Lip Sync, Cinema, Marketing, Workflows, Agents, Design Agent, Apps, MCP & CLI) directly in your browser — no Node.js, no setup. Sign up for a free account to start generating. The hosted version is always up to date with the latest models.

Follow the creator for updates


⬇️ Download Desktop App

One-click installers — no Node.js or terminal required.

PlatformDownload
macOS Apple Silicon (M1/M2/M3/M4)Open Generative AI-1.0.9-arm64.dmg
macOS Intel (x64)Open Generative AI-1.0.9.dmg
Windows (x64)Open Generative AI Setup 1.0.9.exe
Linux (Ubuntu x64)v1.0.9 release (.AppImage / .deb), or build locally with npm run electron:build:linux.

All releases: github.com/Anil-matcha/Open-Generative-AI/releases

macOS Installation Guide

Because the app is not notarized by Apple, macOS Gatekeeper will block it on first launch. Follow these steps:

Step 1 — Mount the DMG and drag the app to /Applications

Step 2 — Open Terminal and run:

bash
xattr -cr "/Applications/Open Generative AI.app"

Step 3 — Right-click the app in /Applications → click Open → click Open again on the dialog

You only need to do this once. After that, the app opens normally.

Alternative (no Terminal):

  1. 1Try to open the app — macOS will block it
  2. 2Go to System Settings → Privacy & Security
  3. 3Scroll down to find "Open Generative AI was blocked"
  4. 4Click Open AnywayOpen

Windows Installation — SmartScreen warning fix

Windows SmartScreen may show a warning because the installer is not code-signed:

  1. 1Click More info on the SmartScreen dialog
  2. 2Click Run anyway

The app will install silently to %LocalAppData% with a Start Menu shortcut.

Ubuntu / Linux Installation

Linux artifacts are available when building with Electron Builder:

bash
# Build Linux installers (AppImage + .deb)
npm run electron:build:linux

Generated files are written to the release/ folder:

  • AppImage — portable, run directly after making executable:chmod +x "release/Open Generative AI-*.AppImage" ./release/Open\ Generative\ AI-*.AppImage
  • .deb — install on Debian/Ubuntu:sudo apt install ./release/open-generative-ai_*_amd64.deb

If AppImage fails to start on older systems, install libfuse2:

bash
sudo apt install libfuse2

Ubuntu 24.04+ / AppArmor sandbox restriction

Ubuntu 24.04 and later enable a kernel security policy (apparmor_restrict_unprivileged_userns) that blocks Chromium's user-namespace sandbox. If the app fails to start silently or crashes immediately, you have two options:

Option A — Recommended: install the .deb instead. The .deb package ships an AppArmor profile that grants the required permission automatically on install with no system-wide changes.

Option B — Temporary system fix (AppImage users):

bash
sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0

This lasts until next reboot. To make it permanent:

bash
echo 'kernel.apparmor_restrict_unprivileged_userns=0' | sudo tee /etc/sysctl.d/99-userns.conf

Open Generative AI is a free, open-source AI image, video, cinema, and lip sync studio that brings creative workflows to everyone. No content filters, no prompt rejections, no guardrails — just full creative freedom. Powered by Muapi.ai, it supports text-to-image, image-to-image, text-to-video, image-to-video, and audio-driven lip sync generation across models like Flux, Nano Banana, Midjourney, Kling, Sora, Veo, Seedream, Infinite Talk, LTX Lipsync, Wan 2.2, and more — all from a sleek, modern interface you can self-host and customize.

Why Open Generative AI instead of other AI Video Platforms?

  • No filters — no content filters, no nanny guardrails, no prompt rejections
  • Free & open-source — no subscription, no vendor lock-in
  • Self-hosted — your data stays on your machine, full creative control
  • 200+ models — text-to-image, image-to-image, text-to-video, image-to-video, lip sync
  • Multi-image input — feed up to 14 reference images into compatible models
  • Lip Sync Studio — animate portraits or sync lips to any audio with 9 dedicated models
  • Extensible — add your own models, modify the UI, build on top of it

For a deep dive into the technical architecture and the philosophy behind the "Infinite Budget" cinema workflow, see our comprehensive guide and roadmap.

⚡ Local Model Inference (Desktop App Only)

The desktop app supports two independent local engines. Pick whichever fits the machine you actually run on:

EngineWhat it isBest for
sd.cpp (bundled)C++ engine from stable-diffusion.cpp, runs on the same machine as the app. Metal GPU on Apple Silicon, CUDA/Vulkan/ROCm on Linux/Windows.Image-only models. Works on Mac M-series.
Wan2GP (BYO server)HTTP client to a user-run Wan2GP server. The server runs Python + PyTorch on a CUDA/ROCm GPU; the desktop app only sends prompts and receives results.Video models (Wan 2.2, Hunyuan, LTX) and large image models (Flux, Qwen-Image). NVIDIA/AMD GPU required on the server; the desktop app itself can run on a Mac.

Both engines share the same UI: open Settings → Local Models to configure each.

Engine 1 — sd.cpp (bundled)

ModelTypeSizeNotes
Z-Image TurboDiffusion Transformer2.5 GB + 2.7 GB aux8-step turbo. Heavy on memory.
Z-Image BaseDiffusion Transformer3.5 GB + 2.7 GB aux50-step high-quality. Heavy on memory.
Dreamshaper 8SD 1.52.1 GB20-step versatile. Lightest tested option on Mac.
Realistic Vision v5.1SD 1.52.1 GB25-step photorealistic
Anything v5SD 1.52.1 GB20-step anime/illustration
SDXL Base 1.0SDXL6.9 GB30-step high-res

Z-Image models require two shared auxiliary files (downloaded once, shared across both models):

  • Qwen3-4B Text Encoder — 2.4 GB
  • FLUX VAE — 335 MB

How to use:

  1. 1Open Settings → Local Models in the desktop app
  2. 2Install the sd.cpp inference engine (one click — auto-downloaded)
  3. 3Download your chosen model (and auxiliary files for Z-Image)
  4. 4In Image Studio, click the ⚡ Local toggle next to the model selector
  5. 5Select your local model and generate — no API key needed

All downloads happen inside the app. Nothing is installed system-wide.

By default, sd.cpp stores the engine, model weights, and temporary downloads under Electron's app data directory. Common paths are:

  • macOS: ~/Library/Application Support/open-generative-ai/local-ai
  • Windows: %APPDATA%\open-generative-ai\local-ai
  • Linux: ~/.config/open-generative-ai/local-ai

To keep multi-GB model weights on another drive, set OPEN_GENERATIVE_AI_LOCAL_AI_DIR before launching the desktop app. The app will create bin/, models/, and tmp/ inside that directory, and Settings -> Local Models shows the resolved model folder. Local engine output and download errors are written to the app process console, so launch from Terminal or PowerShell when you need troubleshooting logs.

Engine 2 — Wan2GP (remote Gradio server)

The app does not bundle Python or model weights for Wan2GP. You run Wan2GP yourself on a machine with a CUDA or ROCm GPU and point the desktop app at its URL.

bash
# On your GPU machine
git clone https://github.com/deepbeepmeep/Wan2GP
cd Wan2GP
./install.sh                          # or install.bat on Windows
python wgp.py --listen --server-name 0.0.0.0   # binds to all interfaces

Then in the desktop app: Settings → Local Models → Wan2GP server, paste the URL (e.g. http://192.168.1.42:7860), click Test, then Save. The Wan2GP models become available — image models in Image Studio, video models reachable via the same generation API (Image Studio rejects video output explicitly; full Video Studio wiring is on the roadmap).

ModelTypeNotes
Flux.1 DevImage1024px, 28 steps
Qwen ImageImage1024px, 30 steps
Wan 2.2 (T2V / I2V)VideoSlow on consumer GPUs
Hunyuan VideoVideoHigh-quality T2V
LTX VideoVideoFastest video option

Why a separate server? Wan2GP's runtime (Sage attention, flash-attn, AWQ/GGUF kernels) is CUDA-only — there is no MPS / Apple Silicon path. Treating it as a remote server lets a Mac-only user keep the desktop app while offloading inference to a Linux/Windows GPU box, a gaming PC on the LAN, or a rented RunPod/vast.ai instance.

Local inference is only available in the desktop app. The hosted web version always uses cloud APIs.

Hardware Notes

  • sd.cpp runs on CPU (all platforms) and Metal GPU on Apple Silicon (M1/M2/M3/M4); CUDA/Vulkan/ROCm on Linux/Windows.
  • Metal GPU acceleration is built into the macOS desktop binary — significantly faster than CPU-only.
  • Recommended for sd.cpp Z-Image: 16 GB RAM (7.4 GB weights + 2.4 GB compute buffer). On a base 8 GB M-series Mac, Z-Image is known to hang the system — stick to SD 1.5 there.
  • For SD 1.5 on M2: expect ~1–2 s/step with the Metal dylib active. If you see ~10 s/step instead, the binary may have fallen back to CPU — see verification below.

Verifying the SD 1.5 path (the fastest sanity test on Mac)

If you want to confirm sd.cpp is installed correctly without going through the UI, you can drive sd-cli directly. This is the same binary the app uses.

bash
# 1. App data layout (created on first app launch)
APP_DATA="${OPEN_GENERATIVE_AI_LOCAL_AI_DIR:-$HOME/Library/Application Support/open-generative-ai/local-ai}"
ls "$APP_DATA/bin"     # sd-cli, libstable-diffusion.dylib
ls "$APP_DATA/models"  # whatever you've downloaded

# 2. Grab a small SD 1.5 model directly (Dreamshaper 8, ~2 GB)
curl -L --fail --progress-bar \
  -o "$APP_DATA/models/DreamShaper_8_pruned.safetensors" \
  "https://huggingface.co/Lykon/DreamShaper/resolve/main/DreamShaper_8_pruned.safetensors"

# 3. Run a single 512x512 / 12-step inference
DYLD_LIBRARY_PATH="$APP_DATA/bin" "$APP_DATA/bin/sd-cli" \
  -m "$APP_DATA/models/DreamShaper_8_pruned.safetensors" \
  -p "a serene mountain lake at sunrise, oil painting" \
  -o /tmp/sd15-test.png \
  --steps 12 -H 512 -W 512 --cfg-scale 7.5 --seed 42 \
  --sampling-method euler_a

A healthy run on Apple Silicon prints total params memory size = 1969.78MB (VRAM 1969.78MB, RAM 0.00MB) (Metal-backed) and produces a coherent 512×512 PNG. If VRAM is 0.00MB instead, the dylib is CPU-only — check otool -L "$APP_DATA/bin/libstable-diffusion.dylib" | grep -i metal and reinstall the engine from Settings → Local Models if Metal is missing.


✨ Features

  • Image Studio — Generate images from text prompts (50+ text-to-image models) or transform existing images (55+ image-to-image models). Switches model set automatically based on whether a reference image is provided. Quality and resolution controls visible for models that support them.
  • Local Inference — Two engines: sd.cpp (bundled, runs on Mac/Win/Linux with Metal/CUDA/Vulkan/ROCm) for SD 1.5, SDXL, and Z-Image; and Wan2GP (BYO Gradio server) for Flux, Qwen-Image, and video models (Wan 2.2, Hunyuan, LTX). Configure both in Settings → Local Models.
  • Multi-Image Input — Upload up to 14 reference images for compatible edit models (Nano Banana 2 Edit, Flux Kontext Dev, GPT-4o Edit, and more). Multi-select picker with order badges, batch upload, and a "Use Selected" confirmation flow.
  • Video Studio — Generate videos from text prompts (40+ text-to-video models) or animate a start-frame image (60+ image-to-video models). Same intelligent mode switching as Image Studio.
  • Audio Studio — Generate and edit AI audio/music from text prompts.
  • AI Clipping — Auto-clip and extract highlights from longer video content.
  • Vibe Motion Studio — Motion/animation generation studio for stylized video effects.
  • Lip Sync Studio — Animate portrait images or sync lips on existing videos using audio. 9 dedicated models across two modes: portrait image + audio → talking video, and video + audio → lipsync video.
  • Body Swap (Recast) Studio — Swap/recast a subject's body or appearance in an image or video.
  • Cinema Studio — Interface for photorealistic cinematic shots with pro camera controls (Lens, Focal Length, Aperture)
  • Marketing Studio — Generate ad and marketing-ready creative variations from a single input.
  • Workflow Studio — Build and run multi-step AI pipelines visually. Chain image, video, and audio models into automated flows. Browse community templates, create your own with a node-based editor, and run them via an interactive playground.
  • Agent Studio — Multi-turn creative agent that plans and executes generation tasks conversationally.
  • Design Agent Studio — Canvas-based autonomous design agent for iterative visual work.
  • Explore Apps — Directory of app templates and use-cases built on the same model catalog.
  • AI Influencer Studio — Tools for creating and managing consistent AI persona/influencer content.
  • Upload History — Reference images are uploaded once and stored locally. A picker panel lets you reuse any previously uploaded image across sessions — no re-uploading.
  • Smart Controls — Dynamic aspect ratio, resolution/quality, and duration pickers that adapt to each model's capabilities (including t2i models with resolution or quality options)
  • Generation History — Browse, revisit, and download all past generations (persisted in browser storage)
  • Image & Video Download — One-click download of generated outputs in full resolution
  • API Key Management — Secure API key storage in browser localStorage (never sent to any server except Muapi)
  • Responsive Design — Works seamlessly on desktop and mobile with dark glassmorphism UI

🖼️ Image Studio — Dual Mode

The Image Studio automatically switches between two model sets:

ModeTriggerModelsPrompt
Text-to-ImageDefault (no image)50+ t2i models (Flux, Nano Banana 2, Seedream 5.0, Ideogram, GPT-4o, Midjourney…)Required
Image-to-ImageReference image uploaded55+ i2i models (Kontext, Nano Banana 2 Edit, Seedream 5.0 Edit, Seededit, Upscaler…)Optional

Newly Added Models

ModelTypeKey Features
Nano Banana 2Text-to-ImageGoogle Gemini 3.1 Flash Image · Resolution 1K/2K/4K · Google Search enhancement · aspect ratio auto
Nano Banana 2 EditImage-to-ImageUp to 14 reference images · Resolution 1K/2K/4K · Google Search enhancement
Seedream 5.0Text-to-ImageByteDance · Quality basic/high · 8 aspect ratios · up to 4K
Seedream 5.0 EditImage-to-ImageByteDance · Natural language style transfer · Quality basic/high
MiniMax Image 01Text-to-ImageMiniMax · 8 aspect ratios · up to 4 images per request · 1500 char prompt

Multi-Image Input

Models that accept multiple reference images expose a multi-select picker when active:

ModelMax Images
Nano Banana 2 Edit14
Nano Banana Edit10
Flux Kontext Dev I2I10
Kling O1 Edit Image10
GPT-4o Edit / GPT Image 1.5 Edit10
Bytedance Seedream Edit v4 / v4.510
Vidu Q2 Reference to Image7
Flux 2 Flex/Pro Edit8
Nano Banana Pro Edit8
Flux Kontext Pro/Max I2I2
Wan 2.5/2.6 Image Edit2–3
Qwen Image Edit Plus / 25113
GPT-4o Image to Image5
Flux 2 Klein 4b/9b Edit4

When a multi-image model is selected the upload trigger switches to multi-select mode:

  • Checkboxes with order numbers — images are sent to the model in the order you select them
  • Batch upload — pick multiple files at once from your file dialog
  • Count badge on the trigger shows how many images are active; a + badge appears when more slots are available
  • "Use Selected" button confirms and closes the picker

🎬 Video Studio — Dual Mode

The Video Studio follows the same pattern:

ModeTriggerModelsPrompt
Text-to-VideoDefault (no image)40+ t2v models (Kling, Sora, Veo, Wan, Seedance 2.0, Hailuo, Runway…)Required
Image-to-VideoStart frame uploaded60+ i2v models (Kling I2V, Veo3 I2V, Runway I2V, Wan I2V, Seedance 2.0 I2V, Midjourney I2V…)Optional

Newly Added Models

ModelTypeKey Features
Seedance 2.0Text-to-VideoByteDance · Aspect ratios 16:9 / 9:16 / 4:3 / 3:4 · Duration 5 / 10 / 15s · Quality basic/high
Seedance 2.0 I2VImage-to-VideoByteDance · Animate images into video · Up to 9 reference images · Aspect ratios 16:9 / 9:16 / 4:3 / 3:4 · Duration 5 / 10 / 15s · Quality basic/high
Seedance 2.0 ExtendVideo ExtensionByteDance · Seamlessly continue any Seedance 2.0 generation · Preserves style, motion & audio · Optional continuation prompt · Duration 5 / 10 / 15s · Quality basic/high
Grok Imagine T2VText-to-VideoxAI · Duration 6 / 10 / 15s · Modes: fun / normal / spicy · Aspect ratios 9:16 / 16:9 / 2:3 / 3:2 / 1:1
Grok Imagine I2VImage-to-VideoxAI · Duration 6 / 10 / 15s · Modes: fun / normal / spicy · Cinematic motion from still images
MiniMax Hailuo 02 / 2.3 Standard & ProText-to-Video / Image-to-VideoMiniMax · Full HD video · Multiple aspect ratios · Fast variant included

🎙️ Lip Sync Studio

The Lip Sync Studio generates audio-driven talking videos using 9 models across two input modes:

ModeTriggerDescription
Portrait ImageDefaultUpload a portrait image + audio file → animated talking video
VideoSwitch to Video modeUpload an existing video + audio file → lipsync video

Image-based Models (Portrait Image + Audio → Video)

ModelEndpointResolutionsPrompt
Infinite Talkinfinitetalk-image-to-video480p, 720pOptional
Wan 2.2 Speech to Videowan2.2-speech-to-video480p, 720pOptional
LTX 2.3 Lipsyncltx-2.3-lipsync480p, 720p, 1080pOptional
LTX 2 19B Lipsyncltx-2-19b-lipsync480p, 720p, 1080pOptional

Video-based Models (Video + Audio → Lipsync Video)

ModelEndpointResolutionsPrompt
Sync Lipsyncsync-lipsync
LatentSynclatentsync-video
Creatify Lipsynccreatify-lipsync
Veed Lipsyncveed-lipsync
Infinite Talk V2Vinfinitetalk-video-to-video480p, 720pOptional

How it works:

  1. 1Select Portrait Image or Video mode using the toggle
  2. 2Upload your portrait image (or video) using the image/video upload button
  3. 3Upload your audio file using the audio upload button
  4. 4Optionally enter a prompt to guide the motion style
  5. 5Select a model and resolution (where supported), then click Generate

Generation history is saved separately in lipsync_history and pending jobs resume automatically on page reload.

🔀 Workflow Studio

The Workflow Studio lets you build and run multi-step AI pipelines without writing code.

Key capabilities:

  • Templates — Start from pre-built workflows (image chains, video pipelines, and more)
  • My Workflows — Save and manage your own custom pipelines
  • Community — Browse and run workflows published by other users
  • Node-based Builder — Drag-and-drop visual editor to connect models and route outputs between steps
  • Playground — Run any workflow interactively with a form UI; results render inline
  • API execution — Every workflow is also callable via the Muapi API

💡 Want to add workflows to your own app? Check out Vibe Workflow — the open-source workflow engine powering this feature. Drop it into any project.

🎥 Cinema Studio Controls

The Cinema Studio offers precise control over the virtual camera, translating your choices into optimized prompt modifiers:

CategoryAvailable Options
CamerasModular 8K Digital, Full-Frame Cine Digital, Grand Format 70mm Film, Studio Digital S35, Classic 16mm Film, Premium Large Format Digital
LensesCreative Tilt, Compact Anamorphic, Extreme Macro, 70s Cinema Prime, Classic Anamorphic, Premium Modern Prime, Warm Cinema Prime, Swirl Bokeh Portrait, Vintage Prime, Halation Diffusion, Clinical Sharp Prime
Focal Lengths8mm (Ultra-Wide), 14mm, 24mm, 35mm (Human Eye), 50mm (Portrait), 85mm (Tight Portrait)
Aperturesf/1.4 (Shallow DoF), f/4 (Balanced), f/11 (Deep Focus)

📁 Upload History & Picker

Every image you upload is saved locally (URL + thumbnail) so you never upload the same file twice:

  • Click the upload button to open the reference image picker
  • Previously uploaded images appear in a 3-column grid with thumbnails
  • Single-image models — click a thumbnail to instantly select and close
  • Multi-image models — toggle multiple thumbnails (shown with order numbers), then click Use Selected
  • Upload new images with the Upload files button (supports multi-file selection in multi-image mode)
  • Remove individual images from history with the ✕ button
  • History persists across browser sessions (stored in localStorage)

🚀 Quick Start

Prerequisites

Setup

Most users want the desktop app, not this dev path. If you just want to run Open Generative AI on your machine, download a prebuilt installer instead — no Node.js required. The instructions below are for contributors building from source.

Pick the entry point that matches your goal:

  • Desktop app (Electron)npm run electron:dev
  • Hosted web version (Next.js)npm run dev
bash
# Clone the repository (with submodules — required for the workflow + agent packages)
git clone --recurse-submodules https://github.com/Anil-matcha/Open-Generative-AI.git
cd Open-Generative-AI

# If you already cloned without --recurse-submodules, run this once:
# git submodule update --init --recursive

# Install dependencies + build workspace packages (studio, workflow, agents).
# This step is REQUIRED — `npm install` alone is not enough; the workspaces
# need to be built before either dev script will work.
npm run setup

# Then start ONE of:
npm run electron:dev   # Desktop app (Electron + Vite) — recommended
npm run dev            # Hosted web version (Next.js) → http://localhost:3000

You'll be prompted to enter your Muapi API key on first use (skip the key if you only plan to use local models).

Troubleshooting — Couldn't find a 'pages' directory: this means Next.js can't see the app/ folder. Confirm you're running npm run dev from the repo root (the directory that contains app/, package.json, and next.config.mjs), and that you cloned with submodules. Re-run npm run setup if packages/Vibe-Workflow or packages/agents are empty.

Production Build

bash
npm run build
npm run start

Desktop App Build

Build native desktop apps with Electron:

bash
# macOS (DMG — Intel + Apple Silicon)
npm run electron:build

# Windows (NSIS installer — x64 + ARM64)
npm run electron:build:win

# Linux (AppImage + DEB — x64)
npm run electron:build:linux

# Both platforms in one pass
npm run electron:build:all

Installers are output to the release/ folder. Pre-built binaries are also available on the Releases page.

🏗️ Architecture

The app is a Next.js monorepo with a shared packages/studio component library.

code
Open-Generative-AI/
├── app/                        # Next.js App Router
│   ├── layout.js               # Root layout (Tailwind, fonts)
│   ├── page.js                 # Redirects → /studio
│   └── studio/
│       └── page.js             # Studio page — renders StandaloneShell
├── components/
│   ├── StandaloneShell.js      # Tab nav + BYOK (API key from localStorage)
│   └── ApiKeyModal.js          # API key entry modal
├── packages/
│   └── studio/                 # Shared React component library
│       └── src/
│           ├── index.js        # Exports: ImageStudio, VideoStudio, AudioStudio, ClippingStudio, VibeMotionStudio, LipSyncStudio, RecastStudio, CinemaStudio, MarketingStudio, WorkflowStudio, AgentStudio, DesignAgentStudio, AppsStudio, AiInfluencerStudio, McpCliStudio
│           ├── models.js       # 400+ model definitions (single source of truth)
│           ├── muapi.js        # API client (named exports, apiKey as first param)
│           └── components/
│               ├── ImageStudio.jsx    # Dual-mode t2i/i2i studio
│               ├── VideoStudio.jsx    # Dual-mode t2v/i2v studio
│               ├── LipSyncStudio.jsx  # Portrait/video + audio → talking video
│               ├── CinemaStudio.jsx   # Pro studio with camera controls
│               └── WorkflowStudio.jsx # Multi-step pipeline builder & playground
├── next.config.mjs             # transpilePackages: ['studio']
├── tailwind.config.js
└── package.json                # workspaces: ["packages/studio"]

The packages/studio library is also consumed by the hosted version on muapi.ai — model updates made in packages/studio/src/models.js apply to both the self-hosted app and the hosted version automatically.

🔌 API Integration

The app communicates with Muapi.ai using a two-step pattern:

  1. 1SubmitPOST /api/v1/{model-endpoint} with prompt and parameters
  2. 2PollGET /api/v1/predictions/{request_id}/result until status is completed

Authentication uses the x-api-key header. During development, a Vite proxy handles CORS by routing /api requests to https://api.muapi.ai.

File uploads use POST /api/v1/upload_file (multipart/form-data) and return a hosted URL that is passed to image-conditioned models. For multi-image models the full images_list array is forwarded to the API in one request.

Lip sync jobs use the same two-step pattern: a dedicated processLipSync() method accepts image_url or video_url alongside audio_url, dispatches to the model's endpoint, and polls until the output video URL is available.

🎨 Supported Model Categories

CategoryCountExamples
Text-to-Image70+Flux Dev, Nano Banana 2, Seedream 5.0, Ideogram v3, Midjourney v7, GPT-4o, SDXL
Image-to-Image70+Nano Banana 2 Edit (×14), Flux Kontext Pro, GPT-4o Edit, Seededit v3, Upscaler, Background Remover
Text-to-Video85+Kling v3, Sora 2, Veo 3, Wan 2.6, Seedance 2.0, Seedance 2.0 Extend, Seedance Pro, Hailuo 2.3, Runway Gen-3
Image-to-Video120+Kling v2.1 I2V, Veo3 I2V, Runway I2V, Seedance 2.0 I2V, Midjourney v7 I2V, Hunyuan I2V, Wan2.2 I2V
Video-to-Video35+Video effects, AI Clipping, Vibe Motion, video-conditioned edits
Lip Sync15Infinite Talk I2V, Wan 2.2 Speech to Video, LTX 2.3 Lipsync, LTX 2 19B Lipsync, Sync, LatentSync, Creatify, Veed, Infinite Talk V2V
Body Swap / Recast3Subject/appearance recast across image and video
Audio15+Text-to-music, remix, and audio editing models

(Counts verified against packages/studio/src/models.js — total 420+ models across these 8 categories, plus additional models surfaced through Marketing, Agent, and Design Agent studios.)

🛠️ Tech Stack

  • Next.js 14 — App Router, server components, fast dev server
  • React 18 — Studio UI components
  • Tailwind CSS v3 — Utility-first styling
  • npm workspaces — Monorepo with shared packages/studio library
  • Muapi.ai — AI model API gateway

🤔 How is this different from other AI Video Platforms?

Open Generative AI is a community-driven, open-source alternative that provides similar creative capabilities without the closed ecosystem:

Other providersOpen Generative AI
CostSubscription-basedFree (open-source)
Content filtersYes — prompts blocked or alteredNone
RestrictionsPlatform guardrails enforcedFull creative freedom
ModelsProprietary400+ open & commercial models
Multi-image inputLimitedUp to 14 images per request
Lip syncNo9 models, image & video modes
Hosted versionSubscriptionFree at muapi.ai/open-generative-ai
Self-hostingNoYes
CustomizableNoFully hackable
Data privacyCloud-basedYour data stays local
Source codeClosedMIT licensed

📄 License

MIT

🙏 Credits

Built with Muapi.ai — the unified API for AI image and video generation models.


Deep Dive: For more details on the "AI Influencer" engine, upcoming "Popcorn" storyboarding features, and the future of this project, read the full technical overview.


Looking for a free, open-source AI Video Platform? Open Generative AI is an open-source AI image and video generation studio — with no content filters that you can self-host, customize, and extend.

#ai-art-generator#ai-image-generation#ai-video-generation#creative-tools#fal-ai-alternative#flux-1#generative-ai#image-to-video#javascript#kling-ai#lipsync#midjourney-alternative