talkiesSelf-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 12 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.
Install via ClawdBot CLI:
clawdbot install psyb0t/talkiesGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Sends data to undocumented external endpoint (potential exfiltration)
POST → http://localhost:8000/unloadCalls external URL not in known-safe list
https://github.com/psyb0t/docker-talkiesUses known external API (expected, informational)
api.openai.comAudited May 29, 2026 · audit v1.0
Generated Aug 23, 2026
Record team meetings and upload audio via the API to get accurate transcripts with speaker diarization (L/R channels). Generate SRT subtitles for video archives, enabling searchable meeting minutes and compliance records.
For content creators, podcasters, and video producers, automatically generate subtitles in multiple languages using Canary models for translation. Seamlessly integrate with existing OpenAI-based pipelines for fast turnaround.
Use TTS with Kokoro-82M or Qwen3-TTS to generate natural-sounding responses for IVR systems or chatbots. The plugin's OpenAI-compatible endpoints allow rapid prototyping with minimal code changes.
With explicit user consent, create custom voices for individuals with speech impairments using Qwen3-TTS voice cloning. Enable personalized communication aids that preserve the user's identity.
Stream live audio from events or broadcasts via WebSocket for real-time transcription, enabling live captioning for deaf/hard-of-hearing audiences. The low-latency streaming endpoint supports immediate on-screen text.
Offer speech-to-text and text-to-speech as a managed service with monthly or annual subscriptions. Tiered pricing based on usage limits (minutes processed) and features (e.g., diarization, voice cloning).
Charge per API call or per audio minute processed. Ideal for developers who need occasional access without commitment. Implement metering on the talkies server to track usage and bill accordingly.
License the talkies service as a white-label solution to other companies (e.g., SaaS platforms, call centers) to integrate into their own products. Provide the container and support for customization, charging a licensing fee.
💬 Integration Tip
Ensure your TALKIES_URL is correctly set and the server is secured with a bearer token. Use the provided OpenAI-compatible endpoints to minimize code changes.
Scored Aug 23, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...