kai-realtime-voiceReal-time voice streaming via MiniMax WebSocket API. Use for low-latency voice conversations and streaming audio generation.
Install via ClawdBot CLI:
clawdbot install ogdegenblaze/kai-realtime-voiceGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Sends data to undocumented external endpoint (potential exfiltration)
POST → https://api.minimax.io/v1/t2a_v2Calls external URL not in known-safe list
https://api.minimax.io/v1/t2a_v2AI Analysis
The skill's external API call to api.minimax.io is consistent with its stated purpose of real-time voice generation and uses a documented, legitimate service. While it sends user-provided text externally, this is an expected function for a TTS skill, and no hidden instructions, credential harvesting, or obfuscation are evident in the provided definition.
Audited Apr 18, 2026 · audit v1.0
Generated Oct 3, 2026
A SaaS company embeds Kai Realtime Voice into its help widget to let customers speak with an AI support agent in real time. The low-latency WebSocket streaming keeps conversations natural, and responses can be played back instantly or saved for QA review.
Game studios use Kai Realtime Voice to give NPCs dynamic, synthesized dialogue instead of pre-recorded lines. Players hear character voices streamed on the fly, enabling branching conversations and emergent storytelling.
An edtech platform powers conversational practice sessions where learners speak with an AI tutor that responds instantly with natural voice. Streaming audio minimizes awkward pauses and makes drills feel like real conversations.
Developers integrate Kai Realtime Voice to read dynamic page content, notifications, and form guidance aloud for visually impaired users. The streaming approach delivers speech as content loads rather than waiting for full page render.
Content creators stream long-form text through Kai Realtime Voice to rapidly generate voice drafts of podcasts or audiobooks. They can then save the audio to files and edit before final production.
Charge customers per character or per minute of streamed audio processed through the Kai Realtime Voice wrapper. Usage is metered through the underlying MiniMax API key, enabling transparent consumption-based pricing.
License the integration scripts and WebSocket handling as an SDK that product teams embed into their own apps. Customers pay a recurring license fee for access to updates, support, and multi-platform packaging.
Offer a fully hosted conversational agent built on Kai Realtime Voice, including prompt design, latency tuning, and voice selection. Clients pay a subscription for the managed service rather than building it themselves.
💬 Integration Tip
Ensure the MINIMAX_API_KEY is present in the environment and install the Python websockets library before running the shell script. Start with the --test flag to confirm WebSocket connectivity, then use --stream for live audio generation.
Scored Oct 3, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.