virtual-voice-aiWires a real microphone through an AI brain (STT → LLM → TTS) and routes the output to a virtual audio cable so apps like Google Meet hear the processed voic...
Install via ClawdBot CLI:
clawdbot install suhas12345685-pro/virtual-voice-aiGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Potentially destructive shell commands in tool definitions
exec(Calls external URL not in known-safe list
https://console.deepgram.comUses known external API (expected, informational)
api.anthropic.comAudited Apr 18, 2026 · audit v1.0
Generated Oct 1, 2026
A Node.js developer pipes their microphone through Deepgram STT, an LLM for tone or persona rewriting, and a TTS voice, then sends the resulting PCM into a virtual cable so Google Meet, Zoom, or Discord transmits the AI voice. Useful for streamers, sales reps, and creators who want a consistent synthetic persona without specialized hardware.
Contact centers intercept agent audio, run live transcription and LLM assistance, and synthesize a standardized brand voice back into the same call path. The pipeline enforces script compliance and consistent tone while keeping CRM transcription in sync.
Users with motor or speech disabilities speak short phrases into a microphone, and the system reconstructs fluent, natural-sounding speech through an LLM and TTS, injecting it into video conferencing apps via virtual audio. This provides a real-time assistive layer without requiring platform-specific plugins.
Meeting participants speak in their native language; the pipeline transcribes, translates with an LLM, and speaks the result in a cloned or standard voice into the conference call. Remote teams can hear near-real-time translation through Google Meet or Teams.
Producers route recorded or live mic input through the pipeline to produce synthetic narration in a chosen voice, writing decoded PCM directly to an audio interface for recording. The kill switch allows instant manual takeover during live sessions.
Offer a hosted dashboard where streamers and podcasters configure STT, LLM persona prompts, and TTS voices, then receive a local virtual audio endpoint. Pricing tiers by monthly processing minutes and number of saved voice profiles.
License the pipeline as an SDK/API to contact centers and enterprise collaboration teams, including compliance logging, SSO, and on-prem deployment options. Integration into existing telephony or meeting infrastructure drives seat-based contracts.
Sell metered API access to the pipeline components (device discovery, resampling, sentence chunking, TTS routing) so other developers can build voice apps without managing ffmpeg or virtual cable quirks. Usage-based pricing per minute of processed audio.
💬 Integration Tip
Validate virtual cable installation and ffmpeg PATH setup before writing any code, then run each pipeline stage (device list, capture, STT, chunker, TTS, PCM write) independently and verify audio at each step. Keep the kill switch IPC-based and test it early so the host app can reliably halt streaming during live calls.
Scored Oct 1, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.