mlx-audio-serverLocal 24x7 OpenAI-compatible API server for STT/TTS, powered by MLX on your Mac.
Install via ClawdBot CLI:
clawdbot install guoqiao/mlx-audio-serverGrade Good — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://github.com/guoqiao/skills/blob/main/mlx-audio-server/mlx-audio-server/SKAudited Apr 17, 2026 · audit v1.0
Generated Mar 1, 2026
Podcasters and video creators can use this skill to transcribe audio or video files locally on their Mac without relying on cloud services, ensuring privacy and reducing costs. It's ideal for generating subtitles, show notes, or repurposing content into text-based formats like blog posts.
Schools and universities can deploy this on Mac mini servers to provide speech-to-text and text-to-speech services for students with disabilities, such as converting lecture recordings to text or creating audio versions of study materials. It offers a low-cost, on-premise solution that complies with data privacy regulations.
Developers building AI or voice applications can use this skill as a local, OpenAI-compatible API server to test speech recognition and synthesis features without internet dependency. It accelerates prototyping for apps like voice assistants, transcription tools, or interactive media on Apple Silicon devices.
Small teams can run this skill on a shared MacBook to transcribe internal meetings or customer calls locally, keeping sensitive discussions secure and avoiding subscription fees. The output can be used for minutes, action items, or training documentation.
Offer a free version with basic STT/TTS models and charge for premium features like advanced models, higher accuracy, or commercial licensing. This targets developers and small businesses looking for cost-effective, privacy-focused alternatives to cloud APIs.
Partner with Apple resellers to pre-install this skill on Mac mini or MacBook devices sold as dedicated transcription or accessibility workstations. This provides an out-of-the-box solution for industries like education or healthcare, with support and maintenance contracts.
Provide consulting and integration services to large organizations needing tailored STT/TTS solutions, such as integrating with existing workflows or training custom models. This leverages the local, secure nature of the skill for compliance-heavy sectors like finance or legal.
💬 Integration Tip
Ensure ffmpeg and jq are installed via brew for audio processing, and use the provided scripts as examples to integrate STT/TTS into custom applications via the local API server.
Scored Apr 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Secure, offline, OpenAI-compatible local Whisper ASR endpoint for OpenClaw. Features faster-whisper (large-v3-turbo), built-in privacy with no cloud telemetr...
Turn your AI into JARVIS. Voice, wit, and personality — the complete package. Humor cranked to maximum.
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch contro...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。