argmax-cliOn-device speech-to-text (Whisper) + text-to-speech (Qwen3-TTS) CLI. Runs on the Apple Neural Engine (ANE), Apple's low power, dedicated ML inference chip. M...
Install via ClawdBot CLI:
clawdbot install ZachNagengast/argmax-cliGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://github.com/argmaxinc/WhisperKitAudited Apr 17, 2026 · audit v1.0
Generated Mar 22, 2026
Agents automatically transcribe customer voice messages from support channels into text for ticketing systems, then generate voice replies in the customer's language using TTS. This enables 24/7 multilingual support without cloud API costs.
Agents process recorded meeting audio files on-device to transcribe discussions, extract action items via LLM, and create voice summaries. Ideal for confidential business meetings where data privacy is critical.
Educators use agents to convert lecture recordings into transcribed notes and generate audio versions of study materials in multiple languages. Supports students with disabilities or language barriers offline.
Medical professionals dictate patient notes via audio, which agents transcribe locally for EHR integration, ensuring HIPAA compliance. TTS can generate patient instructions in preferred languages.
Podcasters automate transcription of episodes for show notes and use TTS to create voiceovers or translations in different languages. Runs entirely offline to avoid subscription fees.
Sell perpetual licenses to businesses for on-premise installation, with support and updates included. Targets industries like healthcare and finance where data sovereignty is mandatory.
Offer a free tier for basic transcription/TTS usage, with premium features like advanced models or API server access. Monetize through subscriptions while keeping core functionality offline.
License the skill to hardware manufacturers (e.g., Apple, IoT devices) for embedding into smart assistants or recording tools. Provides a privacy-focused alternative to cloud services.
💬 Integration Tip
Ensure audio files are saved to accessible temp paths and use the small model for fast agent responses in real-time workflows.
Scored Apr 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...