voice-listener智能唤醒“小龙虾”,启用百度高准确度语音识别,持续监听并自动输入语音内容,支持“停止”暂停输入。
Install via ClawdBot CLI:
clawdbot install fanqing203/voice-listenerGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://vop.baidu.com/server_apiUses known external API (expected, informational)
aip.baidubce.comAudited Apr 17, 2026 · audit v1.0
Generated Mar 21, 2026
This skill enables users with mobility impairments to control their computers via voice commands without needing to manually type or click. After activation with the wake word, it continuously inputs speech into any application, facilitating tasks like writing documents, browsing the web, or sending emails.
Support agents can use this skill to transcribe customer calls in real-time, automatically inputting speech into CRM or ticketing systems. The wake word activation allows for seamless switching between listening and pausing, improving efficiency during multitasking or note-taking.
Language learners can practice pronunciation and speaking by having their speech automatically transcribed into text for review. The continuous listening mode after activation helps in immersive practice sessions, with the stop word allowing easy control over recording intervals.
Professionals in business or legal settings can use this skill to transcribe meetings or interviews hands-free. By activating with the wake word, it captures all speech and inputs it into note-taking apps, reducing manual effort and ensuring accurate records.
Content creators, such as writers or video producers, can dictate scripts or ideas directly into editing software. The skill's high-accuracy recognition and continuous input mode streamline the creative process, allowing for faster ideation and drafting without keyboard interruptions.
Offer the skill for free with limited daily uses, leveraging Baidu's free tier, and charge for premium features like higher usage limits, custom wake words, or advanced analytics. This model attracts users with the free offering and monetizes through upgrades.
License the skill to businesses for internal use, such as in call centers or offices, with tailored configurations and support. Provide volume discounts and integration services to ensure seamless adoption within corporate workflows.
Sell the skill's codebase or API access to developers who can rebrand and integrate it into their own applications. Offer customization options and technical support to help them build voice-enabled features without developing from scratch.
💬 Integration Tip
Ensure proper microphone setup and quiet environment for optimal wake word detection, and regularly update Baidu API credentials to avoid token expiration issues.
Scored Jun 17, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...