article-tts拍照或文字转音频:文章照片 OCR 提取文字,或直接接收文字,生成 Microsoft Edge TTS 语音,支持中英文、自动转写、语速调节、逐句拆分。| Capture article photos (OCR) or plain text, generate natural audio via Edge TT...
Install via ClawdBot CLI:
clawdbot install 54meteor/article-ttsGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated May 9, 2026
Convert printed documents or digital text into spoken audio to assist visually impaired users in reading articles, books, or notes. The skill supports OCR for physical documents and direct text input, enabling access to written content through natural speech.
Students learning English or Chinese can convert article text or phrases into audio with adjustable speed and accurate pronunciation. The sentence-splitting feature allows focused listening practice on individual sentences.
Content creators can quickly generate audio versions of written articles or books using TTS, with customizable voices and speeds. The OCR capability enables digitizing physical texts into audio format without manual transcription.
Professionals can photograph whiteboards or notes during meetings and receive instant audio summaries. The skill extracts text via OCR and produces spoken output, saving time on manual note-taking.
Combine with translation tools to convert foreign language articles into TTS audio in English or Chinese. Useful for travelers or researchers needing spoken content from non-native materials.
Offer basic OCR and TTS with limited audio length or voice options for free. Charge per additional characters, premium voices, or batch processing for high-volume users.
License the skill as an SDK for integration into existing apps (e.g., e-readers, news platforms) to provide TTS features. Charge per integration or per active user.
Package the skill as a branded language learning assistant for schools or language centers. Customize voices and vocabulary lists, charge per seat or institutional subscription.
💬 Integration Tip
Ensure Tesseract and language packs are installed per platform. For production, consider caching frequently used TTS voices or pre-generating common phrases to reduce latency.
Scored Jun 29, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...