sensevoice-transcribeTranscribe audio files (WAV/MP3/M4A/FLAC) to timestamped text using SenseVoice-Small + FSMN-VAD. Supports single-file and batch mode with VAD-anchored per-se...
Install via ClawdBot CLI:
clawdbot install ylongw/sensevoice-transcribeGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated Mar 22, 2026
Transcribe university lectures or seminars from recorded audio files, generating timestamped transcripts for students to review key points. Ideal for Mandarin-language courses where accurate Chinese transcription is crucial, and the VAD-anchored timestamps help segment content by topic shifts.
Process interviews or field recordings for news outlets, converting audio into text with timestamps to easily locate quotes and segments. The high accuracy and low hallucination rate ensure reliable transcription for fact-checking and content creation in Mandarin-speaking regions.
Transcribe business meetings or conferences from audio recordings, producing organized transcripts for minutes and action items. The batch mode supports handling multiple files from daily logs, useful for companies with regular Mandarin meetings that require archival and review.
Convert patient-doctor audio consultations into text records with timestamps for medical documentation and follow-up. The model's native Mandarin optimization ensures accurate transcription of specialized terminology, aiding in compliance and record-keeping in healthcare settings.
Assist podcast creators by transcribing episodes to generate show notes, subtitles, or searchable content. The performance efficiency allows quick processing of long recordings, while emotion tags can provide insights into audience engagement for Mandarin-language podcasts.
Offer a cloud-based transcription service with API access, charging users per minute of audio processed or through monthly subscriptions. Target small businesses and individuals needing reliable Mandarin transcription, with premium features like batch processing and Discord integration.
License the skill package to large organizations for internal use, such as universities or media companies, with custom support and integration services. Provide on-premise deployment options for data security, leveraging the model's efficiency and accuracy for high-volume transcription needs.
Develop a free desktop application with basic transcription features, monetized through advanced capabilities like batch mode, progress tracking, and emotion analysis. Attract individual users and upsell to professionals requiring enhanced functionality for Mandarin audio projects.
💬 Integration Tip
Ensure Python 3.12 and the required packages are installed in a virtual environment, and leverage the batch script for automated daylog processing to save time.
Scored Jun 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...