stepfun-step-audio-r1-1Use StepFun Chat Completions with model step-audio-r1.1 for non-streaming speech turns that can send text with optional local audio input and save the return...
Install via ClawdBot CLI:
clawdbot install praanmichael/stepfun-step-audio-r1-1Grade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://api.stepfun.com`Audited Apr 17, 2026 · audit v1.0
Generated May 7, 2026
Generate natural-sounding voiceovers for videos, podcasts, or e-learning modules using StepFun's step-audio-r1.1 model. Input text prompts and receive high-quality audio files in WAV format.
Enhance a voice-based chatbot by allowing users to upload or record audio clips. The skill processes the audio along with text prompts to generate spoken responses, enabling interactive customer support or personal assistants.
Create audio-based language exercises where learners listen to native-like speech from text prompts. The skill can also take learner's spoken input for pronunciation practice and provide corrective feedback.
Convert written content into speech for visually impaired users. The skill can be integrated into apps or websites to read aloud articles, notifications, or user interfaces in a chosen voice.
Use the dry-run and print-json flags to test and debug API payloads without incurring costs. Ideal for developers building audio-enabled features who need to validate request structure before full integration.
Offer monthly or annual subscription plans for users who need high-volume text-to-speech generation, such as content creators or educators. Provide tiered pricing based on number of requests or audio duration.
Integrate the skill as a microservice exposed via an API marketplace, charging per API call (e.g., per audio generation). Developers consume the API and pay based on usage volume.
License the voice generation capability to enterprises for internal use (e.g., customer service bots, employee training). Customize voice profiles and integration, charging a one-time setup fee plus annual support.
💬 Integration Tip
Set the STEPFUN_API_KEY environment variable or inject it via OpenClaw skill config. Use the --dry-run flag to thoroughly test your request parameters before sending real API calls, and check environment tools like ffmpeg or afconvert are available for audio conversion.
Scored Jun 29, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...