elevenlabs-stt — independently scanned and version-tracked by SaferSkills.
SaferSkills independently audited elevenlabs-stt (Agent Skill) and scored it 100/100 (green). The audit ran 55 deterministic rules across Security, Supply Chain, Maintenance, Transparency, and Community; it found 0 high-severity and 0 lower-severity findings. The full rule-by-rule trace and per-finding evidence are below. Free, methodology-open.
Findings & checks · 0 flagged
Every scanned point with the score it earned and what moved between them.
First recorded scan — no prior version to compare against.
The primary manifest — the file an agent reads to learn what this artifact does.
Install the belt CLI skill: npx skills add belt-sh/cliHigh-accuracy transcription with Scribe models via inference.sh CLI.

Requires inference.sh CLI (belt). Install instructionsbelt login
# Transcribe audio
belt app run elevenlabs/stt --input '{"audio": "https://audio.mp3"}'| Model | ID | Best For |
|---|---|---|
| Scribe v2 | scribe_v2 | Latest, highest accuracy (default) |
| Scribe v1 | scribe_v1 | Stable, proven |
belt app run elevenlabs/stt --input '{"audio": "https://meeting-recording.mp3"}'belt app run elevenlabs/stt --input '{
"audio": "https://meeting.mp3",
"diarize": true
}'Detect laughter, applause, music, and other non-speech events:
belt app run elevenlabs/stt --input '{
"audio": "https://podcast.mp3",
"tag_audio_events": true
}'belt app run elevenlabs/stt --input '{
"audio": "https://spanish-audio.mp3",
"language_code": "spa"
}'belt app run elevenlabs/stt --input '{
"audio": "https://conference.mp3",
"model": "scribe_v2",
"diarize": true,
"tag_audio_events": true,
"language_code": "eng"
}'Get precise word-level and character-level timestamps by aligning known text to audio. Useful for subtitles, lip-sync, and karaoke.
belt app run elevenlabs/forced-alignment --input '{
"audio": "https://narration.mp3",
"text": "This is the exact text spoken in the audio file."
}'{
"words": [
{"text": "This", "start": 0.0, "end": 0.3},
{"text": "is", "start": 0.35, "end": 0.5},
{"text": "the", "start": 0.55, "end": 0.65}
],
"text": "This is the exact text spoken in the audio file."
}# 1. Transcribe video audio
belt app run elevenlabs/stt --input '{
"audio": "https://video.mp4",
"diarize": true
}' > transcript.json
# 2. Use transcript for captions
belt app run infsh/caption-videos --input '{
"video_url": "https://video.mp4",
"captions": "<transcript-from-step-1>"
}'90+ languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, Hindi, Russian, Turkish, Dutch, Swedish, and many more. Leave language_code empty for automatic detection.
# ElevenLabs TTS (reverse direction)
npx skills add inference-sh/skills@elevenlabs-tts
# ElevenLabs dubbing (translate audio)
npx skills add inference-sh/skills@elevenlabs-dubbing
# Other STT models (Whisper)
npx skills add inference-sh/skills@speech-to-text
# Full platform skill (all 250+ apps)
npx skills add inference-sh/skills@infsh-cliBrowse all audio apps: belt app store --category audio
~30 seconds. Free. No account. Every finding cites a rule and a line of evidence.