Speech AI - Pronunciation, STT & TTS
What this MCP does
Supports English pronunciation assessment, multilingual speech transcription, text-to-speech synthesis, and phoneme reference data for language learning.
Tools
Input schema
{'type': 'object', 'required': ['audio_base64', 'text'], 'properties': {'text': {'type': 'string', 'maxLength': 10000, 'description': 'The reference English text that the speaker was expected to read aloud.'}, 'audio_base64': {'type': 'string', 'maxLength': 20000000, 'description': 'Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats.'}, 'audio_format': {'type': 'string', 'default': 'wav', 'description': "Audio format hint â\x80\x94 one of 'wav', 'mp3', 'ogg', 'webm'."}}}
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'properties': {}}
Input schema
{'type': 'object', 'required': ['text'], 'properties': {'text': {'type': 'string', 'maxLength': 5000, 'description': 'English text to convert to speech. Max 5000 characters.'}, 'speed': {'type': 'number', 'default': 1.0, 'description': 'Speech speed multiplier (0.5 = half speed, 2.0 = double).'}, 'voice': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "Voice ID (e.g. 'af_heart', 'am_adam'). Uses default if omitted."}}}
Input schema
{'type': 'object', 'required': ['audio_base64'], 'properties': {'audio_base64': {'type': 'string', 'maxLength': 20000000, 'description': 'Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.'}, 'audio_format': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "Audio format hint â\x80\x94 'wav', 'mp3', 'ogg', 'flac', 'webm'. Auto-detected if omitted."}, 'include_timestamps': {'type': 'boolean', 'default': True, 'description': 'If true, include word-level start/end times and confidence.'}}}
Input schema
{'type': 'object', 'required': ['audio_base64'], 'properties': {'diarize': {'type': 'boolean', 'default': False, 'description': 'Enable speaker diarization to identify who spoke each word.'}, 'language': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "Language code (e.g. 'en', 'es', 'zh'). Auto-detected when omitted."}, 'audio_base64': {'type': 'string', 'maxLength': 20000000, 'description': 'Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.'}}}
Recent tool changes
Similar MCP servers
Riddle
Builds, edits, publishes, embeds, and analyzes quizzes, polls, forms, personality tests, and related question banks.
api
Creates, manages, AI-generates, and renders quiz and flashcard videos, including templates and downloadable video outputs.
PodLearn
Searches podcasts and YouTube, transcribes episodes, searches transcripts, generates structured lessons, and creates LinkedIn pos…
Tutorializer Video Tutorials
Manages localized tutorials, voiceover segments, pronunciation rules, and registered rendered tutorial videos.
Escola de Rádio TV & Web (ER+)
Provides read-only access to a Brazilian radio school's courses, classes, teachers, podcasts, blog, services, and availability.
Brainiall Pronunciation
Assesses English pronunciation, transcribes audio, identifies speakers by voiceprint and synthesizes English speech in multiple v…
QuizBase
Searches, filters, samples, and manages a large multilingual catalog of trivia questions with categories, topics, provenance, and…
Soluble(s) — Podcast de journalisme de solutions
Searches French solutions-journalism podcast episodes and transcripts, including child-friendly versions and practical actions.