Servidor MCP

Brainiall Pronunciation

com.brainiall/pronunciation
Educación Medios y contenido Público y accesible MCP 2025-11-25

Qué hace este MCP

Assesses English pronunciation, transcribes audio, identifies speakers by voiceprint and synthesizes English speech in multiple voices.

assess_pronunciation
Assess Pronunciation
Assess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats. text: The reference English text that the speaker was expected to read aloud. audio_format: Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'. Defaults to 'wav'. Returns: dict with keys: - overallScore (int 0-100): Overall pronunciation quality - sentenceScore (int 0-100): Sentence-level fluency and accuracy - words (list): Per-word scores, each containing: - word (str): The word - score (int 0-100): Word pronunciation score - phonemes (list): Per-phoneme scores with IPA/ARPAbet notation - decodedTranscript (str): What the model heard (ASR transcript) - transcript (str): Reference text - confidence (float 0-1): Scoring confidence - warnings (list[str]): Quality warnings if any - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'required': ['audio_base64', 'text'], 'properties': {'text': {'type': 'string', 'maxLength': 10000, 'description': 'The reference English text that the speaker was expected to read aloud.'}, 'audio_base64': {'type': 'string', 'maxLength': 20000000, 'description': 'Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats.'}, 'audio_format': {'type': 'string', 'default': 'wav', 'description': "Audio format hint â\x80\x94 one of 'wav', 'mp3', 'ogg', 'webm'."}}}
check_pronunciation_service
Check Pronunciation Service
Check if the Brainiall Pronunciation service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API version
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'properties': {}}
check_stt_service
Check STT Service
Check if the Brainiall Speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the speech-recognition model is loaded - version (str): API version
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'properties': {}}
check_tts_service
Check TTS Service
Check if the Brainiall Voice service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the synthesis model is loaded - version (str): API version
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'properties': {}}
check_whisper_service
Check Brainiall Speech Pro Service
Check if the Brainiall Speech Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Brainiall Speech Pro engine is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version - modelName (str): Engine version identifier
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'properties': {}}
get_phoneme_inventory
Get Phoneme Inventory
Get the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: - arpabet (str): ARPAbet symbol (e.g. 'AA', 'TH') - ipa (str): IPA notation - example (str): Example word containing the phoneme - category (str): vowel, consonant, or diphthong
Solo lectura Idempotente
Esquema de entrada
{'type': 'object', 'properties': {}}
list_tts_voices
List TTS Voices
List all available Brainiall Voice synthesis voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice ID
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'properties': {}}
synthesize_speech
Synthesize Speech
Generate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - am_adam, am_michael (American male) - bf_emma, bf_isabella (British female) - bm_george, bm_lewis, bm_daniel (British male) Args: text: English text to synthesize (1-5000 characters). voice: Voice ID. See list above. Defaults to 'af_heart'. speed: Speed multiplier from 0.5 to 2.0 (default: 1.0). Returns: dict with keys: - audio_base64 (str): Base64-encoded WAV audio (16-bit PCM, 24kHz) - duration_ms (str): Audio duration in milliseconds - voice (str): Voice ID used - text_length (str): Input text character count - processing_ms (str): Synthesis time in milliseconds
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'required': ['text'], 'properties': {'text': {'type': 'string', 'maxLength': 5000, 'description': 'English text to convert to speech. Max 5000 characters.'}, 'speed': {'type': 'number', 'default': 1.0, 'description': 'Speech speed multiplier (0.5 = half speed, 2.0 = double).'}, 'voice': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "Voice ID (e.g. 'af_heart', 'am_adam'). Uses default if omitted."}}}
transcribe_audio
Transcribe Audio
Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detected from magic bytes if omitted. include_timestamps: Whether to include word-level timing (default: true). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str): The transcribed word - start (float): Start time in seconds - end (float): End time in seconds - confidence (float 0-1): Word-level confidence - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, audio length, model version - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'required': ['audio_base64'], 'properties': {'audio_base64': {'type': 'string', 'maxLength': 20000000, 'description': 'Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.'}, 'audio_format': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "Audio format hint â\x80\x94 'wav', 'mp3', 'ogg', 'flac', 'webm'. Auto-detected if omitted."}, 'include_timestamps': {'type': 'boolean', 'default': True, 'description': 'If true, include word-level start/end times and confidence.'}}}
transcribe_audio_pro
Transcribe Audio Pro
Transcribe audio with Brainiall Speech Pro — multilingual transcription. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_base64: Base64-encoded audio (WAV, MP3, OGG, FLAC, WebM). language: Language code. Auto-detected if omitted. Supports 99 languages. diarize: Enable speaker diarization (default: false). When true, each word includes a speaker label (e.g. SPEAKER_00, SPEAKER_01). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str), start (float), end (float), confidence (float 0-1) - speaker (str|null): Speaker label when diarize=true - speakers (dict|null): Speaker info with count and labels - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, language, languageProbability - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'required': ['audio_base64'], 'properties': {'diarize': {'type': 'boolean', 'default': False, 'description': 'Enable speaker diarization to identify who spoke each word.'}, 'language': {'anyOf': [{'type': 'string'}, {'type': 'null'}], 'default': None, 'description': "Language code (e.g. 'en', 'es', 'zh'). Auto-detected when omitted."}, 'audio_base64': {'type': 'string', 'maxLength': 20000000, 'description': 'Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.'}}}
voice_id_enroll
Enroll Voiceprint
Enroll a voiceprint for a speaker from ~2s of clear speech. Repeat with more clips to strengthen it. Only an irreversible embedding is stored — never the raw audio. Returns: dict with keys: speaker_id (str), n_samples (int), enrolled (bool).
Acceso externo
Esquema de entrada
{'type': 'object', 'required': ['audio', 'speaker_id', 'group_id'], 'properties': {'audio': {'type': 'string', 'description': 'Base64-encoded WAV with >= 2s of clear speech'}, 'group_id': {'type': 'string', 'maxLength': 256, 'description': 'The group/namespace this speaker belongs to'}, 'speaker_id': {'type': 'string', 'maxLength': 256, 'description': 'Your identifier for this speaker'}}}
voice_id_identify
Identify Speaker (1:N)
1:N identification — rank everyone enrolled in the group against this clip. Returns: dict with keys: candidates (list of {speaker_id, similarity}, best first).
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'required': ['audio', 'group_id'], 'properties': {'audio': {'type': 'string', 'description': 'Base64-encoded WAV of the clip to identify'}, 'top_k': {'type': 'integer', 'default': 5, 'maximum': 50, 'minimum': 1, 'description': 'How many candidate speakers to return'}, 'group_id': {'type': 'string', 'maxLength': 256, 'description': 'The group/namespace to search within'}}}
voice_id_list_speakers
List Enrolled Speakers
List the speakers enrolled in a group. Returns: dict with keys: speakers (list of {speaker_id, n_samples, ...}).
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'required': ['group_id'], 'properties': {'group_id': {'type': 'string', 'maxLength': 256, 'description': 'The group/namespace to list'}}}
voice_id_verify
Verify Speaker (1:1)
1:1 verification — is this clip the enrolled speaker? Returns: dict with keys: similarity (float), match (bool), threshold (float).
Solo lectura Acceso externo Idempotente
Esquema de entrada
{'type': 'object', 'required': ['audio', 'speaker_id', 'group_id'], 'properties': {'audio': {'type': 'string', 'description': 'Base64-encoded WAV of the clip to check'}, 'group_id': {'type': 'string', 'maxLength': 256, 'description': 'The group/namespace'}, 'threshold': {'anyOf': [{'type': 'number'}, {'type': 'null'}], 'default': None, 'description': 'Optional cosine-similarity threshold (defaults to a tuned value); higher = stricter'}, 'speaker_id': {'type': 'string', 'maxLength': 256, 'description': 'The enrolled speaker to verify against'}}}
Añadido
check_whisper_service
21 de September de 2026 a las 02:47
Añadido
transcribe_audio_pro
21 de September de 2026 a las 02:47
Añadido
check_tts_service
21 de September de 2026 a las 02:47
Añadido
list_tts_voices
21 de September de 2026 a las 02:47
Añadido
synthesize_speech
21 de September de 2026 a las 02:47
Añadido
voice_id_list_speakers
21 de September de 2026 a las 02:47
Añadido
voice_id_identify
21 de September de 2026 a las 02:47
Añadido
voice_id_verify
21 de September de 2026 a las 02:47
Añadido
voice_id_enroll
21 de September de 2026 a las 02:47
Añadido
check_stt_service
21 de September de 2026 a las 02:47
Añadido
transcribe_audio
21 de September de 2026 a las 02:47
Añadido
get_phoneme_inventory
21 de September de 2026 a las 02:47
Añadido
check_pronunciation_service
21 de September de 2026 a las 02:47
Añadido
assess_pronunciation
21 de September de 2026 a las 02:47