Speech AI - Pronunciation, STT & TTS is listed as active in the registry but did not answer our last check. It exposes 10 tools. Last commit 3 Aug 2026.
Pronunciation scoring, speech-to-text, and text-to-speech for language learning
Over the last week it answered 91.2% of our checks. We check every 15 minutes, so you hear about the next outage within the hour — not from your users.
Endpoint below is the one we actually reach during checks — not the one copied from a README. Last verified 14 min ago.
claude mcp add speech-ai --transport http https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp
{
"mcpServers": {
"speech-ai": {
"url": "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp"
}
}
}
[mcp_servers.speech-ai]
url = "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp"
{
"mcpServers": {
"speech-ai": {
"url": "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp"
}
}
}
{
"mcpServers": {
"speech-ai": {
"url": "https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp"
}
}
}
Read directly from the server with tools/list, grouped by what they act on.
If a tool disappears, we record the date.
transcribe_audio
transcribe_audio_pro
check_tts_service
list_tts_voices
assess_pronunciation
get_phoneme_inventory
check_pronunciation_service
check_stt_service
synthesize_speech
check_whisper_service
| URL | Transport | State | Latency | Checked |
|---|---|---|---|---|
| https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcp | streamable-http | answering | 20000 ms | 14 min ago |
Speech-to-text transcription for AI agents.
AI speech-to-text for public TikTok videos: SRT, VTT, word timings, speaker labels, 90+ languages.
Text extraction, keyword extraction, language detection, and chunking for RAG
Self-learning memory for AI coding agents with pattern detection and confidence scoring.
AI voice generation: text-to-speech and voice cloning from any MCP client.
Local business intel for AI agents: audits, lead scoring, tech stack, prospecting.
Self-learning operational context layer for AI sports agents. Profile 001: golf.
Text-to-speech, speech-to-text, audio-to-face lipsync, and motion-capture clips for 3D agents.
Answers built from our own checks of this server.