Grok Voice STT 1.0
Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.
BENCHMARKS
Not in the Epoch Capabilities Index. Epoch benchmarks what it can run, so this is a gap in coverage, not a poor result.
RELEASE DATEAug 4, 2026
WEIGHTSClosed weights
CONTEXT WINDOW15K
MAX OUTPUT15K
REASONINGNo
TOOL CALLINGNo
INPUTSaudio
OUTPUTStext
Specifications from Models.dev. Availability and limits may differ by provider. Benchmark scores come from Epoch AI under CC BY 4.0.