Fields
| Field | Type | Required | Description | Example |
|---|---|---|---|---|
channel | Optional[int] | ➖ | Zero-based audio channel index for the word, present when the provider transcribes channels separately | 0 |
confidence | Optional[float] | ➖ | Provider confidence for the word from 0 to 1, present when the provider returns per-word confidence | 0.98 |
end | float | ✅ | Word end time in seconds | 0.4 |
speaker | Optional[int] | ➖ | Speaker index for the word, present when the provider returns diarization data | 0 |
speaker_label | Optional[str] | ➖ | Provider speaker label for the word, present when the provider labels speakers with a string | speaker_0 |
start | float | ✅ | Word start time in seconds | 0 |
type | Optional[components.STTWordType] | ➖ | Kind of entry; omitted or “word” for spoken words, “audio_event” for non-speech sounds the provider tags with timestamps | word |
word | str | ✅ | The transcribed word, or the event tag such as “(laughter)” when type is audio_event | Hello |