Speech to Text - Streaming

Stream real-time audio transcription with our SDKs. This guide covers how to implement live audio streaming for immediate speech-to-text conversion in your applications.

Prerequisites

Before you begin, make sure you have:

  • An aiOla API key (get one here)
  • Python 3.10+ (for Python SDK) or Node.js 18+ (for TypeScript SDK)
  • Microphone access (for live audio streaming)

Installation

pip install 'aiola[mic]'

Step 1: Set up authentication

First, generate an access token and create your client:

For comprehensive authentication details, security considerations, and token management strategies, see our Authentication Guide.

import os
from aiola import AiolaClient
# Generate access token
result = AiolaClient.grant_token(
api_key=os.getenv('AIOLA_API_KEY') or 'YOUR_API_KEY'
)
# Create client using the access token
client = AiolaClient(
access_token=result.access_token
)

Step 2: Initialize streaming connection

Create a streaming connection with event handlers:

from aiola import AiolaClient, MicrophoneStream
from aiola.types import LiveEvents
# Create streaming connection
connection = client.stt.stream(lang_code='en')
# Set up event handlers
@connection.on(LiveEvents.Transcript)
def on_transcript(data):
print('Transcript:', data.get('transcript', data))
@connection.on(LiveEvents.Connect)
def on_connect():
print('Connected to streaming service')
@connection.on(LiveEvents.Disconnect)
def on_disconnect():
print('Disconnected from streaming service')
@connection.on(LiveEvents.Error)
def on_error(error):
print('Streaming error:', error)

Step 3: Start streaming with microphone

Start the streaming connection and pipe microphone audio:

# Connect to the streaming service
connection.connect()
try:
# Capture audio from microphone using the SDK's MicrophoneStream
with MicrophoneStream(
channels=1,
samplerate=16000,
blocksize=4096,
) as mic:
print("Listening... Speak into your microphone")
mic.stream_to(connection)
# Keep the main thread alive
while True:
try:
import time
time.sleep(0.1)
except KeyboardInterrupt:
print('Keyboard interrupt')
break
except KeyboardInterrupt:
print('Keyboard interrupt')
finally:
connection.disconnect()

Custom audio sources

For custom audio sources instead of microphone:

import asyncio
async def stream_audio_file():
# Connect to streaming service
connection = client.stt.stream(lang_code='en')
@connection.on(LiveEvents.Transcript)
def on_transcript(data):
print('Transcript:', data.get('transcript', data))
connection.connect()
# Stream audio file in chunks
with open('audio_file.wav', 'rb') as audio_file:
chunk_size = 4096
while True:
chunk = audio_file.read(chunk_size)
if not chunk:
break
# Send audio chunk
connection.send(chunk)
# Small delay to simulate real-time streaming
await asyncio.sleep(0.1)
# Close connection
connection.disconnect()
# Run the async function
asyncio.run(stream_audio_file())

Advanced streaming options

Keyword detection

Enable keyword detection during streaming:

# Create connection with keyword detection
connection = client.stt.stream(
lang_code='en',
keywords={
"postgres": "PostgreSQL",
"k eight s": "Kubernetes"
}
)

Multiple language support

Stream with different languages:

# Supported languages: en, de, fr, es, pr, zh, ja, it
connection = client.stt.stream(lang_code='es') # Spanish

Error handling

Implement robust error handling for streaming:

try:
connection = client.stt.stream(lang_code='en')
@connection.on(LiveEvents.Error)
def on_error(error):
print(f"Streaming error: {error}")
# Implement reconnection logic here
@connection.on(LiveEvents.Disconnect)
def on_disconnect():
print("Connection lost. Attempting to reconnect...")
# Implement reconnection logic
connection.connect()
except Exception as e:
print(f"Failed to initialize streaming: {e}")

Complete working example

Here’s a complete Python example that combines all steps:

import os
from aiola import AiolaClient, MicrophoneStream
from aiola.types import LiveEvents
def live_streaming():
try:
# Step 1: Generate access token, save it
result = AiolaClient.grant_token(
api_key=os.getenv('AIOLA_API_KEY') or 'YOUR_API_KEY'
)
# Step 2: Create client using the access token
client = AiolaClient(
access_token=result.access_token
)
# Step 3: Start streaming
connection = client.stt.stream(
lang_code='en'
)
@connection.on(LiveEvents.Transcript)
def on_transcript(data):
print('Transcript:', data.get('transcript', data))
@connection.on(LiveEvents.Connect)
def on_connect():
print('Connected to streaming service')
@connection.on(LiveEvents.Disconnect)
def on_disconnect():
print('Disconnected from streaming service')
@connection.on(LiveEvents.Error)
def on_error(error):
print('Streaming error:', error)
connection.connect()
try:
# Capture audio from microphone using the SDK's MicrophoneStream
with MicrophoneStream(
channels=1,
samplerate=16000,
blocksize=4096,
) as mic:
mic.stream_to(connection)
# Keep the main thread alive
while True:
try:
import time
time.sleep(0.1)
except KeyboardInterrupt:
print('Keyboard interrupt')
break
except KeyboardInterrupt:
print('Keyboard interrupt')
except Exception as error:
print('Error:', error)
finally:
connection.disconnect()
if __name__ == "__main__":
live_streaming()

Best practices

  1. Audio Quality: Use 16kHz sample rate, mono channel for optimal results
  2. Chunk Size: 4096 bytes is recommended for real-time performance
  3. Error Handling: Always implement reconnection logic for production use
  4. Resource Cleanup: Properly disconnect streaming connections when done
  5. Audio Input: Handle audio input sources and permissions appropriately
  6. Latency: Consider buffering strategies for smoother transcription

Supported audio formats

For streaming, the following formats work best:

  • PCM 16-bit (recommended)
  • WAV uncompressed
  • Raw audio at 16kHz sample rate

Next steps

Now that you’ve implemented streaming transcription, you can:

Browser Examples

For web applications, check out our complete browser microphone streaming example: