Analyze audio to identify speakers
Generate text from audio with instructions
Generate text from audio and instructions
Transcribe audio into text