Image, Audio & Multimodal Generation SDKs
Generate images, transcribe and synthesize speech, and understand the growing multimodal generation tooling landscape.
Introduction
Generative AI extends well beyond text. Image generation, speech-to-text, and text-to-speech each have their own dedicated SDKs, following the same "install, call, get a result" pattern as every text-based tool earlier in this course.
- How to generate an image from a text prompt with the OpenAI Images API.
- How to transcribe audio to text with Whisper.
- How to synthesize natural-sounding speech with ElevenLabs.
A Real-Life Analogy First
Think about how a human uses more than one sense to understand and act in the world — eyes to see, ears to hear, a voice to speak. Every tool in this course up to now has effectively given an AI application only one sense: reading and writing text. This lesson adds the others: image generation is a "hand that can draw," Whisper is an "ear that can listen," and ElevenLabs is a "voice that can speak" — each one a separate, specialized capability that can be combined with the text abilities from earlier lessons.
Image Generation
Use case: the same openai SDK from lesson 5 also exposes an image generation endpoint (DALL·E) — a text prompt in, an image URL or base64 file out.
from openai import OpenAI
client = OpenAI()
response = client.images.generate( model="dall-e-3", prompt="A minimalist icon of an open book made of glowing circuit lines", size="1024x1024", n=1,)
print(response.data[0].url)Click Run to see what this code prints.
Speech-to-Text with Whisper
Use case: Whisper (available both as an open-source Hugging Face model and as a hosted OpenAI API endpoint) transcribes spoken audio into text, and can also translate non-English speech directly into English text in one step.
from openai import OpenAI
client = OpenAI()
with open("meeting-recording.mp3", "rb") as audio_file: transcript = client.audio.transcriptions.create( model="whisper-1", file=audio_file, )
print(transcript.text)Click Run to see what this code prints.
Text-to-Speech with ElevenLabs
Use case: ElevenLabs specializes in high-quality, natural-sounding text-to-speech, offering a wide range of voices as well as custom voice cloning — commonly used for voice assistants, audiobook narration, and accessibility features.
pip install elevenlabsfrom elevenlabs import ElevenLabs
client = ElevenLabs() # reads ELEVENLABS_API_KEY from the environment
audio = client.text_to_speech.convert( voice_id="21m00Tcm4TlvDq8ikWAM", text="Your order has shipped and will arrive within 3 to 5 business days.", model_id="eleven_multilingual_v2",)
with open("confirmation.mp3", "wb") as f: for chunk in audio: f.write(chunk)Click Run to see what this code prints.
Chaining these three together — Whisper to transcribe a user's spoken question, an LLM (lesson 5) to generate an answer, and ElevenLabs to speak it back — is the exact pattern behind most voice assistant products, like giving the AI a full set of "ears, brain, and voice" working in sequence.
Choosing a Modality Tool
| Need | Tool |
|---|---|
| Generate an image from a text description | OpenAI Images API (DALL·E), or open-weight Diffusers models via Hugging Face |
| Convert spoken audio to text | Whisper (hosted API or open-source model) |
| Convert text to natural-sounding speech | ElevenLabs, or OpenAI's text-to-speech endpoint |
| Understand an image alongside text in one prompt | A multimodal model like Gemini (lesson 6) or GPT-4o |
Common Mistakes
- Assuming image or audio generation is free of cost concerns the way a small text call is — media generation is typically priced higher per request.
- Not checking usage rights on generated images before using them commercially — check the provider's current terms.
- Sending very large audio files directly without checking a provider's file size or duration limits, which vary by endpoint.
Best Practices
- Cache or store generated media rather than regenerating identical requests repeatedly.
- Set explicit size, quality, and format parameters instead of relying on defaults, since media generation cost scales with these settings.
- For voice products, test with real accents and background noise conditions, not just clean studio audio.
Frequently Asked Questions
Yes, for open-weight image models (like Stable Diffusion via Hugging Face Diffusers) — the same LoRA/PEFT concepts from lesson 17 apply to image models as well as text ones.
Whisper supports many languages with strong accuracy, and can translate directly to English text, though quality varies by language and audio clarity.
OpenAI, Google, and several other providers also offer text-to-speech endpoints — ElevenLabs is called out here for its particularly natural-sounding voices and cloning features, not because it is the only option.
Conceptually yes — mobile photo apps typically use a smaller, optimized version of similar generative image technology, often run on-device (similar in spirit to lesson 18's local inference) rather than calling a cloud API, for speed and privacy.
Key Takeaways
- Image, speech-to-text, and text-to-speech each have dedicated SDKs following the same install-and-call pattern as text tools.
- The OpenAI SDK alone covers text, images, and audio transcription through different endpoints.
- ElevenLabs specializes in natural-sounding text-to-speech and voice cloning.
- Chaining transcription, an LLM, and speech synthesis together builds a full voice assistant pipeline.
Summary
Multimodal generation extends everything covered so far in this course beyond text — the same install-a-package, call-an-endpoint pattern applies whether the output is a sentence, an image, or a voice recording.
- You can generate an image from a text prompt.
- You can transcribe audio to text with Whisper.
- You can synthesize natural-sounding speech with ElevenLabs.