- All Courses
- Building AI Apps with LLMs
- Shipping to Production
- Managing cost and latency
Managing cost and latency
Cost is prompt design, model choice, and how often you call. Measure per-request tokens before you try to optimise anything.
- 13m
- Intermediate

Video by Trevor Spires · YouTube
Check your understanding
3 questions. Unlocks when you finish the video.
Overview
What this lesson covers
- Attribute spend to specific calls and prompts
- Trim prompts and route by difficulty
- Set timeouts and fall back gracefully
This lesson sits in Shipping to Production, part of Building AI Apps with LLMs. It assumes what came before it and leads directly into the next lesson in the module.
In this lesson you will:
- Attribute spend to specific calls and prompts
- Trim prompts and route by difficulty
- Set timeouts and fall back gracefully
Resources
Your notes will live here
Note taking is not available yet. Nothing you type would be saved, so the tab stays read-only for now.