Video by Hugging Face via YouTube

Prompt caching can cut the input cost of long AI agent sessions dramatically—but only when your harness preserves reusable prompt prefixes. This video explains what prompt caching actually stores, why agent costs compound across turns, how provider behavior differs, and the implementation mistakes that invalidate your cache.
—
🔗 *Links*
– Written tutorial: https://alejandro-ao.com/tutorials/prompt-caching/
– Tau coding agent: https://github.com/huggingface/tau
– Pi coding agent: https://github.com/badlogic/pi-mono
– Hugging Face Inference Providers: https://huggingface.co/docs/inference-providers/index
– OpenAI prompt caching: https://platform.openai.com/docs/guides/prompt-caching
– Anthropic prompt caching: https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching
– Gemini context caching: https://ai.google.dev/gemini-api/docs/caching
—
👋 *Connect with me*
– My website: https://alejandro-ao.com/
– X (Twitter): https://x.com/_alejandroao
– LinkedIn: https://www.linkedin.com/in/alejandro-ao/
—
🤓 *Topics Covered*
– How prompt caching works
– AI agent token cost reduction
– Cache-friendly agent harness design
—
⏱️ *Timestamps*
0:00 Why long AI agent sessions get expensive
0:45 What prompt caching actually caches
5:55 Provider pricing and cache discounts
9:19 Prompt caching best practices
15:40 Summary and cache monitoring