All Courses
Advanced
LiteLLM AI Gateway
Platform-engineering the LiteLLM AI gateway — from a local Ollama setup with fallback you can trigger, through five free-tier providers behind one endpoint, to AWS Bedrock in both dev and production shapes with virtual keys, response cache, and an honest gap table.
What you'll learn
Deploy a local LiteLLM + Ollama gateway with fallback you can trigger on demand
Validate free-tier provider keys directly with Python before any gateway YAML
Route five hosted free-tier providers behind one OpenAI-compatible endpoint
Prove fallback on a healthy route with mock_testing_fallbacks and per-request overrides
Stand up the smallest verifiable AWS Bedrock dev setup and isolate failures with four curls
Operate the production-shape Bedrock stack with Postgres, Redis, healthchecks, virtual keys, and response cache
Mint per-team virtual keys with budgets, TPM/RPM limits, and scoped model access
Plan the laptop → real server promotion with an honest "what is NOT here" gap table
Course content
20 sections · 136 articles · 3h 30m total
Course Roadmap & Gateway PatternArticle
15m
OpenAI-Compatible API as the Stable ContractComing soon
Soon
Provider Economics in 2026Coming soon
Soon
LiteLLM vs Cloudflare, Kong, Portkey, Self-BuiltComing soon
Soon
Course Setup, Repo Layout & Four-Curl MethodComing soon
Soon
Production-Grade Is a Checklist, Not a BinaryComing soon
Soon
LiteLLM + Ollama with FallbackArticle
Code35m
The Four-Curl Debugging Method, AppliedComing soon
Soon
Why Postgres Ends Up in the Local StackComing soon
Soon
Direct Provider Validation (Pre-Gateway)Article
Code25m
Five Providers, One EndpointArticle
Code40m
Routing Strategies Deep DiveComing soon
Soon
Declared Fallback Chains in YAMLComing soon
Soon
Per-Request Fallback OverrideComing soon
Soon
Streaming & the Mid-Stream Fallback LimitComing soon
Soon
Provider-Specific GotchasComing soon
Soon
Why the Dev Profile Exists SeparatelyComing soon
Soon
Smallest Verifiable Bedrock SetupArticle
Code30m
Seven Bedrock Failure Modes by LikelihoodComing soon
Soon
Bedrock Chat — Nova, Claude Haiku, Claude SonnetComing soon
Soon
Bedrock Embeddings — Titan & CohereComing soon
Soon
Bedrock Streaming & IAM PermissionComing soon
Soon
Aliasing Bedrock Models to OpenAI DefaultsComing soon
Soon
Azure OpenAI — Deployment Names & VersionsComing soon
Soon
Google Vertex AIComing soon
Soon
Anthropic Direct (Non-Bedrock)Coming soon
Soon
Cohere, Mistral & Second-Tier ProvidersComing soon
Soon
The Provider Matrix — When Each One WinsComing soon
Soon
Three-Service Topology — Gateway + Postgres + RedisComing soon
Soon
Healthchecks & depends_on: service_healthyComing soon
Soon
STORE_MODEL_IN_DB & the Admin UIComing soon
Soon
Production-Shape Bedrock StackArticle
Code45m
The Checklist Framing & Gap TableComing soon
Soon
Laptop → Real Server Promotion ChecklistComing soon
Soon
IAM-Role-Ready AuthComing soon
Soon
Three Kinds of Keys — Master, Operator, AppComing soon
Soon
Per-Key Budgets, TPM/RPM & Model ScopingComing soon
Soon
Team-Scoped Model Access & Least PrivilegeComing soon
Soon
Spend Tracking & Reconciliation GapsComing soon
Soon
Pre-Call Estimation vs Post-Call ReconciliationComing soon
Soon
Showback & ChargebackComing soon
Soon
The Finance Conversation in Engineer-SpeakComing soon
Soon
Exact-Match Redis CacheComing soon
Soon
Semantic Cache with Similarity ThresholdComing soon
Soon
Cache Invalidation — TTL, Flush, Version-KeyedComing soon
Soon
Measuring the Delta — p50/p95, Cost, Hit RateComing soon
Soon
Honest Cache Limits per Workload ShapeComing soon
Soon
What to Log at the Gateway LayerComing soon
Soon
Shipping Logs to Loki / Datadog / CloudWatchComing soon
Soon
Prometheus Exporter & Four Standard DashboardsComing soon
Soon
Langfuse for Trace-Level VisibilityComing soon
Soon
OpenTelemetry ExportComing soon
Soon
Alerting on Cliffs That MatterComing soon
Soon
The On-Call Runbook for an AI GatewayComing soon
Soon
PII Detection & Redaction (Presidio)Coming soon
Soon
Prompt-Injection Detection InboundComing soon
Soon
Content Filtering OutboundComing soon
Soon
Per-Key Guardrail PoliciesComing soon
Soon
SOC2 / HIPAA / GDPR PostureComing soon
Soon
Audit Logging for Regulated WorkloadsComing soon
Soon
Data Residency, Made ConcreteComing soon
Soon
Bearer Tokens vs JWT ValidationComing soon
Soon
OIDC — Okta, Auth0, Cognito, Azure ADComing soon
Soon
End-User Identity → Virtual Key MappingComing soon
Soon
SaaS Multi-Tenancy Key PatternsComing soon
Soon
BYOK — Bring-Your-Own-KeyComing soon
Soon
Per-Tenant Budgets, Model Access, AuditComing soon
Soon
A/B Testing & Shadow TrafficComing soon
Soon
Canary Deployments for Model SwapsComing soon
Soon
Latency-Aware RoutingComing soon
Soon
Cost-Aware RoutingComing soon
Soon
Circuit Breakers for Provider OutagesComing soon
Soon
Multi-Region / Multi-Account FailoverComing soon
Soon
Blast-Radius Isolation Across AccountsComing soon
Soon
Embeddings + Vector DB (pgvector / Qdrant / Pinecone)Coming soon
Soon
Full RAG Loop Behind the GatewayComing soon
Soon
Image Models — DALL-E, Stable DiffusionComing soon
Soon
Audio — Whisper, ElevenLabs, PollyComing soon
Soon
Vision & Multimodal ParityComing soon
Soon
Tool / Function-Calling Parity Across ProvidersComing soon
Soon
vLLM as a Serving BackendComing soon
Soon
Text Generation Inference (TGI)Coming soon
Soon
GPU Autoscaling (KEDA, Karpenter)Coming soon
Soon
Model Swap Without DowntimeComing soon
Soon
Hybrid Cloud + On-Prem Behind One GatewayComing soon
Soon
Model Context Protocol (MCP) HubComing soon
Soon
Agent Frameworks Behind the GatewayComing soon
Soon
Realtime / WebSocket APIsComing soon
Soon
Streaming Function-Call Wire FormatsComing soon
Soon
LiteLLM Config Schema Validation in CIComing soon
Soon
Dry-Run Testing for Model-List ChangesComing soon
Soon
Staged Rollout — Dev → Staging → ProdComing soon
Soon
GitOps with Flux / ArgoCDComing soon
Soon
Disaster Recovery and BackupsComing soon
Soon
Rolling Back a Bad Model Swap in < 5 MinComing soon
Soon
Why the Gateway Is on the Hot PathComing soon
Soon
k6 & Locust Scripts Against the ProxyComing soon
Soon
Profiling the Postgres CeilingComing soon
Soon
Profiling the Redis CeilingComing soon
Soon
When Horizontal Scaling Starts to MatterComing soon
Soon
Capacity-Planning ConversationComing soon
Soon
The Brief — B2B SaaS AI FeatureComing soon
Soon
Full Stack — TLS + RDS + ElastiCache + ObsComing soon
Soon
IAM Model & Audit TrailComing soon
Soon
Deployment — ECS / EKS / Single EC2Coming soon
Soon
Load Test, Failover Drill, Rollback DrillComing soon
Soon
90-Day Operational ReviewComing soon
Soon
Submitting the Capstone ArtifactsComing soon
Soon
Positioning as an LLM Platform EngineerComing soon
Soon
System Design Interviews for AI GatewaysComing soon
Soon
Writing the Right Kind of Public ContentComing soon
Soon
The About Blurb That Signals AuthorityComing soon
Soon
Building a Public Portfolio RepoComing soon
Soon
Roadmap & Gap List (Overview)Article
10m
Local Gateway + Fallback (Ollama)Article
Multi-Provider Free-Tier GatewayArticle
AWS Bedrock Dev ProfileArticle
AWS Bedrock Prod ProfileArticle
Cost Governance with Virtual Keys + BudgetsArticle
Semantic CachingArticle
Observability + Per-Key AnalyticsArticle
PII Redaction + GuardrailsArticle
Multi-Region / Multi-Account FailoverArticle
Streaming + Function-Calling ParityArticle
A/B Testing & Shadow TrafficArticle
Auth Beyond Bearer TokensArticle
SaaS Multi-Tenancy and BYOKArticle
Embeddings + Vector DB for RAGArticle
Image, Audio, and Multimodal RoutingArticle
Self-Hosted LLM Cluster (vLLM / TGI)Article
CI/CD and Config-as-Code for LiteLLMArticle
MCP Routing Through the GatewayArticle
Realtime / WebSocket APIsArticle
Agent Frameworks Behind the GatewayArticle
Cost Showback/Chargeback AutomationArticle
Load Testing & Capacity PlanningArticle
Prerequisites
- Basic Linux command line knowledge
- Docker and Docker Compose fundamentals
- Understanding of JSON and REST APIs
Multi-container orchestration used in every hands-on exercise.