Multi-Provider Free Gateway

GitHub

Five hosted free-tier providers behind one OpenAI-compatible endpoint, with declared fallback chains, per-request overrides, streaming fallback, and an opt-in chat UI.

40m15m reading25m lab

Project Files

lite-llm-free-models
Explorer
docker-compose.yml
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
name: lite-llm-free-models

services:
  litellm:
    image: ghcr.io/berriai/litellm:main-stable
    container_name: lite-llm-free-models-proxy
    env_file:
      - .env
    command: ["--config", "/app/config.yaml", "--host", "0.0.0.0", "--port", "4000"]
    environment:
      DATABASE_URL: postgresql://${POSTGRES_USER:-litellm}:${POSTGRES_PASSWORD:-litellm_password}@postgres:5432/${POSTGRES_DB:-litellm}
      STORE_MODEL_IN_DB: "True"
    ports:
      - "${LITELLM_PORT:-4004}:4000"
    volumes:
      - ./litellm_config.yaml:/app/config.yaml:ro
    depends_on:
      postgres:
        condition: service_healthy
    restart: unless-stopped

  postgres:
    image: postgres:16-alpine
    container_name: lite-llm-free-models-postgres
    environment:
      POSTGRES_USER: ${POSTGRES_USER:-litellm}
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD:-litellm_password}
      POSTGRES_DB: ${POSTGRES_DB:-litellm}
    volumes:
      - postgres_data:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U ${POSTGRES_USER:-litellm} -d ${POSTGRES_DB:-litellm}"]
      interval: 5s
      timeout: 5s
      retries: 10
    restart: unless-stopped

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: lite-llm-free-models-ui
    profiles: ["ui"]
    depends_on:
      - litellm
    environment:
      OPENAI_API_BASE_URL: http://litellm:4000/v1
      OPENAI_API_KEY: ${LITELLM_MASTER_KEY}
      WEBUI_AUTH: "False"
      ENABLE_OLLAMA_API: "False"
    ports:
      - "${OPEN_WEBUI_PORT:-3004}:8080"
    volumes:
      - open-webui-data:/app/backend/data
    restart: unless-stopped

volumes:
  postgres_data:
  open-webui-data:
YAMLUTF-8
Ln 583 files

What you're running

One LiteLLM container exposes six model aliases on port 4004. One optional Open WebUI container (gated by a Compose profile) exposes a chat UI on port 3004.

AliasProviderModel
free-chatGeminigemini-2.5-flash (primary route)
free-chat-groqGroqllama-3.3-70b-versatile
free-chat-openrouterOpenRouteropenai/gpt-oss-20b:free
free-chat-cloudflareCloudflare Workers AI@cf/meta/llama-3.3-70b-instruct-fp8-fast
free-chat-huggingfaceHugging FaceQwen/Qwen2.5-7B-Instruct
free-dummy-chat(broken on purpose)gemini/not-a-real-model
Apps call free-chat only. The gateway decides whether it lands on Gemini, Groq, OpenRouter, Cloudflare, or HF — and the per-provider aliases exist for pinning when you want to.

What this is NOT

  • Not production — no TLS, no rate limiting, no secret manager
  • Not a multi-region SLO — fallback hides single-provider hiccups, not a global outage of all five
  • Not a benchmark — no quality, latency, or cost comparison here
  • Not a "best free model" guide — these are common free-tier defaults that happen to work today
Production shape (virtual keys, budgets, cache, observability) is covered in Phase 6 — Bedrock Prod Profile.

The routing config that does the work

router_settings:
  routing_strategy: simple-shuffle
  num_retries: 1
  timeout: 60
  fallbacks:
    - free-chat: [free-chat-groq, free-chat-openrouter, free-chat-cloudflare, free-chat-huggingface]
    - free-dummy-chat: [free-chat, free-chat-groq, free-chat-openrouter, free-chat-cloudflare, free-chat-huggingface]

Routing lives in YAML, not in application code. That is the whole point of an AI gateway.

Deploy

  1. 1 Download All the project files into a new folder
  2. 2 Copy env.example to .env and fill in the provider keys you proved in Phase 3
  3. 3 Start the gateway:
cp env.example .env

docker compose --env-file .env -f docker-compose.yml up -d
💡

Tip: Only fill in the providers you have keys for. The gateway boots even with subset keys — the aliases for missing providers will fail at call-time, not at startup.

Verify the alias list:

set -a
source .env
set +a

curl "http://localhost:${LITELLM_PORT:-4004}/v1/models" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY"

Expected — six entries:

free-chat
free-chat-groq
free-chat-openrouter
free-chat-cloudflare
free-chat-huggingface
free-dummy-chat
ℹ️

Note: /v1/models only proves the config loaded. It does not prove any provider key works. That's the next call.


Lab 1: Pin a single provider

Each provider-specific alias bypasses the routing pool and pins the request to that one provider. Useful for debugging, cost control, licence choice, or comparing answers across providers.

curl "http://localhost:${LITELLM_PORT:-4004}/v1/chat/completions" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "free-chat-cloudflare",
    "messages": [{"role": "user", "content": "Reply in one sentence: pinned to Cloudflare."}]
  }'
Repeat for each provider you have a key for. The cheap, boring proof that the gateway can reach each one.

Lab 2: Fallback via the dummy alias

free-dummy-chat points at a model name Gemini will reject. The first hop always fails, and the router walks the fallback list until something answers.

curl "http://localhost:${LITELLM_PORT:-4004}/v1/chat/completions" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "free-dummy-chat",
    "messages": [{"role": "user", "content": "Say fallback worked in one sentence."}]
  }'

A successful response shows the underlying provider model that actually answered, in the model field. The client never had to know the first route failed.

Lab 3: Fallback on a healthy route — mock_testing_fallbacks

The dummy alias works once. The cleaner pattern for demos, talks, and CI is the request-time flag — it forces the primary to fail on an otherwise-healthy route, with no config change:

curl "http://localhost:${LITELLM_PORT:-4004}/v1/chat/completions" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "free-chat",
    "mock_testing_fallbacks": true,
    "messages": [{"role": "user", "content": "Say second provider answered in one sentence."}]
  }'

You can also override the fallback list per request:

curl "http://localhost:${LITELLM_PORT:-4004}/v1/chat/completions" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "free-chat",
    "fallbacks": ["free-chat-groq"],
    "mock_testing_fallbacks": true,
    "messages": [{"role": "user", "content": "Say Groq answered in one sentence."}]
  }'
The per-request override wins over the YAML default.

Lab 4: Streaming still falls back

Same flag, with stream: true:

curl -N "http://localhost:${LITELLM_PORT:-4004}/v1/chat/completions" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "free-chat",
    "stream": true,
    "mock_testing_fallbacks": true,
    "messages": [{"role": "user", "content": "stream one short sentence about fallback"}]
  }'
You'll get standard SSE data: {...} chunks ending in data: [DONE], served by a fallback provider — even though the primary was forced to fail before the first chunk.
⚠️

Warning: What this does not do — mid-stream provider switch after chunks have already started flowing. Once a provider has sent its first token, the request stays on that provider. That's a fair and intentional limit.


Two UIs, two different questions

http://localhost:4004/ui/     LiteLLM admin UI    "is the gateway healthy?"
http://localhost:3004         Open WebUI chat     "can a human actually chat through it?"

LiteLLM admin UI

The proxy serves its own admin UI at /ui/. Log in with the master key (used as both username and password). With Postgres in this stack, the UI shows:
  • The six aliases from /v1/models
  • A Test Key playground for one-off chat calls against any alias
  • Request logs for the running process
  • Virtual key management (covered in detail in Phase 6)

Open WebUI chat — opt-in via Compose profile

Open WebUI is a ChatGPT-style browser UI. It's wired under a ui profile so the default docker compose up -d stays lean:

docker compose --env-file .env --profile ui up -d
Open http://localhost:3004. The model picker auto-discovers every alias from the gateway's /v1/models. In one window you can:
  • Pick free-chat for smart routing with fallback
  • Pick any free-chat- to pin one provider
  • Pick free-dummy-chat and watch the chat reply normally even though the first hop failed inside the gateway
This UI is "can a human actually chat through this and feel the fallback recovery for themselves." Demoing fallback by curl is fine for engineers — demoing it in a chat window is what makes a stakeholder nod.

Common Issues

Open WebUI shows unhealthy but the page works

The Open WebUI image ships a healthcheck that assumes auth is on. With WEBUI_AUTH=False (the default here for friction-free local testing) the probe can fail even though the UI serves HTTP 200. Ignore the healthcheck status — load the page and confirm directly.

Provider NotFoundError on free-chat-cloudflare

You set CLOUDFLARE_API_TOKEN but not CLOUDFLARE_ACCOUNT_ID. Both are required.

free-chat keeps landing on the same provider

routing_strategy: simple-shuffle is random per-request, not round-robin. Across many requests the load spreads. Across a handful of requests it can look sticky.

408 Timeout from Hugging Face

Free HF inference has cold starts. Re-run, or remove HF from the fallback chain for demos.

Next Steps