Reading LiteLLM Through One Real Request
I recently started reading LiteLLM’s source code.
At first, I thought it was just a Python SDK that gives you one interface for OpenAI, Anthropic, Gemini, Kimi, and other providers.
That’s true, but it turns out to be a pretty incomplete picture.
After running the proxy locally and stepping through a real request in the debugger, LiteLLM feels much more like an LLM Gateway. It is not only forwarding requests. Around every model call, there is a lot of production stuff happening:
Authentication
Model routing
Provider protocol translation
Retry / fallback
Cost tracking
Logging and callbacks
Caching, rate limits, budgets, and more
The repository is huge. The litellm/ package alone has more than two thousand Python files and close to 600k lines of code.
So I did not try to read the whole thing.
Instead, I picked one path:
Start LiteLLM Proxy
→ Send POST /v1/chat/completions
→ Set breakpoints
→ Follow the request all the way to the provider
→ See how the response comes back
The request was intentionally simple:
{
"model": "kimi-k2.5",
"messages": [
{
"role": "user",
"content": "Reply with OK only"
}
]
}
That one request gave me a much clearer mental model of LiteLLM.
LiteLLM has two ways to use it
You can use LiteLLM as a Python SDK:
Your Python app
→ litellm.completion()
→ OpenAI / Kimi / Claude
Or you can run it as a gateway:
Your services
→ HTTP request
→ LiteLLM Proxy
→ Router
→ Provider
The SDK version is great for simple apps.
The Proxy version is more like a shared model platform. Your Java service, Python service, agent service, or frontend does not need to know the real provider key or even the real provider.
They only call LiteLLM.
LiteLLM decides things like:
Which model should handle this request?
Which provider should be used?
Should this be retried?
Should it fall back to another model?
How much did it cost?
Which user or team made the call?
That is why calling it a gateway makes sense.
Where does /v1/chat/completions enter?
The entry point is in:
litellm/proxy/proxy_server.py
The function is:
async def chat_completion(...)
This is the handler for:
POST /v1/chat/completions
It does not directly call Kimi or OpenAI.
Its job is more like a Spring Boot controller:
Receive the HTTP request
→ Authenticate the API key
→ Read the JSON body
→ Build internal request data
→ Hand the request to the next layer
The real request orchestrator is in:
litellm/proxy/common_request_processing.py
The key method is:
ProxyBaseLLMRequestProcessing.base_process_llm_request()
I think of it as the coordinator for one full LLM request:
Pre-call processing
→ Route the request
→ Call the model
→ Wait for the result
→ Run logging / cost / callbacks
→ Return the HTTP response
Why does LiteLLM do so much work before calling a model?
At first, I saw all the request cleanup and metadata work and thought: this is a lot before even calling the model.
Then it made sense.
You cannot trust fields from an HTTP request.
A client could send something like this:
{
"metadata": {
"user_api_key_user_id": "admin"
}
}
If the gateway trusted that value for permissions, budgets, or auditing, it would be a serious problem.
So LiteLLM does something closer to this:
Remove client-controlled internal fields
→ Read the real user/team identity from authentication
→ Write trusted metadata into the internal request context
This is one of the biggest differences between an SDK call and a gateway.
With an SDK call, your application is usually the trusted caller.
With a gateway, many different services and users may call it. So identity, permissions, budgets, and audit data have to be handled by the gateway itself.
data is not just the HTTP request body
In the debugger, I saw a data object that looked roughly like this:
{
"model": "kimi-k2.5",
"messages": [...],
"metadata": {
"user_api_key_user_id": "...",
...
}
}
This is more than the original JSON body.
It behaves more like a request context that travels through the whole call chain:
Request parameters
+ authenticated identity
+ request ID
+ routing information
+ logging context
+ cost attribution data
At first, passing a big dictionary through many layers felt messy.
But it makes more sense when you realize the same information is needed by the Router, provider adapter, logging hooks, budget checks, cache logic, and async tasks.
If every layer had to depend directly on FastAPI’s request object or authentication object, the coupling would be much worse.
Proxy and Router are not the same thing
The request later reaches:
litellm/proxy/route_llm_request.py
The main function is:
route_request(...)
This layer makes an important decision:
If llm_router exists
→ call the Router
Otherwise
→ call litellm.acompletion() directly
So LiteLLM does not always require a Router.
A simple flow can be:
Application
→ LiteLLM
→ One provider
A production flow can be:
Client
→ Proxy
→ Router
→ Multiple deployments
→ Provider
The Router is not responsible for sending the HTTP request itself.
Its job is to decide where the request should go.
For example:
Client requests smart-chat
↓
Router sees multiple deployments for smart-chat
↓
Router chooses one
↓
If it fails, retry or fallback
That is why the Router is one of the core pieces of the project.
Model, deployment, and provider are different things
These terms were confusing at first.
This is how I think about them now:
Client model name
→ The name sent by the client, like smart-chat
Model group
→ A logical group used by the Router
Deployment
→ One actual runnable model configuration
Provider
→ moonshot / openai / anthropic / gemini
Actual model
→ kimi-k2.5 / claude-sonnet-... / gpt-...
A request might look like this:
Client: smart-chat
↓
Router: choose deployment A
↓
Deployment A: moonshot/kimi-k2.5
↓
Provider: moonshot
↓
Actual model: kimi-k2.5
The Router answers:
Which deployment should I use?
Provider resolution answers:
Which provider adapter should handle this deployment?
Those are separate problems.
How does LiteLLM detect the provider?
The main logic lives in:
litellm/litellm_core_utils/get_llm_provider_logic.py
The main function is:
get_llm_provider(...)
It returns something like:
(
model,
custom_llm_provider,
api_key,
api_base
)
I originally expected logic like this:
if model.startswith("claude"):
return "anthropic"
But it is more careful than that.
The rough priority is:
1. Explicit provider configuration in the deployment
2. provider/model prefix
3. custom_llm_provider
4. api_base inference
5. Exact match against built-in model lists
6. Generalized fallback rules
The safest way is still to be explicit:
anthropic/claude-...
moonshot/kimi-k2.5
gemini/gemini-...
Because a model name is not enough to identify a real integration path.
For example, Claude can be called through:
Anthropic’s own API
AWS Bedrock
Google Vertex AI
The model is still Claude, but the endpoint, credentials, request format, and billing path are different.
How does LiteLLM unify different providers?
This is probably the most important part of the design.
LiteLLM does not force every provider into one giant universal JSON request format.
Instead, it normalizes two boundaries:
Before the provider:
model / messages / tools / stream / temperature / max_tokens
After the provider:
ModelResponse / Usage / Exception
The middle is handled by a provider adapter.
Unified request semantics
↓
Provider adapter
↓
Provider-specific HTTP request
↓
Raw provider response
↓
Provider adapter
↓
Unified response object
The base contract is in:
litellm/llms/base_llm/chat/transformation.py
The two important methods are:
transform_request()
transform_response()
The names are pretty direct:
transform_request()
Unified LiteLLM input → Provider request body
transform_response()
Raw provider response → LiteLLM response object
Anthropic is a good example of why this layer exists.
In OpenAI-style APIs, the system prompt is usually inside the messages array:
{
"messages": [
{"role": "system", "content": "You are a helpful assistant"}
]
}
Anthropic expects the system prompt separately. It also has its own content block format for messages, tool calls, tool results, and thinking.
So LiteLLM’s Anthropic adapter has to do things like:
Move system messages out of messages
Transform normal messages
Transform tool calls and tool results
Handle thinking rules
Handle Anthropic-specific headers
Kimi/Moonshot is much closer to OpenAI-compatible APIs, so its adapter can reuse much more of the OpenAI transformation logic.
That was a useful insight for me:
Similar protocol
→ reuse a shared adapter
Different protocol
→ add provider-specific transformations
What does LiteLLM get after the provider responds?
After the provider returns, LiteLLM gives the upper layers a:
ModelResponse(...)
Not a raw Kimi response. Not an Anthropic SDK object.
A unified ModelResponse.
It contains fields like:
id
model
choices
usage
So upper layers can always do:
response.choices[0].message.content
response.usage.prompt_tokens
response.usage.completion_tokens
They do not need provider-specific if statements everywhere.
LiteLLM also does not throw away provider-specific fields.
For example, reasoning content or reasoning token usage can still be preserved, even though the common response structure stays stable.
Why are cost and logging handled after the model returns?
At first, it seems like cost tracking could happen as soon as the request enters the gateway.
But you cannot know the real cost yet.
You only know what the user asked for. You do not know:
Which deployment was actually selected
Which provider was actually called
Whether retry happened
Whether fallback happened
How many input/output/reasoning tokens were used
Whether cache was involved
So the real post-call flow is closer to:
Provider response
→ ModelResponse / Usage
→ Read deployment and cost information
→ Update request status
→ Run success hooks / callbacks
→ Add response headers
→ Return response to the client
That is where the actual observability and billing data becomes reliable.
My current mental model of LiteLLM
If I had to describe LiteLLM in one sentence now:
LiteLLM is an LLM gateway that hides provider differences behind adapters and centralizes routing, retries, fallbacks, authentication, cost tracking, and observability.
It is not just about avoiding a few SDK integrations.
It solves the problem you get once model calls are spread across multiple services:
Every service stores provider keys
Every service implements retries
Every service calculates tokens and cost differently
Every service handles OpenAI / Claude / Gemini differently
Calling an SDK directly is fast in the beginning.
But once LLM calls become shared infrastructure, a gateway starts to make a lot more sense.
What I want to read next
Today I followed the main happy path.
There are still a lot of parts I only understand at an architecture level, not at source-code level:
How Router chooses between deployments
How load balancing strategies are implemented
Where rate limits are checked
How cooldown prevents repeatedly calling bad deployments
How cache and Redis work across multiple proxy instances
How streaming is normalized into OpenAI-style SSE
The next thing I want to read is deployment selection inside the Router.
That should connect retry, fallback, load balancing, and cooldown into one real production call path.