yash@jain:~$

../ yash@jain:~$ cat /archive/bifrost-vs-litellm.md

buildThe plumbing

The gateway everyone recommends is 51,000 stars of someone else's code

LiteLLM is the default answer when you ask which AI gateway to run. I ran the numbers on it anyway — and picked the Go binary with a fraction of the stars instead.

Ask any forum which AI gateway to run and you’ll get one answer. LiteLLM. It has 51,000 stars, it supports more providers than I’ll ever need, and it’s the safe recommendation to give a stranger. I went the other way — and it wasn’t the security story that decided it: a poisoned update had already put backdoored versions of it on the public package registry, and I covered that when I made the switch. This is the rest of the evaluation.

An AI gateway is the one door all your AI traffic walks through. Every request from every tool and model passes through one piece of software. Which makes that software either the most boring thing you run, or the most dangerous. I wanted boring.

The benchmark was not close. I put both gateways on the same server, same keys, same models, and fired identical requests at each. In nine of ten tests, the Go binary won. Averaged across the run, it came back about three and a half times faster on ordinary requests, six times faster to the first streamed word, and four times faster with five requests at once. The other gateway runs on Python, and Python has a known ceiling under concurrency — published testing puts it falling over around a thousand requests a second. My server will never see that load. But a tool that buckles under pressure it never faces is telling you something.

Then there’s the shape of the thing. LiteLLM is Python, so installing it pulls in a long chain of packages written by people I’ve never met. The Go binary is a single compiled file with no package registry behind it. A big dependency chain is just a bigger front door. And remember what this software holds: every API key I own, in one place. It’s the most valuable thing on my server.

The Go gateway gave me a second job for free. Bifrost isn’t only a model router. It also gathers every tool my agents use — search, GitHub, my own memory system, seven servers in total, 86 tools — behind one endpoint. Instead of each tool juggling its own tangle of connections, they connect once and get everything. That collapsed a mess I’d been managing by hand into one address.

And it solved a key problem I didn’t know I had. I hold multiple API keys for the same provider, and Bifrost supports that natively rather than through a hack. That matters more than it sounds, because providers cache the fixed front half of a conversation per account — split your traffic evenly across two keys and you halve your cache hit rate. So I run one primary key carrying the load and a second cold for failover. For an agent holding long conversations, a cache hit rate in the low nineties is close to the ceiling anyway; the rest of the misses come from the conversation growing, not from anything you can tune.

I’m not an engineer, so I’ll be blunt. LiteLLM wins on features, provider breadth, and the comfort of a crowd. If you already run a Python shop, or need a provider Bifrost doesn’t list, that crowd is worth something real. I traded it for a smaller thing that fails less often in ways I can’t see.

SMB Applicability Score: Bifrost 4/5. Worth it for a solo operator running their own gateway who cares about a small blast radius and streaming that starts fast; skip it if you’re already standardized on Python, or you need the widest provider list.

The gateway is the least glamorous part of the stack, and every request goes through it. Pick the one whose failure modes you can explain to yourself.

Related reading

← cd /archive