Every business owner has had this moment: a customer tells you something’s broken, and your stomach drops — because they knew before you did. That’s what running an AI stack without monitoring feels like. The thing answering your customers sits behind a curtain, and the first person to notice it fell over is the person paying you.
I run my own AI infrastructure on two rented servers. Four free tools tell me whether it’s alive. Not a custom platform. Four pieces of software that already exist and already work.
Prometheus is the notebook that never stops writing. Every 15 seconds it asks my AI gateway — the piece of software every request passes through — for a page of numbers: how many requests are in flight, how many tokens went in and out, how many failed, how long the first word of a reply took to arrive. Fifteen seconds matters. At an hourly reading, a burst of slow answers is just an average that looks fine. At 15 seconds, you can see the spike, and you can see it while it’s happening. It keeps 15 days of history, and it costs me under a gigabyte of disk.
Grafana is the wall of dials. Prometheus writes the numbers down; Grafana draws them. Mine is a 13-panel dashboard: output speed, success rate, time to first token, latency per provider, and a plain grid that says whether each of my provider keys is up. It refreshes on the same 15-second beat. Before this, when something felt slow, I had no vocabulary for it — was it me, the model, or the provider? Now I open one page and the answer is already on screen. That’s the difference between an opinion and a fact.
Uptime Kuma is the smoke alarm. Grafana tells me how well things are going. Uptime Kuma tells me whether they’re going at all. It pings my services on a schedule and, when one stops answering, it messages me on Slack. That’s the whole point of this post: I learn my stack is down from a notification, not from a client. It runs on the box that doesn’t carry my workloads, so if that box dies, the alarm keeps ringing instead of dying with it.
Glance is the one screen I actually look at. A home dashboard with my VPS stats, my running containers, and the feeds I read, all on one page. Nothing on it is clever. It’s just the first thing I open, and it turns “how are things going” from a task into a glance.
Monitoring is the cheapest trust you can buy. The tools are free, and the monitoring containers plus the gateway they watch fit in about 380MB of memory. They answer the question every customer silently asks: is this thing actually up? I can’t prove reliability by saying the word. I can only notice trouble first, fix it quietly, and let them never find out.
The non-engineer’s version is four prebuilt tools, not a custom observability platform. I didn’t build a monitoring system. I installed four things other people maintain and pointed them at my own stack.
If you can’t see it, you can’t trust it — and neither can your customers.