Put a Meter on Your Team’s AI Usage

How to use LiteLLM virtual keys, budgets, and spend tracking to see exactly where your team’s OpenAI and Claude money goes, with a setup you can run in five minutes.The Problem: An Unmetered AI BillHere is a story that probably sounds familiar.Someone on the team creates an OpenAI or Anthropic API…

How to use LiteLLM virtual keys, budgets, and spend tracking to see exactly where your team’s OpenAI and Claude money goes, with a setup you can run in five minutes.The Problem: An Unmetered AI BillHere is a story that probably sounds familiar.Someone on the team creates an OpenAI or Anthropic API key. It goes into a .env file. Then it gets pasted into Slack "just for now," copied onto a teammate's laptop, and baked into a staging script. Six weeks later the invoice arrives, and it is three times what anyone expected.Who spent it? Which project? Was it a genuine workload, or a test loop someone forgot to stop on Friday evening?Nobody can say. One shared key means one shared bill, and a shared bill means no accountability and no way to fix the problem next month.Electricity works because every home has a meter. Your AI usage should too.Takeaway: if every developer shares one API key, you cannot control what you cannot measure.The Idea: A Gateway Between Your Team and the ProvidersThe fix is to stop letting people talk to OpenAI, Anthropic, or Google directly. Instead, everyone talks to one gateway that you run, and the gateway talks to the providers.Developer / App → LiteLLM Gateway → OpenAI / Claude / Gemini / ...That is exactly what LiteLLM does. It is an open-source, OpenAI-compatible AI gateway. Your apps keep using the OpenAI SDK they already know; they just point at your gateway’s URL instead of the provider’s.Because every request passes through one place, the gateway can:Issue virtual keys, one per developer, project, or app, so real provider keys never leave the gatewayAttach a budget, rate limit, and model access list to each keyTrack spend automatically for every requestThis is a crowded and growing category in 2026. Comparison write-ups list LiteLLM alongside gateways like Bifrost, Kong AI Gateway, Cloudflare AI Gateway, and OpenRouter, and the common pitch is the same: without a shared layer, there is no place to enforce budgets or even see where tokens are going.LiteLLM’s particular niche is being MIT-licensed and self-hostable, so your prompts and keys stay in your own infrastructure. (There is also a separate commercial enterprise tier for things like SSO and audit logs, while the core gateway is free to run yourself.)Takeaway: A gateway turns “everyone calls the provider directly” into “everyone calls us,” and that one change is what makes metering possible.Installing the Meter: Start LiteLLM in Two CommandsYou need Docker installed. Then, in an empty folder:curl -sSLO https://github.com/BerriAI/litellm/raw/main/docker/docker-compose.quickstart.ymlprintf 'LITELLM_MASTER_KEY=sk-%s\nLITELLM_SALT_KEY=sk-%s\n' "$(openssl rand -hex 32)" "$(openssl rand -hex 32)" > .envdocker compose -f docker-compose.quickstart.yml up -dWhat just happened:You downloaded the quickstart compose file. It defines two services: the gateway on port 4000 and a Postgresdatabase. Read it before you run it; it is short, and you can edit it afterward to pin a specific release tag.You generated two secrets into a .env file.You started everything in the background.The Postgres database matters more than it looks. It is where your models, virtual keys, and spend logs live. No database means no spend tracking, which would defeat the whole point of this article.Two keys, two very different jobs:Keep the .env file. The salt key has no in-place rotation. If you regenerate it later, every provider credential already stored in the database becomes unreadable until you re-enter it.Takeaway: Two commands and a .env file give you a gateway plus the database that remembers every request.Connect Your Models in the Admin UIEverything else happens in the browser.Open http://localhost:4000/uiLog in with username admin and your LITELLM_MASTER_KEY as the passwordGo to Models + Endpoints and open the Add Model tabPick a provider, choose the models you want to expose, and paste your provider API keyClick Test Connect to verify the key, then Add ModelLiteLLM ships with each provider’s model catalog, so you select models from a list instead of typing names, and the pricing is already mapped. That built-in pricing is what lets it turn token counts into dollars later.Screenshot idea: The Add Model tab with a provider selected, and the All Models list showing pricing.Prefer environment variables to pasting keys in the UI? Add the provider key to the litellm service in the compose file (for example OPENAI_API_KEY: ${OPENAI_API_KEY}) and enter os.environ/OPENAI_API_KEY in the API key field instead of the raw key.Then go to the Playground, pick your model, and send a message. If a reply comes back with latency and token counts, your gateway is working end to end. The Get Code button there generates the equivalent API call for your language.Takeaway: The real provider key goes in once, on your side, and nobody else needs to see it again.Give Each Teammate Their Own KeyThis is the heart of the whole setup.Go to Virtual Keys, click + Create New Key, give it a name, and click Create Key.Name keys after the person or the project: priya-dev, checkout-service, nightly-eval-job. The name is what you will read in the spend report later, so make it something a human recognizes at a glance.Copy the key immediately. It is shown only once.Now hand that virtual key to your developer, not the real OpenAI or Anthropic key. Because the gateway is OpenAI-compatible, they change two things in their existing code: the base URL and the key.curl http://localhost:4000/v1/chat/completions \ -H 'Authorization: Bearer sk-' \ -H 'Content-Type: application/json' \ -d '{ "model": "", "messages": [{"role": "user", "content": "Say hello in five words."}] }'With the OpenAI Python SDK it is the same idea:from openai import OpenAIclient = OpenAI( api_key="sk-", base_url="http://localhost:4000",)response = client.chat.completions.create( model="", messages=[{"role": "user", "content": "Say hello in five words."}],)print(response.choices[0].message.content)Here is why this beats sharing the real key:Attribution. Every request is tied to a named key, so spend has an owner.Containment. If a key leaks or someone leaves, you revoke that one key. Nothing else breaks, and you do not have to rotate the real provider key.Safety. The real provider key lives only inside the gateway; developers never see it.Takeaway: One key per person or project turns an anonymous bill into a list of names.Set the Limits: Budgets, Rate Limits, and Model AccessA meter that only watches is useful. A meter that can also cut the power is better.Each virtual key can carry its own budget, rate limits, and model access list. In practice, that lets you say things like:“Interns get a small monthly budget and can only call the cheaper models.”“The nightly evaluation job may run heavier models, but is rate limited so a loop cannot flood the API.”“Only the production service key can use the most expensive model.”A runaway script then hits a wall you chose in advance, instead of quietly running up a bill nobody notices until the invoice arrives. Industry write-ups on AI gateways also point out that hidden overhead, such as retries and duplicate calls, can add a meaningful amount on top of raw API fees, which is one more reason to set a hard ceiling per key rather than trust everyone to self-police.Set these when you create a key on the Virtual Keys page. LiteLLM’s docs on Budgets + Rate Limits and Model Access cover team-level controls and the full list of options if you want to go further than per-key limits.Takeaway: Limits are what turn monitoring into control.Reading the Meter: Where the Money WentBecause every request flowed through a named key, the answer to “who spent what?” is now a lookup instead of an investigation.LiteLLM records the tokens and cost of each request against the key that made it. Open the Admin UI’s usage and spend views and you can see which keys and models are driving cost, broken down per key and per model.Try this exercise once you have a few days of data:Sort keys by spend. Is the top key what you expected?Look at which models that key calls. Is it using an expensive model for a simple job?Compare its request count to its spend. A few very costly requests and a huge number of cheap ones need different fixes.That last question is where the savings hide. Usage keeps shifting toward reasoning models and long-context coding workloads, which are token-hungry, so the cost of one careless call is higher than it used to be.Screenshot idea: the spend view sorted by key, with one obviously high spender highlighted.Takeaway: Once usage has names attached, cost conversations become specific instead of vague.Gotchas Worth KnowingA few things that will save you an afternoon:1. No database, no budgets. LiteLLM can run without a database using just a config file, and that is fine if all you need is an OpenAI-compatible proxy. But virtual keys themselves need a database (requests carrying one fail with No connected db.), and a global max_budget is not enforced without one — the proxy just logs a single warning at startup and then keeps serving requests past your limit. If a budget is part of how you control spend, run with the database from the quickstart, not the config-only path.2. Guard the master key. Anyone holding it has full admin access. Keep it out of source control and rotate it if it ever leaks.3. Do not lose the salt key. Changing it makes every stored provider credential unreadable until you re-enter it.4. Pin a version. The compose file can be edited to pin a specific release tag. Do that before this becomes something people depend on.5. Localhost is for learning. For a real team, you need a proper deployment behind HTTPS, not a laptop running Docker Compose.Takeaway: Most problems come from running without the database or losing a key you were told to keep.Wrap-Up: Your Next StepsIn about five minutes, you now have a gateway on port 4000, a provider connected, a named virtual key per developer, and every request tied to a person or project.From here:Go to production. LiteLLM’s Production Deployment guide covers Helm, Terraform, and Kubernetes on AWS, GCP, and Azure, and the production checklist covers hardening and tuning.Set real budgets. Start generous, watch a week of data, then tighten.Add more providers. The same gateway can front OpenAI, Anthropic, and others, so one meter covers all of them.AI spend is heading in one direction. The teams that stay comfortable with it will not be the ones who spend the least, but the ones who can always answer the simple question: who spent what, and was it worth it?If you set this up, I would love to hear what your first spend report revealed. Leave a comment and tell me about the most surprising key on your list.This story is published under the Generative AI publication. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories. Let’s shape the future of AI together!Put a Meter on Your Team’s AI Usage was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

Source: Generative AI Pub — Published — Category: Image AI

🔗 Read full article on Generative AI Pub →