model call orchestration gateway · cost governance

AWS itself tells you to build an LLM gateway.
Veltrix already is that gateway.

22 models, 9 providers. Veltrix decides who answers each call, measures the cost of the request, and refuses it when the budget runs out.

You turn it on by swapping one line. And you start by measuring: a 4-week audit, no cost and no card.

base_url = "api.veltrixplatform.com/v1" ← the only change
the bill that dissolves · 30 days measured
measured
74.5%
of calls Veltrix acted on
measured
29.2%
answered without touching a provider
counterfactual
50.2%
cut where the router swapped the model
counterfactual
14.0%
floor across the whole bill
30 days, 4,957 calls, Veltrix's own account. Real traffic, not a simulation — and not a customer's.
  • OpenAI
  • Anthropic
  • Google
  • DeepSeek
  • Qwen
  • Zhipu
  • Moonshot
  • MiniMax
  • Groq
  • vLLM
  • Bedrock
  • Azure OpenAI
  • Copilot
  • Cursor
  • Perplexity
  • Grok
01 · the product

Four decisions per call. Watch each one happen.

rerouted to a cheaper model
64.0%
OpenAI
gpt-4o
US$ 2,50 / 1M
default
Anthropic
sonnet 4.5
US$ 3,00 / 1M
Google
2.5 flash
US$ 0,30 / 1M
DeepSeek
chat
US$ 0,27 / 1M
vLLM
your GPU
cost/hour
chosen
capability
context
quality
latency
cost
criticality
2,246 of 3,512 calls outside the cache, with nothing declared by you. Deterministic, and every choice stays in the log.
02 · the cut

Your number depends on your traffic. Put yours in.

R$ 40,000
1k500k
order of magnitude: ~222,222 calls / month
The order of magnitude assumes an average of R$ 0.18 per call — our own assumption to give scale, not a measurement of your traffic.
floor across the whole bill (14.0%)
R$ 5,600 / mo
includes everything the router left alone
where the router acted (50.2%)
R$ 20,080 / mo
only on calls where it swapped the model
A projection from our measured percentages, not a promise about your traffic. If you already run everything on cheap models, there is little to cut — and we say so before you sign.
04 · where it runs

Our cloud, yours, or pointing at your own GPU.

mode 01 · saasready

Our cloud

One line in your client. Nothing to deploy.

mode 02 · hybridready

Your model, Veltrix in the cloud

Your GPU answers; the cap and the attribution still come from us.

mode 03 · self-hostedpartial

Veltrix inside your infra

The embedding comes from a local model, not from OpenAI. Never run end to end on a clean machine — and we say so.

vLLM · TGI · Ollama — the case nobody solves

The fallback chain stops being a try/catch and becomes policy.

In the order you wrote it, never reordered by price — and every trigger is counted.

cost/hour

an own model has no per-token price: you declare the GPU cost and we attribute it per second occupied

counterfactual

the machine would cost the same hour if nobody called — idle GPU is shown separately

Where Veltrix has evidence — and where it cannot make a claim yet.

Coverage is not a generic savings promise. Every surface states its available signal and what that signal supports.

Browser

Estimated usage in compatible integrations

Estimated

Browser events; varies by provider

Desktop

Compatible tool in the foreground

Presence

Optional Local Collector on macOS

Terminal

Recognized CLI running

Presence

Local allowlist; no arguments or paths

API through Gateway

Tokens, cost, attribution, and per-call decision

Measured

Traffic through the Veltrix endpoint

Local presence never becomes tokens, cost, or quota. Complete financial measurement only exists in traffic routed through Gateway.

Start where you are. Move up when you need control.

Conversion does not depend on promising savings before there is measurable traffic.

Lens · discover

Map browser, desktop, and terminal with explicit sources.

Install Lens

Gateway · measure and govern

Route API calls to measure real cost and apply rules.

See Gateway pricing

Teams · respond together

Give spending, policies, and team coverage an owner.

See Teams
05 · against the hyperscaler

They optimize inside their own cloud. Nobody shows the spend between clouds.

Fighting Bedrock or Foundry on routing would be losing on purpose. Click each row — the source is their own documentation.

AWS answers "No" to per-prompt cost in Cost Explorer. Minimum granularity is usage type, per day. With Veltrix it is a database row — auditable.

AWS · Bedrock cost mgmt FAQ ↗

Microsoft writes that Azure OpenAI has no hard cap. AWS Budgets blocks via IAM/SCP, hours late, and denies everything — not per tenant.

Microsoft Learn · Azure OpenAI quotas ↗

AgentCore does call OpenAI and Gemini — but the docs say the provider bills you directly. That spend never reaches Cost Explorer.

AWS · Bedrock AgentCore docs ↗

AWS's own FAQ answers: "Not from the Amazon Bedrock side. […] Enforce tagging in a shared client or LLM gateway." In other words: they tell you to build exactly what Veltrix already is.

AWS · Bedrock cost mgmt FAQ ↗
06 · the audit

Every number has a query behind it. Including when the answer is "we don't measure that".

request_logs · 30d
SELECT count(*)                               AS chamadas,
       round(sum(estimated_cost::numeric), 4) AS pago_usd,
       round(sum(cost_saved::numeric), 4)     AS economizado_usd
FROM   request_logs
WHERE  created_at > now() - interval '30 days'
  AND  cache_hit = false
  AND  status_code < 400;
the four states of a number
  • measured
    non-null column over a declared base
  • counterfactual
    comparison against a scenario that did not happen
  • not measured
    the sensor did not write — absence, never a zero-height bar
  • not eligible
    we measured, and the lever correctly did not act

"There was nothing to compress" is success; "the sensor went down" is an incident. In most tools both become the same green 0%.

07 · pricing

Flat subscription. No percentage of what you spend.

Your budget line doesn't move when traffic grows, and all the savings stay with you.

Lens

R$ 0
chrome extension

A capacity wallet per provider, with source and confidence per reading.

Gateway

main
R$ 179
/ month

Cache, routing across 9 providers, a cap that blocks, cost per request and per feature, alerts on Slack, Discord and webhook.

Teams

R$ 219
/ seat · min. 3

Everything in Gateway, per seat, with roles (admin, dev, finance), budgets and quota approval.

next step

Start with the audit. It's yours.

Four weeks measuring your real traffic, without changing a line of your application. At the end, an auditable report for whoever decides — including if the conclusion is that you don't need us.

veltrix © 2026model call orchestration · 9 providers + your vLLM