Our cloud
One line in your client. Nothing to deploy.
22 models, 9 providers. Veltrix decides who answers each call, measures the cost of the request, and refuses it when the budget runs out.
You turn it on by swapping one line. And you start by measuring: a 4-week audit, no cost and no card.
That gap is what makes orchestration worth it. Each dot is an active model, placed by its real price. Your own endpoint joins as one more provider.
One line in your client. Nothing to deploy.
Your GPU answers; the cap and the attribution still come from us.
The embedding comes from a local model, not from OpenAI. Never run end to end on a clean machine — and we say so.
The fallback chain stops being a try/catch and becomes policy.
In the order you wrote it, never reordered by price — and every trigger is counted.
an own model has no per-token price: you declare the GPU cost and we attribute it per second occupied
the machine would cost the same hour if nobody called — idle GPU is shown separately
Coverage is not a generic savings promise. Every surface states its available signal and what that signal supports.
Estimated usage in compatible integrations
EstimatedBrowser events; varies by provider
Compatible tool in the foreground
PresenceOptional Local Collector on macOS
Recognized CLI running
PresenceLocal allowlist; no arguments or paths
Tokens, cost, attribution, and per-call decision
MeasuredTraffic through the Veltrix endpoint
Local presence never becomes tokens, cost, or quota. Complete financial measurement only exists in traffic routed through Gateway.
Conversion does not depend on promising savings before there is measurable traffic.
Map browser, desktop, and terminal with explicit sources.
Install LensRoute API calls to measure real cost and apply rules.
See Gateway pricingGive spending, policies, and team coverage an owner.
See TeamsFighting Bedrock or Foundry on routing would be losing on purpose. Click each row — the source is their own documentation.
AWS answers "No" to per-prompt cost in Cost Explorer. Minimum granularity is usage type, per day. With Veltrix it is a database row — auditable.
AWS · Bedrock cost mgmt FAQ ↗Microsoft writes that Azure OpenAI has no hard cap. AWS Budgets blocks via IAM/SCP, hours late, and denies everything — not per tenant.
Microsoft Learn · Azure OpenAI quotas ↗AgentCore does call OpenAI and Gemini — but the docs say the provider bills you directly. That spend never reaches Cost Explorer.
AWS · Bedrock AgentCore docs ↗AWS's own FAQ answers: "Not from the Amazon Bedrock side. […] Enforce tagging in a shared client or LLM gateway." In other words: they tell you to build exactly what Veltrix already is.
AWS · Bedrock cost mgmt FAQ ↗SELECT count(*) AS chamadas,
round(sum(estimated_cost::numeric), 4) AS pago_usd,
round(sum(cost_saved::numeric), 4) AS economizado_usd
FROM request_logs
WHERE created_at > now() - interval '30 days'
AND cache_hit = false
AND status_code < 400;"There was nothing to compress" is success; "the sensor went down" is an incident. In most tools both become the same green 0%.
Your budget line doesn't move when traffic grows, and all the savings stay with you.
A capacity wallet per provider, with source and confidence per reading.
Cache, routing across 9 providers, a cap that blocks, cost per request and per feature, alerts on Slack, Discord and webhook.
Everything in Gateway, per seat, with roles (admin, dev, finance), budgets and quota approval.
Four weeks measuring your real traffic, without changing a line of your application. At the end, an auditable report for whoever decides — including if the conclusion is that you don't need us.