Built for production AI teams

OpenAI-compatible LLM gateway for lower-cost inference

Route requests across open-weight models like Qwen, DeepSeek, Llama, and Mistral with team budgets, usage logs, fallback routing, and regional deployment options.

No training on customer data Prompt logging can be disabled US/EU regional deployment
1 URL Drop-in OpenAI SDK migration path
5+ models Open-weight routes for cost and quality tiers
0 training Customer prompts are not used for model training
Team caps Budgets, API keys, logs, and model-level limits
Why this exists

LLM costs grow faster than product revenue

Most teams start with a single model API. Then usage grows, costs become unpredictable, API keys get shared across projects, and nobody knows which customer, feature, or model is driving the bill.

OpenRelay AI gives teams one OpenAI-compatible gateway to route requests, set limits, track usage, and switch models without rewriting the app.

Reduce inference cost without changing product logic.
Know which API key, project, model, and customer drives spend.
Keep sensitive routes in approved regions with clear controls.
Gateway controls

One API surface for model routing, budgets, and logs

Built for teams shipping AI products, not chatbot hobby projects. Keep provider choice flexible while giving finance and engineering one clean usage picture.

Drop-in OpenAI compatibility

Use your existing OpenAI SDK and change the base URL. Keep your app logic while routing to lower-cost models.

Model routing and fallback

Route simple tasks to cheaper models and reserve stronger models for complex requests. Add fallback when a provider is slow, down, or rate-limited.

Team budgets and API keys

Create project-level API keys, set monthly budgets, cap model usage, and stop one feature from burning the entire budget.

Usage logs without lock-in

Track requests, tokens, latency, cost, model, provider, and API key. Disable prompt and output logging when privacy matters.

Regional deployment options

Use hosted endpoints or deploy private instances in your preferred region. Keep sensitive data out of regions your customers do not approve.

Clear data handling

Customer prompts are not used for model training. China-based routes are opt-in and disclosed before use.

Qwen Cost-efficient multilingual routes
DeepSeek Reasoning and coding tiers
Llama Open-weight baseline
Mistral EU-friendly options
BYOK Bring provider keys when needed
Beta pricing

Start small, then add team controls as usage grows

Use a simple subscription plus service credits model for the beta. Credits are used only for OpenRelay AI API consumption and cannot be transferred or withdrawn.

Developer

$19 /mo

For builders testing lower-cost models.

  • OpenAI-compatible API
  • Personal API keys
  • Usage dashboard
  • Email support
Join beta

Business

$499 /mo

For teams that need controls, reporting, and reliability.

  • Unlimited projects
  • Audit logs
  • Higher rate limits
  • BYOK support
Request access

Private

Custom

For dedicated infrastructure and data residency.

  • Dedicated endpoint
  • US/EU/Singapore regions
  • Prompt logging controls
  • SLA options
Talk to us
Migration path

Switch one route first, then scale the savings

The beta is built around a practical migration review. Pick one endpoint, test cheaper routes, measure latency and quality, and then expand where the numbers work.

1

Share one endpoint

Tell us the task, current model, monthly tokens, and privacy constraints.

2

Map model routes

Test cheaper open-weight models and define fallback rules for quality.

3

Change base URL

Use your existing OpenAI SDK and point one environment at OpenRelay AI.

4

Measure real usage

Track spend, latency, success rate, and model quality before rollout.

FAQ

Clear answers before you route production traffic

Is this a chatbot?

No. OpenRelay AI is an API gateway for teams building AI products and internal tools.

Do I need to rewrite my app?

Usually no. If you already use the OpenAI SDK, change the base URL and model name.

Do you train on customer data?

No. Customer prompts and outputs are not used for model training.

Are requests sent to China?

Not by default. Hosted routes use approved providers and regions. Any route involving a China-based provider must be explicitly enabled.

Can prompt logging be disabled?

Yes. Teams can disable prompt and output logging while keeping metadata such as tokens, cost, latency, and model.

What happens if a model fails?

You can configure fallback routes to another model or provider.

Cut LLM costs without losing control

Join the beta and get migration help for your first OpenAI-compatible route. We are prioritizing teams with real production usage and clear cost visibility problems.