Products
Pricing

Network Security Pricing

Network DDoS Protection
Partners
About Myra
Login
Emergency Demo

Updated 11 October 2026

What is an AI gateway (LLM gateway)?

As of 11 October 2026

An AI gateway, also called an LLM gateway, is a reverse proxy that sits between your applications and the model APIs they call. Your code talks to one endpoint with one token. The gateway holds the provider keys, decides which model actually answers, enforces budgets and rate limits, inspects prompts and responses, and writes a log entry you can query later.

If three teams already call OpenAI, Anthropic and a self-hosted model with keys pasted into environment variables, you are the target audience. The gateway turns that sprawl into a single choke point where platform and security teams can set policy once.

The jobs a gateway takes over

The first job is credentials. Applications authenticate against the gateway with a scoped token, and the real provider keys stay in a vault on the server side. Revoking a leaked application token no longer means rotating the OpenAI key for everyone.

The second is translation. Most gateways expose an OpenAI-compatible API and convert requests into each provider's native format, so switching from GPT to Claude or Gemini becomes a change to the model string.

Then come the controls that are hard to bolt onto every client separately. Spend caps stop a runaway agent loop before the invoice does. Rate limits protect shared quotas. Routing rules send a share of traffic to another model, and fallback chains retry against a second provider when the first returns 5xx errors. Guardrails check content in both directions, from regex filters to classifiers for jailbreaks and prompt injection, which OWASP lists as the first risk in its LLM Top 10.

Finally, there is the record of what happened. A request log with model, tokens, cost and latency per call is what finance, incident response and auditors will ask you for.

Three names for one category

“AI gateway” and “LLM gateway” are used interchangeably. “LLM proxy” is the older and narrower term; it usually means the forwarding part without much policy. “AI firewall” puts the emphasis on inspection: blocking prompt injection, toxic output or data leaving the company. Some products are sold under one label and do all of it, others only cover one slice. Judge them by what sits in the request path.

Open source proxy or managed gateway

The category has well-known open source projects. LiteLLM describes its AI Gateway as “a self-hosted server” that gives applications “one OpenAI-compatible endpoint for 100+ LLM providers”. Portkey's gateway sits on GitHub under the MIT licence and calls itself “A blazing fast AI Gateway with integrated guardrails.” Kong describes its AI Gateway as a “Connectivity and governance layer for modern AI-native applications”, and according to Cloudflare, its AI Gateway lets you “gain visibility and control over your AI apps”.

A self-hosted open source proxy is often the better fit. That holds if you want to read and patch the code, if your team wants to operate the gateway itself inside its own Kubernetes cluster, or if you mainly need a routing layer in front of your own GPUs. The open source projects also cover a very wide range of providers, and because the logs sit in your own database, you decide who can read them. You are free to change every default yourself. In return you own upgrades, the database behind the logs, high availability and the security review of every plugin.

A managed gateway trades that control for operations you do not have to run, and it suits teams that want EU-hosted models and masking from the same vendor. The questions then move to the operator: where it runs, who can read the logs, and what the contract says.

How Myra's gateway handles a request

Myra AI Workspace is built around such a gateway, a multi-tenant reverse proxy that runs on Myra's own CDN infrastructure. Each tenant has one or more named gateways (production, staging), and each gateway has its own provider keys, tokens, guardrails, routing rules, budgets and rate limits. A token issued for one gateway is rejected on another.

Every request passes three phases. The access phase runs before the body is read: authentication, rate limit, budget check and an optional IP allowlist. The content phase covers the response cache, routing, request guardrails, the provider call with retries and fallback, and response guardrails. The log phase writes the structured log entry after the answer has gone out.

From an existing OpenAI client you change the base URL and the key. A minimal call to the OpenAI-compatible endpoint looks like this:

curl https://ai-api.myra.eu/v1/{tenant}/{gateway}/compat/chat/completions \
  -H "Authorization: Bearer <gateway-token>" \
  -H "Content-Type: application/json" \
  -d '{"model": "claude-sonnet-4-6",
       "messages": [{"role": "user", "content": "Hello"}]}'

The gateway resolves the provider from the model name and returns an OpenAI-shaped response, streaming included. Tokens are also accepted in the x-aig-token header, and in x-api-key for Anthropic SDK compatibility.

Behind the endpoint sit external providers such as OpenAI, Anthropic, Google Gemini and Vertex AI, AWS Bedrock, Azure OpenAI and Mistral, plus Myra's own models. Those run on Myra's own EU infrastructure; on paid self-service plans they need the “EU-Gov” add-on. For external providers their own terms apply, and Myra does not use customer data to train AI models.

Budgets exist at three levels (token, gateway, tenant) and the most restrictive one wins. They are spend caps, so “per-token budget” means per auth token, not a count of LLM tokens. Rate limits are sliding windows per gateway and per token, answered with 429 and Retry-After when exceeded.

PII masking is pseudonymisation

The guardrail that matters most for GDPR questions is the PII Protector. Before an external model sees the request, a detection engine running at Myra replaces personal data such as names or IBANs with placeholders of the form [MYRA-REDACT-{TYPE}:…]. The gateway keeps the mapping and can put the original values back into the answer. English and German are fully supported, and other Latin-script languages are detected on a best-effort basis. Masking is available in every plan.

Be precise about what this is. Under Art. 4(5) GDPR it is pseudonymisation: the mapping exists, so the data can be re-identified. Recital 26 treats pseudonymised data as personal data, so masking reduces what a provider receives without making the data anonymous.

Two details are worth checking in the configuration. The default target is "request": the request is masked, but the placeholders in the response stay as they are. Set "target": "both" to get the original values back. And fail_open defaults to true, which lets a request through unmasked if the detection service fails. Set it to false if you would rather block. A project set to “PII protection required”, or a tenant that enforces masking, refuses the request when the masker is down, whatever fail_open says. Requests that go entirely to a Myra model are not masked at all, because they stay on Myra's infrastructure.

Location is a separate switch. EU-region routing can be enforced per gateway or for the whole organisation; when armed, any route outside the EU is refused with 403 data_residency_blocked before a provider is called. It is off by default.

What Myra's gateway is not

It is not open source, and in the standard setup you do not run it yourself: Myra provisions and operates the instance as a managed service. On-premises deployment is available on request, with Myra handling licensing, hardware requirements and deployment support. If you need to fork the proxy, look at the open source projects above.

Several defaults are permissive until you change them. The request log stores prompt and response text by default (log_payloads: true). Set log_payloads: false on the gateway or send x-aig-collect-log-payload: false per request, and the entry with routing, status, token counts and timing is still written without the text; x-aig-collect-log: false skips the entry for that request altogether. A gateway has no rate limit until you configure one, and the response cache is off (cache_ttl is 0). Caching is exact-match only, since semantic caching is currently disabled. A token budget also isn't a hard ceiling for a single request: calls already in flight when the cap is reached may still complete at their real cost.

The service runs on Myra infrastructure with a BSI C5 Type 2 attestation; the application itself is outside that scope. Admin changes and security events are logged in an audit log, which can be made tamper-evident with a hash chain on request. Customer admins have no screen for it: Myra's platform admins read it via API.

The full API reference, including the native provider routes and the error codes, is in the documentation. For the workspace around the gateway, see Myra AI Workspace.

Statements about third-party projects come from their own pages (see Sources), as of 11 October 2026. LiteLLM, Portkey, Kong and Cloudflare are trademarks of their respective owners. Myra Security has no affiliation with them. Nothing on this page is legal advice.

Frequently asked questions

Is an LLM proxy the same as an AI gateway?

In practice, mostly yes. “LLM proxy” stresses the forwarding: one endpoint, many providers, keys kept server-side. “AI gateway” usually adds policy on top, such as budgets, rate limits, guardrails and request logs. Vendors use both terms for the same kind of product, so look at the feature list instead of the label.

Does PII masking make my LLM use GDPR compliant?

It helps, but it does not settle the question on its own. Masking with placeholders that the gateway can restore in the answer is pseudonymisation under Art. 4(5) GDPR, and Recital 26 says pseudonymised data is still information about an identifiable person. You still need a legal basis, a data processing agreement and a decision on which providers may receive the masked text. This page is not legal advice.

Can I use the Anthropic SDK or Claude Code with Myra's gateway?

Yes. Anthropic-native clients use the passthrough route /v1/{tenant}/{gateway}/anthropic/v1/messages, which speaks the Messages API in both directions. For Claude Code you set ANTHROPIC_BASE_URL to that route without the /v1/messages suffix and use a gateway token as the key. The OpenAI-compatible /compat route does not accept Messages requests and answers them with 400.

Where do the models run, and is my data used for training?

That depends on the model you route to. Myra models run on Myra's own EU infrastructure; on a paid self-service plan they need the “EU-Gov” add-on. External providers run on their own infrastructure under their own terms. Myra does not use customer data to train AI models.

Sources