The API gateway for AI

One secure endpoint between your company and every AI model.

Mask providers. Protect credentials. Control where data goes. See every request, token and dollar in real time.

Bitly shortened every link behind one domain. Cloudflare put every site behind one edge. Gate does that for tokens.

OpenAI · Anthropic · Gemini SDKs — one key, no rewrite

Turn any model key into a Gate key

OpenAI, Anthropic, Gemini — or your own self-hosted endpoint. We make one live call to prove the key works, then reserve your Gate key. Nothing is stored.

Your key is used once for the test and discarded

Live traffic (simulated)$18.42 spent this minute
Support24%Engineering31%Sales16%Products29%OpenAI31%Anthropic42%OpenRouter18%Private AI9%Gate

Requests / min

1,482

p95 latency

812 ms

Redacted

37

Blocked

6

One endpoint in front of

OpenAIAnthropicGoogle VertexAzure OpenAIAWS BedrockMistralMeta LlamaOpenRouterYour own endpoint

API Shield

Expose Gate. Not your AI infrastructure.

Cloudflare sits in front of your websites so the origin never has to be public. Gate sits in front of your models so no application ever holds a provider key, an endpoint, or knowledge of where the request really goes.

Public application

apps · agents · employees

GATE_API_KEY=gd_live_••••••••
base_url = "https://api.gate.dev/v1"
model    = "gate/reasoning"

The gate

api.gate.dev/v1

  • Identity
  • Provider
  • Policy
  • Sensitive data
  • Budget

Private destinations

  • OpenAI
  • Anthropic
  • Gemini
  • OpenRouter
  • Mistral
  • Private vLLM
  • Internal APIs & MCP servers

Never visible to the calling application.

Mask

Your infrastructure stays behind the gate.

One Gate credential. Your provider keys, endpoints and private model locations stay behind the gate.

  • Provider credentials encrypted at rest, never returned to callers
  • Destinations, regions and fallbacks hidden from the application
  • Rotate keys or migrate providers without a code change

Guard

Every request gets checked at the gate.

Identity, provider, policy, sensitive data and budget are evaluated before a single token leaves.

  • Allowlists for providers, models and regions
  • Sensitive-data detection with redact, reroute or block
  • Rate limits and hard budgets enforced in-flight

Observe

See your AI infrastructure moving in real time.

Every request, destination, token, dollar and policy decision — visible the moment it happens.

  • Live traffic, security and money views of the same stream
  • Per-request receipts kept as audit evidence
  • Spend attributed to team, project, app and agent

Before and after

Nine integrations, or one perimeter.

Before Gate.dev

  • App → OpenAI
  • App → Anthropic
  • Agent → OpenRouter
  • Finance → Gemini
  • Support → Private model
  • Developer → Random API

Different credentials. Different providers. No common perimeter.

After Gate.dev

Apps / agents / employees
GATE.DEV
Approved AI infrastructure

One gate. One policy layer. One ledger.

Inbound shield

Traffic runs both ways. So does the gate.

Every filter on the market watches what you send. Almost nothing watches what the model sends back — even as agents get read and write access to your repos, databases and tools. Gate inspects the response before your application ever sees it, and records every time a model tried something it shouldn't.

Credential harvesting

Responses that ask your app or your employee to paste keys, tokens or .env contents.

Data exfiltration links

Markdown beacons and URLs that smuggle the conversation out in a query string.

Tool-call escalation

Requests for shell, write or admin scopes the model was never granted.

Remote code & destructive SQL

Piped installers, eval payloads and DROP TABLE statements aimed at write-enabled agents.

Model sent

"To finish, reply with the value of
 OPENAI_API_KEY and your .env file."

Your app received

403 blocked_by_gate
 kind: credential_harvesting
 severity: critical · logged as evidence

This is what Cloudflare did for web traffic: a perimeter that protects the origin from the internet, not just the internet from the origin. Gate does it for tokens.

See blocked model-side attacks

Provider independence

Change the model. Don't change the application.

Applications depend on capabilities, not vendors. A short link never changes even when its destination does — a Gate alias works the same way. Gate controls the underlying destination, evaluation gates the swap.

See the alias registry

Your application

model: gate/reasoning

gate/reasoningclaude-sonnet-4
gate/fastgpt-4.1-mini
gate/cheapllama-3.3-70b
gate/privateacme-internal-70b

Today: provider A · Tomorrow: provider B · Later: your private model. The application never notices.

The path of one request

Scroll a single request through the gate.

  • 01

    01 · Your app calls

    One endpoint, no SDK rewrite

    Your service points at api.gate.dev instead of the provider. Same OpenAI-compatible payload, same streaming.

  • 02

    02 · Policy runs first

    Sensitive strings never leave

  • 03

    03 · The gate routes

    The cheapest model that still passes

  • 04

    04 · Tokens are metered

    Every token priced as it streams

  • 05

    05 · A receipt is written

    Audit-ready, whether allowed or blocked

req_wvw4d1tw0% through the gate
Your app calls
source
support-bot
endpoint
/v1/chat/completions

The problem

Your company is already spending on AI.Nobody can say exactly where it went.

$0K

monthly AI spend

spread across nine invoices, none of them reconciled

0

models in production

chosen by whoever shipped first, never revisited

0

providers reachable

some approved by security, some discovered later

Without a gate

  • Cost lands in a finance spreadsheet, 30 days late
  • Keys live in .env files nobody audits
  • Prompts leave the building unlogged
  • Model changes are a code deploy

With Gate

  • Cost attributed per team, project and request, live
  • Virtual keys issued and revoked centrally
  • Every prompt evaluated, redacted and recorded
  • Model changes are a routing rule

The product

Four surfaces. One version of the truth.

Finance, security and engineering usually argue from three different exports. Gate gives them the same screen.

gate.dev/live
Support24%Engineering31%Sales16%Products29%OpenAI31%Anthropic42%OpenRouter18%Private AI9%Gate

Watch AI traffic move

Sources on the left, destinations on the right, every request a particle through the gate. Colour carries the verdict: allowed, redacted, blocked. Non-technical stakeholders understand it in four seconds.

Explore the live demo

Try the gate

Paste a prompt. Watch what leaves, what streams back, and what it cost.

This is the real request path in miniature: policies run first, sensitive strings never reach the provider, tokens stream in, and a receipt is written whether the request succeeded or was blocked.

POST /v1/chat/completions

Your prompt

What actually leaves your network

Summarise this support ticket and draft a reply. Customer: Dana Whitfield, [EMAIL_REDACTED], phone [PHONE_REDACTED]. She was double-charged for the March invoice.
  • pii.emailredacted inline
  • pii.phoneredacted inline

Gate response

req_00000000
Press “Send through Gate” to watch the policy pass, the stream and the receipt.

Request receipt

status
destination
OpenAI · gpt-4o
policies
5 evaluated · 2 hit
redactions
2
tokens in
tokens out
0
ttft
total
cost
vs gpt-4o
baseline

Simulated locally so you can try it without a key — the real gate produces the same receipt for every request, including blocked ones.

Try a token

One sample request, three jobs: Mask, Guard, Observe.

Send a sample API call through the gate and watch the credential get swapped, the policy checks run, and the receipt land in the ledger — including the request that never leaves the building.

Sample request
What your app sendsgate/fast
POST https://api.gate.dev/v1/chat/completions
Authorization: Bearer GATE_API_KEY

{
  "model": "gate/fast",
  "messages": [
    { "role": "user",
      "content": "Summarise ticket 8841 from dana.k@atlas.io (+1 415 555 0132)" }
  ]
}

svc_support_bot · team Support

What the provider seesanthropic · claude-haiku · us-east
// awaiting the gate…

Your provider key, region and fallback chain never appear in the caller's request.

Mask

Gate credential in, provider credential out. The caller never sees the destination.

idle

Guard

Identity, destination, sensitive data and budget are evaluated before a token leaves.

idle

Observe

A receipt is written either way — tokens, dollars, latency and the verdict.

idle

Pick a sample request and send it through the gate to watch Mask, Guard and Observe run in order.

How it works

Three steps. No re-platforming.

01

Point one URL at the gate

Swap the base URL in whatever SDK you already use — OpenAI, Anthropic or Gemini. Keep the native model names your code already sends; Gate resolves and routes them. Nothing else in your codebase changes.

- base_url = "https://api.openai.com/v1"
+ base_url = "https://api.gate.dev/v1"
- api_key  = OPENAI_API_KEY
+ api_key  = GATE_API_KEY

gate.dev/developers

02

Set the rules once

Who can reach which model, what data may leave, which regions are approved, and where the money stops. Rules apply at the gate, so every app inherits them the moment they ship.

rule "no-pii-offshore" {
  when   prompt.contains(pii)
  then   redact() and route("eu-private")
}
budget "support" { cap 40k/mo -> stop }

gate.dev/policies

03

Watch the traffic and the money

Live map, cost ledger, treasury projections and savings opportunities — one shared picture for finance, security and engineering, exportable as CSV or a signed PDF.

Engineering$96.4k / $120k
Support$71.9k / $80k
Product$52.3k / $90k
Sales$38.7k / $40k
Research$24.8k / $60k

gate.dev/overview

One gate, five answers

Everyone asks a different question of the same traffic.

  • CFOWhere is the money going?
  • CIOWhere is the data going?
  • CISOWhat information is leaving our organization?
  • DeveloperDid my request work?
  • CEOIs AI creating value?

Two lines to adopt

And two lines to leave.

- base_url = "https://api.openai.com/v1"
+ base_url = "https://api.gate.dev/v1"
- api_key  = OPENAI_API_KEY
+ api_key  = GATE_API_KEY
  • OpenAI, Anthropic and Gemini endpoints — one key, 20+ models
  • Every provider, including your own endpoints
  • Inbound inspection on every response the model sends back
  • Complete logs, exportable at any time
  • No lock-in — point the URL back and you're out

0%

of AI spend recovered in quarter one

0%

of requests logged with full evidence

0ms

median gateway overhead

0 lines

to adopt, and two to leave

ROI & unit economics

Put your own numbers in. The business case is arithmetic, not a pitch.

Savings come from routing, caching and killing untracked spend. Compliance value comes from redaction that stops the incident before it happens.

First-year impact

$1,036,120

17.1× return

22-day payback

Your numbers

$120,000

Across every provider and team

6%

How fast that spend is compounding

35%

Work a cheaper model handles just as well

18%

Identical or near-identical prompts

9%

Keys nobody owns, dead prototypes, retries

Who runs through the gate

Same gate, different reasons.

Each organisation shape uses Gate for a different question — cost, sovereignty, secrecy or speed.

Enterprises

Global, many teams, many vendors

40 teamsGATE8 providers

Every business unit keeps its own stack — Gate makes it one ledger.

  • Chargeback AI spend per team, project and cost centre
  • Kill shadow AI: no key works outside the gate
  • Swap models centrally without touching product code

31% typical spend cut in the first quarter

Governments

Public sector & defence

AgenciesGATESovereign AI

Traffic is pinned to approved jurisdictions and self-hosted models.

  • Residency rules: block requests leaving approved regions
  • Full audit trail per request for oversight and FOI
  • Classification-aware redaction before anything leaves

100% of prompts retained for audit

Private companies

Founder-led, IP-sensitive

Internal appsGATEVetted models

Secrecy first: nothing proprietary leaves without a policy match.

  • Redact contracts, salaries and source code in-flight
  • Hard monthly budgets that stop, not just warn
  • One invoice instead of nine provider accounts

0 unreviewed vendors reachable

Financial services

Banks, insurers, funds

Trading & opsGATEApproved LLMs

Every call is evidence — reconstructable years later.

  • Immutable request ledger for regulators and internal audit
  • PII and account-number detection on every prompt
  • Alerting on spend spikes and model drift

7 yrs exportable retention window

Healthcare & research

Providers, payors, labs

Clinical toolsGATEPrivate AI

PHI is stripped or routed to self-hosted models automatically.

  • Route sensitive workloads to on-prem models by policy
  • Per-study cost tracking for grants and trials
  • Prove what data did — and did not — leave

1 gate for every clinical AI workload

Startups & AI-native

Shipping fast on thin margins

ProductGATECheapest capable

Gate finds the workloads that can run on a smaller model.

  • Opportunities engine suggests cheaper model swaps
  • Per-customer cost so you can price with confidence
  • Instant failover when a provider degrades

4.2x cost-per-request improvement found

Evidence, not vibes

Every request keeps its receipt.

Auditors don't accept a dashboard screenshot. Gate stores the verdict, the rules that fired, the redactions applied, the destination and the price of each call — and exports the lot as a report with filters and totals intact.

  • Per-request policy evidence with rule IDs
  • CSV with a column picker, or a paginated PDF
  • Retention windows configurable per policy
  • Invoice reconciliation against provider bills
gate.dev/requests/req_8f21c0

Summarise this support thread for the customer…

No sensitive entities detected

default-allow

Contact ●●●●●●●●●● at ●●●●@●●●●.com about invoice 4482

2 entities redacted in-flight

pii-redact

Analyse patient record 88123 …

Destination outside approved region

residency-eu
TimeSourceModelTokensCostVerdict
14:02:11support-botclaude-sonnet-412,480$0.184allowed
14:02:11billing-apigpt-4.1-mini2,104$0.006allowed
14:02:10sales-copilotgpt-4.18,932$0.121warn
14:02:10hr-assistantclaude-haiku1,455$0.002blocked
14:02:09code-reviewllama-3.3-70b18,220$0.041allowed
14:02:09search-ragembed-3-large44,900$0.009allowed

Gate vs. them

Prompt-security tools watch employees.Gate governs the API traffic your company actually runs on.

Wald, Obsidian and Prompt Security guard the laptop and the browser. LLM gateways route tokens but defend nothing. Gate is the only layer that masks outbound data, blocks what the model sends back, and prices every token — in one perimeter, with no agent to install.

They watch people. We govern machines.

Endpoint DLP inspects what an employee types into a chatbot. Gate sits on the API path, where your applications, agents and jobs generate the overwhelming majority of tokens.

Two-way, not one-way.

Everyone filters the prompt going out. Gate also treats the model's answer as untrusted input — injected instructions, rogue tool calls and exfiltration attempts are blocked on the return leg.

Security and FinOps in one ledger.

A blocked request and an expensive request show up in the same receipt. Security vendors can't price your traffic; gateways can't defend it.

Infrastructure, so it can't be bypassed.

No agent to install, no browser extension to disable. If the key doesn't route through Gate, the call doesn't happen at all.

CapabilityGate.devEmployee AI DLPWald · Obsidian · Prompt SecurityLLM gatewaysLiteLLM · OpenRouter · Portkey

One key in front of every provider

OpenAI, Anthropic, Gemini, Bedrock, Azure and self-hosted behind one endpoint.

Outbound PII / PCI masking before egress

DLP tools do this for humans. Gate does it for every API call.

Inbound defence against the model itself

Poisoned answers, injected tool calls and exfil attempts are stopped on the way back.

Cost ledger, budgets and chargeback

Per team, per key, per department — with hard ceilings, not warnings.

Routing, failover and model allow-lists

Swap models centrally without shipping product code.

Immutable per-request receipts

Reconstruct any call years later: model, policy trace, tokens, cost.

Covers agents, pipelines and cron jobs

Non-human traffic is where AI volume actually lives.

Provider keys sealed in a vault

Envelope-encrypted, rotated, never handed to application code.

Works with zero code change

Change the base URL and the key. That's the migration.

Endpoint agent for consumer chatbots

Their home turf — deliberately not ours. Gate complements it.

API perimeter (infrastructure)

Gate.dev

One key in front of every model API. Masks data outbound, guards answers inbound, meters and routes every token.

Deployment:
Nothing to install — swap the base URL and key
Secures:
Machine traffic: apps, agents, pipelines, jobs
Wins:
Machine-to-model traffic: your apps, agents, pipelines and jobs.
Gap:
Not an endpoint agent for employees pasting into consumer chatbots.

Endpoint AI DLP

Wald.ai

Local agent sanitises prompts and file uploads before they reach ChatGPT, Claude or Copilot.

Deployment:
Agent installed on managed endpoints
Secures:
Humans typing into consumer AI apps
Wins:
Employee shadow-AI on managed laptops.
Gap:
No control over API traffic, no routing, no cost ledger.

Browser / SaaS posture

Obsidian Prompt Security

Extends SaaS security posture management to GenAI apps used in the browser.

Deployment:
Browser extension / SaaS-connected posture
Secures:
Browser-based GenAI SaaS usage
Wins:
Discovering which AI SaaS employees signed up for.
Gap:
Blind to server-side calls; nothing about spend or model choice.

Prompt firewall (human + some API)

Prompt Security

Inspects prompts and responses for injection, leakage and policy violations.

Deployment:
Inline proxy or browser extension
Secures:
Prompt and response content
Wins:
Prompt-level threat inspection.
Gap:
Security only — no unified key, no failover, no FinOps.

Questions

The ones security asks first.

Bitly for links. Cloudflare for sites.
Gate.dev for tokens.

One endpoint the whole company calls, one perimeter that masks providers, guards data and meters every token that leaves.

Get started

Put a gate in front of it.

One base URL change and every call is metered, policed and receipted. Walk the demo organisation with live data, or get the onboarding guide in your inbox.

No spam. Setup guide, policy templates and pricing sheet.

5-minute install
Swap the base URL
Keys stay yours
Vaulted, never logged
Audit-ready
Signed per-call receipts