one gateway.
every model & tool.
fully governed.

The unified control plane for your AI stack. Route to any model. Govern every tool call. Enforce hard spend limits. Audit everything.

rust 2ms added latency, ~5K req/s on a single machine. see the benchmark →
terminal
# one endpoint, any model
curl https://gateway.heimos.ai/v1/chat/completions \
  -H "Authorization: Bearer $HEIMOS_KEY" \
  -d '{
    "model": "claude-sonnet-4-5",
    "messages": [{"role": "user", "content": "hello"}],
    "fallback": ["gpt-5", "gemini-2.5-pro"]
  }'

# automatic failover · hard spend caps · guardrails · full audit trail
200 OK · X-Gateway-Overhead-Us: 1310 · $0.0004 · team:engineering · app:chatbot

ROUTE MODELS AND TOOLS THROUGH A SINGLE POLICY PLANE

UNIFIED GOVERNANCE

one policy plane for models and tools

Every gateway routes models. HeimOS is the only one that routes MCP tool calls through the same policy engine — same auth, same quotas, same guardrails, same audit trail. Your agents are governed the way your completions are.

HeimOS chatbots agents copilots LLM PROVIDERS OpenAI · Anthropic · … SELF-HOSTED Ollama · vLLM · TGI MCP SERVERS tools · data · actions + any OpenAI-compatible
authquotasguardrailsaudit

one policy engine, both paths

FEATURES

everything your AI stack is missing

One integration gives you routing, governance, observability, and security — without touching your application code.

universal api

One OpenAI-compatible endpoint for every provider. Drop in your existing SDK calls — models from OpenAI, Anthropic, Gemini, and self-hosted Ollama respond through the same interface. Zero code changes.

smart routing

Automatic failover, latency-based routing, and weighted load balancing. Define fallback chains per model, so one provider's outage never becomes yours.

hard spend guarantees

Spend is enforced atomically on the hot path — reserved before the request, reconciled after. A cap is a cap. No overshoot in the lag window, no surprise bills.

guardrails engine

PII redaction, prompt injection defense, and content policy. Redaction gates are sequential — sensitive data never leaves the trust boundary. Fail-closed by default.

mcp gateway

Native Model Context Protocol support with tool-level RBAC and argument scanning. Streamable HTTP transport. Your agents get the same spend controls and audit trail as your completions.

durable audit

Every request logged with full metadata — model, tokens, cost, latency, team, app, user. Queryable, exportable, retained per policy. Built for teams who need to answer "who used what, when, and why."

GOVERNANCE WITHOUT THE TAX

control you won't feel

Every other gateway makes you trade speed for control. HeimOS runs auth, quotas, guardrails, and audit on one box — and still gets out of the way.

2ms
your users never feel it
Governance costs less than a network hop.
~5K req/s
one machine, your whole org
Skip the capacity planning meeting.
225MB
no cluster required
A single binary on hardware you already run.
100%
nothing slips past
Every model and tool call, governed and audited.

Numbers from a pinned, reproducible k6 rig — run it yourself. See the full head-to-head benchmark →

HOW IT WORKS

up and running in minutes

HeimOS sits between your apps and your AI providers. Real-time config changes, zero downtime — you won't even notice it's there.

01

point your sdk at heimos

Change one base URL. Your existing OpenAI-compatible code keeps working — no SDK swap, no rewrite.

02

policies apply at the edge

Every request passes through the gateway. Authentication, rate limiting, spend checks, and guardrails apply instantly. Requests that violate your policies are rejected before they cost you money.

03

requests route intelligently

The router picks the best provider based on your rules — primary model, fallback chain, latency targets, cost optimization. If a provider is down, traffic shifts automatically.

04

everything gets logged

Request metadata, token counts, costs, latencies, and guardrail decisions are captured automatically. Durable, queryable, and built for compliance — not just dashboards.

EU AI ACT — ENFORCEMENT BEGINS AUG 2, 2026

compliance lives at the gateway

Full applicability of the EU AI Act lands August 2, 2026. Logging, transparency, and data-protection obligations apply to every AI system you run. HeimOS ships them as configuration, not consulting.

30 / 90 / 365d

per-tenant audit retention

Retention policies set per tenant, enforced in the audit store, with retention proof for your auditors.

policy templates

redaction & logging conformity

Preconfigured guardrail bundles mapped to the Act's logging and transparency articles. Fail-closed by default.

one-click export

auditor-facing reports

Active policies, violations, retention proof, and full audit-trail export — generated on demand.

ENTERPRISE READY

built for teams, not just developers

A hierarchical multi-tenant model designed for real organizations. Org admins set global policies, team leads manage their own budgets, and individual apps get their own API keys with fine-grained permissions.

Organization → Team → App → API Key hierarchy
Role-based access control
Per-team and per-app spend limits
Real-time config sync — no restarts needed
Full admin UI for non-technical team leads
SSO and SCIM provisioning (enterprise)
ORGANIZATION
Acme Corp
TEAM
Engineering
$2,400 / $5,000 budget
TEAM
Product
$890 / $2,000 budget
PRICING

start free. scale when ready.

No markup on model costs — you pay providers directly. HeimOS charges only for the gateway and governance layer.

DEVELOPER
$0
free forever
  • 10,000 requests/month
  • 3-day log retention, 30-day metrics
  • Universal API gateway
  • Automatic fallbacks & load balancing
  • Observability — logs, traces, filters
  • Deterministic guardrails
  • Community support
Start for free
ENTERPRISE
Custom
annual contract
  • Everything in Production
  • Custom log & metrics retention
  • SSO / SAML / SCIM
  • Dedicated support + SLA
  • Private cloud deployment
  • VPC hosting & data isolation
  • SOC 2, GDPR, HIPAA compliance
  • Custom BAA signing
Talk to sales
ABOUT HEIMOS

why we're building this

Every team building with AI hits the same wall: fragmented provider APIs, no visibility into spend, no guardrails, and no audit trail. We've been there — managing keys across providers, waking up to surprise bills, wiring up custom logging for every new model.

HeimOS is the infrastructure layer we wished existed. Named after Heimdall, the all-seeing guardian of the Bifrost, HeimOS stands between your applications and the AI providers they depend on — routing traffic, enforcing policy, and keeping a complete record of everything that passes through.

Get in touch

hello@heimos.ai

ready to ship?

Create your org, mint a key, and route your first request in under 5 minutes.

no credit card required. deploy in under 5 minutes.