The unified control plane for your AI stack. Route to any model. Govern every tool call. Enforce hard spend limits. Audit everything.
# one endpoint, any model curl https://gateway.heimos.ai/v1/chat/completions \ -H "Authorization: Bearer $HEIMOS_KEY" \ -d '{ "model": "claude-sonnet-4-5", "messages": [{"role": "user", "content": "hello"}], "fallback": ["gpt-5", "gemini-2.5-pro"] }' # automatic failover · hard spend caps · guardrails · full audit trail 200 OK · X-Gateway-Overhead-Us: 1310 · $0.0004 · team:engineering · app:chatbot
ROUTE MODELS AND TOOLS THROUGH A SINGLE POLICY PLANE
Every gateway routes models. HeimOS is the only one that routes MCP tool calls through the same policy engine — same auth, same quotas, same guardrails, same audit trail. Your agents are governed the way your completions are.
one policy engine, both paths
One integration gives you routing, governance, observability, and security — without touching your application code.
One OpenAI-compatible endpoint for every provider. Drop in your existing SDK calls — models from OpenAI, Anthropic, Gemini, and self-hosted Ollama respond through the same interface. Zero code changes.
Automatic failover, latency-based routing, and weighted load balancing. Define fallback chains per model, so one provider's outage never becomes yours.
Spend is enforced atomically on the hot path — reserved before the request, reconciled after. A cap is a cap. No overshoot in the lag window, no surprise bills.
PII redaction, prompt injection defense, and content policy. Redaction gates are sequential — sensitive data never leaves the trust boundary. Fail-closed by default.
Native Model Context Protocol support with tool-level RBAC and argument scanning. Streamable HTTP transport. Your agents get the same spend controls and audit trail as your completions.
Every request logged with full metadata — model, tokens, cost, latency, team, app, user. Queryable, exportable, retained per policy. Built for teams who need to answer "who used what, when, and why."
Every other gateway makes you trade speed for control. HeimOS runs auth, quotas, guardrails, and audit on one box — and still gets out of the way.
Numbers from a pinned, reproducible k6 rig — run it yourself. See the full head-to-head benchmark →
HeimOS sits between your apps and your AI providers. Real-time config changes, zero downtime — you won't even notice it's there.
Change one base URL. Your existing OpenAI-compatible code keeps working — no SDK swap, no rewrite.
Every request passes through the gateway. Authentication, rate limiting, spend checks, and guardrails apply instantly. Requests that violate your policies are rejected before they cost you money.
The router picks the best provider based on your rules — primary model, fallback chain, latency targets, cost optimization. If a provider is down, traffic shifts automatically.
Request metadata, token counts, costs, latencies, and guardrail decisions are captured automatically. Durable, queryable, and built for compliance — not just dashboards.
Full applicability of the EU AI Act lands August 2, 2026. Logging, transparency, and data-protection obligations apply to every AI system you run. HeimOS ships them as configuration, not consulting.
Retention policies set per tenant, enforced in the audit store, with retention proof for your auditors.
Preconfigured guardrail bundles mapped to the Act's logging and transparency articles. Fail-closed by default.
Active policies, violations, retention proof, and full audit-trail export — generated on demand.
A hierarchical multi-tenant model designed for real organizations. Org admins set global policies, team leads manage their own budgets, and individual apps get their own API keys with fine-grained permissions.
No markup on model costs — you pay providers directly. HeimOS charges only for the gateway and governance layer.
Every team building with AI hits the same wall: fragmented provider APIs, no visibility into spend, no guardrails, and no audit trail. We've been there — managing keys across providers, waking up to surprise bills, wiring up custom logging for every new model.
HeimOS is the infrastructure layer we wished existed. Named after Heimdall, the all-seeing guardian of the Bifrost, HeimOS stands between your applications and the AI providers they depend on — routing traffic, enforcing policy, and keeping a complete record of everything that passes through.
Create your org, mint a key, and route your first request in under 5 minutes.