Kong gives platform teams a place to govern the traffic between applications, AI models and agent tools. An AI gateway can centralize credentials, routing and usage controls that would otherwise be repeated inside many applications. The buying decision is whether that shared layer makes access and operations clearer without obscuring the differences between model providers, tool permissions and application responsibilities.
- 01The offer API and AI connectivity infrastructure, with Konnect-managed and self-hosted gateway routes.
- 02The fit Teams operating several AI applications that need common access, routing and usage controls.
- 03The boundary Public-source research and a proposed gateway pilot; no traffic or provider credentials were connected.
01 / ProductKong treats AI as traffic that needs identity and policy
The AI Gateway introduction describes a connectivity layer for LLM, Model Context Protocol and agent-to-agent traffic. These interactions share operational needs such as authentication and observability, but their meanings differ. A model request generates output; an MCP invocation may read or change a business record. The gateway configuration must preserve that distinction.
AI Gateway 2.0 provides an entity model managed through Konnect and a documented self-hosted route. On-premises Kong Gateway runs the corresponding Services, Routes, plugins and Consumers. Teams can configure those primitives directly or translate AI Gateway entities with decK. The management model differs even where the traffic-handling primitives are shared, so deployment choice should precede configuration work.
Kong's company page describes its wider API portfolio, including Insomnia and the acquisition of OpenMeter. The AI gateway is part of that broader connectivity portfolio. Existing API management customers may be evaluating an extension to familiar infrastructure, while new customers must also account for the gateway operating model.
02 / AudienceThe strongest fit is a platform serving several AI application teams
Imagine several internal services calling different model providers directly. Each team manages credentials, chooses fallback behavior and tries to attribute cost. That can work at small scale, but shared requirements become hard to enforce consistently. Kong is relevant when a platform team already owns the common access policy and can operate the gateway as a service for application teams.
The Cloudflare blueprint offers an infrastructure comparison for teams considering an edge platform and its AI tooling. The Datadog blueprint is relevant when the immediate problem is visibility across application behavior. Gateway governance and observability overlap, but measuring a request does not necessarily control whether it is allowed to reach a provider.
A single experimental application may not need this extra layer immediately. Adding a gateway creates another component whose availability and configuration matter. Start from a concrete shared requirement, such as provider-key rotation or department-specific access. If no one owns those requirements, centralization can simply concentrate unclear policy in a new place.
03 / WorkflowA proposed shared endpoint pilot should test both allowed and denied calls
For a proposed pilot, select two internal applications with different permitted model access. Give each a distinct consumer identity and route only synthetic requests through the gateway. Keep the first provider and model fixed so the team can verify authentication, usage records and response behavior before adding fallback or optimization policies.
Configure the model provider centrally and keep application-facing identifiers stable. The load-balancing guide describes model targets, routing and several balancing strategies. In the proposed pilot, begin with a simple priority arrangement only after the baseline works. A sophisticated routing rule is harder to debug when the underlying permissions and response contract are still uncertain.
Test that the first application can reach its approved model and cannot reach the restricted alternative. Repeat using the second identity. Record which policy produces a denied response and what the application displays to its user. An access-control rollout is incomplete if every tested request succeeds; the negative case establishes whether the intended boundary is real.
Next, interrupt the primary provider in a controlled test and inspect the fallback. Compare structured-output handling, tool-call behavior and error responses. Two models can accept similar requests while producing meaningfully different outputs. The gateway can route around an outage, but the application must decide which fallback is acceptable for the task and whether the user should be told about degraded behavior.
Add a conservative consumption policy. The rate-limiting documentation distinguishes counter strategies and their operational tradeoffs. A local counter has different consistency characteristics from a shared one. Validate enforcement across the deployment topology actually chosen, including simultaneous requests, rather than assuming a configured number is an exact global ceiling.
If the pilot also exposes a backend API as an MCP tool, create that route separately. The AI MCP Server reference explains that MCP traffic is API-level traffic and that prompt-oriented LLM policies do not automatically apply to it. Test authentication and tool access at this boundary. A prompt guard cannot substitute for a backend authorization check.
Finish with provider-key rotation and a configuration rollback exercise. Application teams should be able to explain what endpoint they call, what identity the gateway sees and where a failed request can be traced. The proposed acceptance criteria are operational: correct access, understandable failures and predictable recovery. No latency or cost improvement is assumed from the gateway's presence alone.
04 / PricingGateway charges and model-provider charges need separate estimates
| Route or component | Published basis | Planning implication |
|---|---|---|
| AI Management Plus | From $25/month plus usage, monthly billing | Starting price is not the complete topology cost |
| Enterprise | Custom annual pricing | Confirm contracted deployment and policies |
| Gateway topology | Serverless, hybrid or dedicated cloud components | Estimate the selected control planes and traffic |
| Fully self-hosted | Separate Gateway Enterprise offer | Obtain explicit commercial and support scope |
Commercial structure from Kong pricing, consulted 3 October 2026. USD headline pricing excludes applicable usage, deployment and add-on components.
The Kong pricing page advertises AI Management Plus from US $25 per month plus usage, and custom annual Enterprise pricing. The comparison includes different control-plane and deployment charges, transaction allowances and model-proxy costs. The headline is therefore a starting component, not the complete price of every deployment topology.
The public table distinguishes serverless, hybrid and dedicated cloud routes; fully self-hosted gateways are described as a separate Gateway Enterprise offer. Establish the intended architecture before comparing quotes. A price for a serverless control plane cannot be transferred directly to a private deployment with a different support and operational scope.
Kong's model cost-management documentation explains how configured rates and provider token data contribute to a request cost figure. That accounting supports policy and reporting, but it is distinct from the gateway subscription. Reconcile the pilot's calculated usage against provider billing before treating it as a reliable departmental chargeback record.
Estimate traffic from both ordinary requests and operational behavior such as retries or evaluation runs. Also identify which paid policies are needed. The pricing comparison lists some AI capabilities as add-ons on Plus, so a low starting plan can become a different purchase once the required controls are included. Request a quote that names the policies and topology the pilot proved useful.
05 / DistinctionsThe distinction is consistency across model and tool connectivity
Kong's value is strongest where applications should share a stable access and governance layer. Central provider configuration can reduce repeated credential handling, while consumer identities give usage records an organizational meaning. That creates a more coherent platform service when teams otherwise solve the same operational problem in different ways.
The separation between model entities and policies is also useful. Provider rates can change independently of a department's spending rule, and routing can change independently of a client-facing alias. This gives platform teams a place to manage infrastructure concerns without requiring every application to redeploy. It still needs disciplined change review because a centralized mistake can affect many consumers at once.
MCP support extends the scope beyond text generation. A company may need to govern the tools an agent calls as carefully as the models it uses. The API-level distinction in the documentation is valuable here: a common gateway does not mean every traffic type should receive an identical policy stack or be evaluated using the same definition of success.
06 / QuestionsCompatibility and policy precision are consequential limits
The first question is provider compatibility for the application features you actually use. Streaming, structured output and tool calling deserve specific checks. A standardized request interface simplifies integration, but it cannot guarantee equivalent model behavior. Preserve a small application fixture set and run it against every approved fallback before changing a production routing rule.
The second is the precision of cost enforcement. Provider-reported tokens, configured rate tables and distributed counter synchronization affect the result. A usage alert or rate limit is not automatically a contractual maximum spend. Identify how concurrent requests, in-flight responses and rate changes are handled in the chosen topology before promising a hard budget to another team.
The third is feature coverage across deployment routes. The current on-premises guide specifies AI Gateway 2.0 and explains that entity mappings are not always one-to-one. For example, an AI MCP Server in upstream-server mode has no self-hosted equivalent. Check required entities and policies against that mapping, then confirm plugin licensing and support in the commercial agreement. A private-network requirement is compatible with a documented route, but does not establish feature parity.
07 / DecisionAdopt the gateway around a shared policy your teams need
Kong is worth evaluating when several AI applications require common access, routing and consumption controls. Pick one requirement that is difficult to enforce today, then demonstrate both its successful path and its deliberate failure. A useful pilot should make an application team's responsibilities clearer and give the platform owner evidence that the common policy is actually applied.
Add model diversity, MCP tools and advanced policies only after the baseline is understandable. The aim is a service that teams can depend on and diagnose, with a bill that reflects the chosen deployment. More routing options are valuable when they solve a measured operating problem; they are not a substitute for clear application contracts.
Several teams call model providers
Pilot separate identities and prove that restricted models are denied.
Agent tools reach business APIs
Test MCP authentication and backend authorization separately from prompt policies.
Private deployment requirement
Map required entities and policies to the supported self-hosted route, then confirm licensing and support.
A business worth understanding.
Suggest your business or one you find interesting. Tell us what you want to understand about its product, positioning, design or workflows.
Suggestions are free. Selection and publication stay with the desk.
- Kong AI Gateway introductionConsulted
- About KongConsulted
- AI Gateway load balancingConsulted
- AI Rate Limiting AdvancedConsulted
- AI MCP ServersConsulted
- Kong pricingConsulted
- AI Gateway model cost managementConsulted
- Configure Kong AI Gateway on-premConsulted


