Enterprise service bus, rebuilt for agentic AI

Your agents reason.
klanex executes.

The asynchronous execution layer between LLM agents and production APIs. Fire a tool-use intent, get an execution ID in milliseconds — klanex owns the retries, backoff, approvals, credentials, and audit trail, so a hallucinated JSON key or a rate limit never crashes your workflow.

Try it in your agent
in 60 seconds, no signup

Grab a free sandbox key with one command — no account, no card — and wire klanex into your agent as MCP tools. Works in Claude Code, Cursor, VS Code, and Windsurf.

Sandbox keys are disposable (Stripe test mode, auto-expire after 7 days). Want the full walkthrough, including how to make the retries fire? Open the sandbox guide →

Ready for production? Create a durable account →

Wiring an LLM straight into production APIs
breaks in three predictable ways

{ }

Hallucinated payloads

One invented JSON key and your five-step workflow dies at step three — with the API's cryptic 400 lost far from the model that caused it.

⏱

Brittle synchronous loops

An agent thinks for 30–60 seconds. Held-open HTTP connections time out, and every timeout takes the agent's operational context with it.

⚠

Naked credentials

Nobody wants raw production API keys floating through a generative environment — or an agent autonomously refunding customers unsupervised.

Decouple reasoning from execution

klanex closes the synchronous loop in milliseconds, then executes on its own terms: queued, retried, supervised, recorded.

1

Validate

Every intent is checked against your JSON Schema. Hallucinations bounce back instantly with an llm_hint the agent uses to fix itself.

→
2

Queue

Valid intents are persisted to the audit trail and queued. Your agent gets 202 + execution_id and moves on with its life.

→
3

Execute

Isolated workers make the real call with per-host circuit breakers, timeouts, and exponential backoff on 429s, 5xxs, and outages. A 200 only counts once its body agrees, and a 403 that means "slow down" gets retried.

→
4

Report

Terminal states fire an HMAC-signed webhook to your backend — or poll the API. Either way, the full history is queryable forever.

Reliability features your agents
can't provide for themselves

Schema gate with self-correction

Invalid payloads are rejected in milliseconds with an llm_hint written to be pasted straight back into the model's context.

Retries & circuit breakers

429 / 5xx / timeouts are absorbed with exponential backoff and per-host breakers. Permanent 4xxs fail fast with a correction hint.

Human-in-the-loop approvals

Flag destructive calls with requires_approval. They pause until a human approves or rejects — right from a Slack button. Approvals never carry over to replays.

Idempotency built in

Pass an idempotency_key and network retries can never double-refund, double-email, or double-anything.

Credentials in a vault

Your API keys are encrypted with Cloud KMS the moment they arrive, decrypted only in worker memory for the duration of the call, and redacted everywhere else.

How it works →

Replay without re-prompting

Payloads are stored byte-exact. After an outage, one bulk-replay call re-runs every failure — no expensive second trip through the LLM, and permanently rejected payloads are skipped automatically.

Full audit trail

Every intent, attempt, decision, and result is recorded and queryable — the compliance story your enterprise buyers will ask about first.

Signed webhooks

Results arrive HMAC-signed with replay protection. SDK verification is one function call, byte-compatible across Go, TypeScript, and Python.

Per-tenant rate limits

Token buckets per API key keep a runaway agent loop from taking the platform — or your budget — down with it.

New: failures hiding in 200 responses

A 200 OK is not always a success.
klanex checks what the body says.

Slack answers a failed call with HTTP 200 and "ok": false. GraphQL puts its errors inside a 200. Plenty of APIs describe the failure in a sentence. Anything that only checks status codes marks those calls as done, and your agent moves on believing the message was sent.

Known conventions, caught by rules

"ok": false, "success": false, an error status, and GraphQL errors with no data fail the execution instantly. No model call, no guesswork.

Everything else, judged by a model

A 200 body that reads like a failure gets one fast yes/no check from TypeSafe's Jev model. Only a confident "this failed" changes the outcome. If the check is unsure or unavailable, the success stands.

The evidence stays on the record

The failed execution keeps the exact response, the message names which check fired, and the llm_hint tells your agent what to fix. The check sees the response body and host, never your payload values or credentials.

New: every rejection, diagnosed

A 403 can mean a dozen things.
klanex tells your agent which one.

Status codes are a poor guide to what an agent should do next. GitHub sends its secondary rate limit as a 403. A 409 can be a duplicate or a lock that clears in a second. A 400 usually names the bad field, each API in its own format. klanex reads the response, works out the cause and the field, and acts on it.

The bad field, by name

The llm_hint points at the exact payload field, like amount or line_items[0].quantity, so the model fixes one value instead of rewriting the whole call. error.diagnosis carries the same answer for your code.

Disguised rate limits get retried

A 403 or 409 whose body says "slow down" or "try again" goes back on the retry queue with backoff instead of failing the task. Your agent does nothing, same as for a plain 429.

Your own duplicates count as done

When an attempt times out after reaching the API, the retry's "already exists" is that attempt's success. klanex marks it succeeded and says why, so the agent never sends it a third time.

Only a confident diagnosis changes what happens. An unclear response gets the plain rejection. The check uses TypeSafe's Jev model and sees the response, the host, and your payload's field names. Never the values, never your credentials.

Plugged into where
your team already works

Configure once with a single API call — credentials are sealed in the KMS vault like everything else.

Slack — approve without leaving the channel

Executions awaiting approval post to Slack with Approve / Reject buttons. Clicks are verified against Slack's request signature, the decision lands in the audit trail as via Slack by @you, and the message updates with the outcome. Optional failure alerts included.

Jira — failures become tickets, automatically

When an execution fails terminally, klanex files an issue in your project: execution ID, target, attempts, error code, and one-line replay instructions. No payload contents ever leave the vault. Nothing to triage by hand at 2am.

⇄

Everything else — signed webhooks & API

Every lifecycle event fires an HMAC-signed webhook, and the full audit trail is queryable through the OpenAPI-specified REST API — wire up Teams, PagerDuty, or your own dashboard.

Model Context Protocol

Adopt klanex from inside
your agent — over MCP

klanex is a Model Context Protocol server. Point any MCP-capable client — Claude Code, Claude Desktop, Cursor, VS Code, Windsurf — at the hosted endpoint and your agent gets the whole reliability engine as native tools. No SDK, no glue code.

The full reliability loop

Six tools cover submit, poll, list, replay, usage, and stored credentials — every guarantee from the REST API, callable by the model itself.

Errors the agent fixes itself

A schema-rejected payload returns as a tool error carrying the llm_hint, so the model self-corrects without ever leaving the conversation.

Zero-install or npm

Use the hosted Streamable HTTP endpoint directly, or npx -y klanex-mcp for stdio-only clients. Free sandbox keys to try it end to end.

Drop it into the agent framework
you already use

Official adapters turn any API call into a native tool for the Vercel AI SDK, the OpenAI Agents SDK, and LangGraph. The model's tool input becomes the payload. klanex owns the retries, credentials, and approvals, and hands back text the model can act on.

import { generateText, stepCountIs } from "ai";
import { Klanex } from "klanex-sdk";
import { klanexTool } from "klanex-sdk/ai";

const klanex = new Klanex({ apiKey: process.env.KLANEX_API_KEY });

const refund = klanexTool(klanex, {
  description: "Refund a Stripe charge",
  inputSchema: z.object({ charge: z.string(), amount: z.number().int() }),
  target: { url: "https://api.stripe.com/v1/refunds", connectionId: "con_…" },
});

await generateText({ model, tools: { refund }, stopWhen: stepCountIs(5), prompt });
import { Agent, run } from "@openai/agents";
import { klanexTool } from "klanex-sdk/openai-agents";

const refund = klanexTool(klanex, {
  name: "create_refund",
  description: "Refund a Stripe charge",
  parameters: z.object({ charge: z.string(), amount: z.number().int() }),
  target: { url: "https://api.stripe.com/v1/refunds", connectionId: "con_…" },
  requiresApproval: true, // a human approves in Slack first
});

await run(new Agent({ name: "Support", tools: [refund] }), prompt);

// Python: from klanex.adapters import openai_agents_tool
from langchain.agents import create_agent
from klanex import AsyncKlanex
from klanex.adapters import langchain_tool

refund = langchain_tool(
    AsyncKlanex(api_key=KLANEX_API_KEY),
    name="create_refund",
    description="Refund a Stripe charge",
    target={"url": "https://api.stripe.com/v1/refunds", "connection_id": "con_…"},
    payload_schema=refund_schema,  # the tool's own parameters
)

agent = create_agent(model, tools=[refund])
await agent.ainvoke({"messages": [("user", prompt)]})
$ curl -s https://api.klanexai.com/v1/executions \
    -H "X-API-Key: klx_…" \
    -d '{
      "target": { "url": "https://api.stripe.com/v1/refunds",
                  "headers": { "Authorization": "Bearer …" } },
      "payload": { "charge_id": "ch_9f2k", "amount": 500 },
      "payload_schema": { "type": "object", "required": ["charge_id", "amount"] },
      "idempotency_key": "refund-ch_9f2k"
    }'

{"execution_id": "exe_4d0c85a6", "status": "QUEUED"}

Rejections come back as fixes

When an API rejects the call, the tool returns the llm_hint, naming the bad field when klanex can tell. The model corrects one value and calls again, inside the same run.

Exactly once per tool call

The framework's tool call ID becomes the idempotency key. A resumed graph or a retried step can call the tool again, and the refund still happens once.

Slow APIs never cause duplicates

If the API is still retrying or a human has not approved yet, the model is told the action is in progress and not to call again. It never guesses and resubmits.

Pay for reliability,
not for tokens

Flat monthly plans with generous execution volume — no inference markup, no seat licenses. You pay per execution; retries and schema-rejected payloads are always free.

Free

$0/mo

For trying klanex.

  • 1,000 executions / month
  • 7-day execution history
  • Up to 5 retries per call
  • 2 requests / second
Start free

Starter

$19/mo

Your first agent in production.

  • 20,000 executions / month
  • Execution replay
  • 30-day history
  • Up to 10 retries · 10 req/s
Start

Enterprise

Volume execution pricing, SSO & custom SLA, custom data retention, and priority support — for teams putting agents near money.

Contact sales

Stop babysitting your agents' API calls.

Start free — 1,000 executions a month, the full reliability engine, no credit card to explore.