Skip to content

LLM API

The LLM API lets your code call a model directly with an Endue API key, in the same request format as the OpenAI and Anthropic APIs, and pays for each call from your credits.

When you want a model rather than an agent: a completion inside your own app, code that already talks to the OpenAI or Anthropic API, or Claude Code. You change the base URL and the key; the rest of your code stays as it is.

If you want an agent to do the work — with its prompt, memory, tools, and connectors — call it through the Agent API instead.

  1. Issue an account-wide key in Settings → Account → API Keys. Copy it when it is shown; it is shown once. A key issued from an agent’s API section is limited to that agent and cannot call models directly.

  2. Keep the key out of your code. Put it in an environment variable, for example ENDUE_API_KEY.

  3. Point your SDK at Endue with this base URL, and call a model from the model list:

    https://platform.endue.ai/api/v1/llm
import os
from openai import OpenAI
client = OpenAI(base_url="https://platform.endue.ai/api/v1/llm", api_key=os.environ["ENDUE_API_KEY"])
res = client.chat.completions.create(
model="openai/gpt-5.4",
messages=[{"role": "user", "content": "Hello"}],
)
print(res.choices[0].message.content)

Settings → Account → LLM API shows the same base URL and examples, ready to copy.

GET https://platform.endue.ai/api/v1/llm/models lists every model you can call, with its context length and its price. No key is needed to read it.

  • Model ids look like openai/gpt-5.4 or x-ai/grok-4.6 — provider, then model.
  • Prices are in credits per token, as decimal strings: prompt for input, completion for output, and input_cache_read where the model discounts cached input. One credit is one US dollar.
  • Some models charge more for very long prompts. Those prices are listed under overrides, with the prompt length they start at.

The same models are browsable on the models page.

Set "stream": true. You receive the answer as server-sent events in the OpenAI chunk format, ending with data: [DONE]. The last chunk before [DONE] carries usage, including what the call cost.

  1. Before the model runs, Endue sets aside enough of your credits to cover the most the call could cost: your prompt plus max_tokens of output. If your balance is lower than that, the model is not called and you get a 402.

  2. When the call finishes, you are charged what it actually cost, and the rest of what was set aside is returned to your balance.

  3. The response tells you the charge. usage.cost is the credits taken for this call, and the X-Endue-Request-Id header identifies the call if you need to ask about it.

If you leave out max_tokens and your balance cannot cover the model’s full output length, the call goes ahead with max_tokens lowered to what your balance covers — as long as that is at least 1,024 tokens. Below that, you get a 402.

LLM API calls are paid from purchased and promotional credits only. Your plan’s allowance is for agents and is not used here. Calls show up on the usage dashboard under the Sources tab, as LLM API.

The same base URL also accepts the Anthropic Messages format. The Anthropic SDK and Claude Code add /v1/messages to the base URL themselves, so the setup is the two lines in the tabs above.

  • Send the key as x-api-key or as Authorization: Bearer.
  • Claude model names work as you would write them for Anthropic — claude-sonnet-4-5, or with a date such as claude-sonnet-4-5-20250929 — as long as that model is on the model list. Any id from the model list works too, including non-Claude models.
  • Streaming uses Anthropic’s event format, and errors use Anthropic’s error shape.
  • /v1/messages/count_tokens returns an estimate of the input tokens. It is not billed, and it is an approximation rather than an exact tokenizer count.

Errors use the shape of the format you called: OpenAI’s {"error": {…}} or Anthropic’s {"type": "error", "error": {…}}.

StatusWhat happenedWhat to do
400The request is malformed, n is more than 1, or the model rejected the inputFix the request; the message says what is wrong
401The key is missing, revoked, or wrongCheck the key and how you send it
402Your credit balance cannot cover the call (insufficient_credits, or billing_error in Anthropic format)Buy credits, or lower max_tokens
403The key is limited to one agentIssue an account-wide key
404The model is not on the model listPick an id from /models
429Too many requestsWait for the number of seconds in Retry-After
502The model provider failedRetry
503Billing or the model list is briefly unavailable. The model was not called and nothing was chargedRetry shortly
  • Only the models on the model list can be called.
  • Two formats are supported: OpenAI Chat Completions and Anthropic Messages. There are no embeddings, image, audio, or Responses endpoints.
  • n must be 1.
  • Fields outside the standard request are ignored — for example a fallback list of other models, provider routing, or plugins. In the Anthropic format, server tools such as web search are refused, and top_k is ignored.
  • The plan allowance is not used; calls need purchased or promotional credits.
  • A key limited to one agent cannot call models.
  • Up to 100 calls per minute, counted per key and per account. A request body can be up to 8 MB.
  • count_tokens is an estimate.