LLM API
The LLM API lets your code call a model directly with an Endue API key, in the same request format as the OpenAI and Anthropic APIs, and pays for each call from your credits.
When to use it
Section titled “When to use it”When you want a model rather than an agent: a completion inside your own app, code that already talks to the OpenAI or Anthropic API, or Claude Code. You change the base URL and the key; the rest of your code stays as it is.
If you want an agent to do the work — with its prompt, memory, tools, and connectors — call it through the Agent API instead.
Make your first call
Section titled “Make your first call”-
Issue an account-wide key in Settings → Account → API Keys. Copy it when it is shown; it is shown once. A key issued from an agent’s API section is limited to that agent and cannot call models directly.
-
Keep the key out of your code. Put it in an environment variable, for example
ENDUE_API_KEY. -
Point your SDK at Endue with this base URL, and call a model from the model list:
https://platform.endue.ai/api/v1/llm
import osfrom openai import OpenAI
client = OpenAI(base_url="https://platform.endue.ai/api/v1/llm", api_key=os.environ["ENDUE_API_KEY"])res = client.chat.completions.create( model="openai/gpt-5.4", messages=[{"role": "user", "content": "Hello"}],)print(res.choices[0].message.content)import OpenAI from 'openai';
const client = new OpenAI({ baseURL: 'https://platform.endue.ai/api/v1/llm', apiKey: process.env.ENDUE_API_KEY });const res = await client.chat.completions.create({ model: 'openai/gpt-5.4', messages: [{ role: 'user', content: 'Hello' }],});console.log(res.choices[0].message.content);curl https://platform.endue.ai/api/v1/llm/chat/completions \ -H "Authorization: Bearer $ENDUE_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"openai/gpt-5.4","messages":[{"role":"user","content":"Hello"}]}'import osimport anthropic
client = anthropic.Anthropic(base_url="https://platform.endue.ai/api/v1/llm", api_key=os.environ["ENDUE_API_KEY"])msg = client.messages.create( model="claude-sonnet-4-5", max_tokens=1024, messages=[{"role": "user", "content": "Hello"}],)print(msg.content[0].text)export ANTHROPIC_BASE_URL=https://platform.endue.ai/api/v1/llmexport ANTHROPIC_API_KEY=$ENDUE_API_KEYclaudeSettings → Account → LLM API shows the same base URL and examples, ready to copy.
Models and prices
Section titled “Models and prices”GET https://platform.endue.ai/api/v1/llm/models lists every model you can call, with its context length and its price. No key is needed to read it.
- Model ids look like
openai/gpt-5.4orx-ai/grok-4.6— provider, then model. - Prices are in credits per token, as decimal strings:
promptfor input,completionfor output, andinput_cache_readwhere the model discounts cached input. One credit is one US dollar. - Some models charge more for very long prompts. Those prices are listed under
overrides, with the prompt length they start at.
The same models are browsable on the models page.
Streaming
Section titled “Streaming”Set "stream": true. You receive the answer as server-sent events in the OpenAI chunk format, ending with data: [DONE]. The last chunk before [DONE] carries usage, including what the call cost.
How a call is paid for
Section titled “How a call is paid for”-
Before the model runs, Endue sets aside enough of your credits to cover the most the call could cost: your prompt plus
max_tokensof output. If your balance is lower than that, the model is not called and you get a402. -
When the call finishes, you are charged what it actually cost, and the rest of what was set aside is returned to your balance.
-
The response tells you the charge.
usage.costis the credits taken for this call, and theX-Endue-Request-Idheader identifies the call if you need to ask about it.
If you leave out max_tokens and your balance cannot cover the model’s full output length, the call goes ahead with max_tokens lowered to what your balance covers — as long as that is at least 1,024 tokens. Below that, you get a 402.
LLM API calls are paid from purchased and promotional credits only. Your plan’s allowance is for agents and is not used here. Calls show up on the usage dashboard under the Sources tab, as LLM API.
Anthropic format and Claude Code
Section titled “Anthropic format and Claude Code”The same base URL also accepts the Anthropic Messages format. The Anthropic SDK and Claude Code add /v1/messages to the base URL themselves, so the setup is the two lines in the tabs above.
- Send the key as
x-api-keyor asAuthorization: Bearer. - Claude model names work as you would write them for Anthropic —
claude-sonnet-4-5, or with a date such asclaude-sonnet-4-5-20250929— as long as that model is on the model list. Any id from the model list works too, including non-Claude models. - Streaming uses Anthropic’s event format, and errors use Anthropic’s error shape.
/v1/messages/count_tokensreturns an estimate of the input tokens. It is not billed, and it is an approximation rather than an exact tokenizer count.
Errors
Section titled “Errors”Errors use the shape of the format you called: OpenAI’s {"error": {…}} or Anthropic’s {"type": "error", "error": {…}}.
| Status | What happened | What to do |
|---|---|---|
400 | The request is malformed, n is more than 1, or the model rejected the input | Fix the request; the message says what is wrong |
401 | The key is missing, revoked, or wrong | Check the key and how you send it |
402 | Your credit balance cannot cover the call (insufficient_credits, or billing_error in Anthropic format) | Buy credits, or lower max_tokens |
403 | The key is limited to one agent | Issue an account-wide key |
404 | The model is not on the model list | Pick an id from /models |
429 | Too many requests | Wait for the number of seconds in Retry-After |
502 | The model provider failed | Retry |
503 | Billing or the model list is briefly unavailable. The model was not called and nothing was charged | Retry shortly |
Limits
Section titled “Limits”- Only the models on the model list can be called.
- Two formats are supported: OpenAI Chat Completions and Anthropic Messages. There are no embeddings, image, audio, or Responses endpoints.
nmust be 1.- Fields outside the standard request are ignored — for example a fallback list of other models, provider routing, or plugins. In the Anthropic format, server tools such as web search are refused, and
top_kis ignored. - The plan allowance is not used; calls need purchased or promotional credits.
- A key limited to one agent cannot call models.
- Up to 100 calls per minute, counted per key and per account. A request body can be up to 8 MB.
count_tokensis an estimate.