Quickstart
llmax.ai speaks the OpenAI API. If your code already talks to OpenAI, you change the base URL and the key — nothing else. No SDK to install, no client to rewrite, and nothing to undo if you ever leave.
The only change
Two values, wherever your client keeps them:
| Base URL | https://run.llmax.ai/v1 |
|---|---|
| API key | sk-… (your llmax.ai key) |
| Model | qwen3.6 |
Keep the key out of your source: the code examples below read it from an environment variable, and the tools take the same two values in a settings screen.
A GET /v1/models always returns the authoritative list of model
IDs your key can reach — worth a look before you hard-code one.
The OpenAI API base URL
Every OpenAI client has one setting that decides where the requests actually go. Out of the box it points at OpenAI's own endpoint:
| OpenAI (default) | https://api.openai.com/v1 |
|---|---|
| llmax.ai | https://run.llmax.ai/v1 |
Both ends speak the same protocol, so swapping one for the other is a single line. The only thing that ever makes this confusing is that each SDK spells the setting differently:
Python (openai) | base_url= |
|---|---|
| JavaScript / TypeScript | baseURL: |
| Environment variable | OPENAI_BASE_URL (older clients: OPENAI_API_BASE) |
| Continue | apiBase |
| Cline, Open WebUI | Base URL |
| Xcode 26 | URL (host only, no /v1) |
Whatever the field is called, the value ends at /v1 and your client appends the rest of
the path itself — /chat/completions, /models, /embeddings.
Pasting a full endpoint URL into a field that expects only the base is the most common way to get a
404.
Coding tools and chat apps
Most people never call the API directly — their editor or chat app does it for them. Anything that accepts a custom OpenAI-compatible endpoint works, which is most of the ecosystem, and you are always filling in the same two fields. No code required.
Cline (VS Code)
| API Provider | OpenAI Compatible |
|---|---|
| Base URL | https://run.llmax.ai/v1 |
| API Key | sk-… |
| Model ID | qwen3.6 |
Continue (VS Code / JetBrains)
In config.json, under models:
{
"title": "llmax.ai",
"provider": "openai",
"model": "qwen3.6",
"apiBase": "https://run.llmax.ai/v1",
"apiKey": "sk-…"
}
Aider
export OPENAI_API_BASE="https://run.llmax.ai/v1"
export OPENAI_API_KEY="sk-…"
aider --model openai/qwen3.6
Xcode 26
Xcode → Settings → Intelligence → Add a Model Provider → Internet Hosted.
| URL | https://run.llmax.ai (no /v1) |
|---|---|
| API Key | sk-… (no "Bearer" prefix) |
Xcode is the one exception to the base URL above: it appends the version segment itself, so pasting the
full https://run.llmax.ai/v1 makes it request /v1/v1/… and every call 404s.
Enter the host alone. It also adds the Bearer prefix to the key for you. Both quirks, and
how to set up a local model instead, are covered in
a longer write-up on the blog.
Open WebUI and other chat clients
Add an OpenAI-compatible connection with the same base URL and key. This is the route for anyone who isn't writing code: the app does the talking, you only paste two values into its settings.
All of the above is squarely within normal use — agents are what the plan is built for. The one thing the terms rule out is unattended automation: scripts or schedulers looping with nobody supervising the session.
Python
Calling the API from your own code is the same story: the official openai package works
as-is, no fork and no wrapper.
# pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://run.llmax.ai/v1",
api_key=os.environ["LLMAX_API_KEY"],
)
resp = client.chat.completions.create(
model="qwen3.6",
messages=[{"role": "user", "content": "Explain actors in Swift, briefly."}],
)
print(resp.choices[0].message.content)
JavaScript and TypeScript
// npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://run.llmax.ai/v1",
apiKey: process.env.LLMAX_API_KEY,
});
const resp = await client.chat.completions.create({
model: "qwen3.6",
messages: [{ role: "user", content: "Explain coroutines in Kotlin, briefly." }],
});
console.log(resp.choices[0].message.content);
curl
curl https://run.llmax.ai/v1/chat/completions \
-H "Authorization: Bearer $LLMAX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6",
"messages": [{"role": "user", "content": "Hello"}]
}'
To list what your key can reach:
curl https://run.llmax.ai/v1/models \
-H "Authorization: Bearer $LLMAX_API_KEY"
Streaming
Set stream: true and read the chunks exactly as you would from OpenAI — the
server-sent-event format is the same, terminator included.
stream = client.chat.completions.create(
model="qwen3.6",
messages=[{"role": "user", "content": "Write a haiku about tokens."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Embeddings and reranking
The /v1/embeddings and /v1/rerank endpoints follow the same shape and the
same authentication. They are billed as add-ons rather than being part of the flat inference plan —
see the terms for how that works.
Rate limits and fair use
Tokens are unlimited; requests are not unmetered chaos. Each plan carries requests-per-minute and
concurrency limits so one account can't saturate the cluster and slow everyone else down. Exceeding
them returns 429, and the right response is the usual one: back off and retry.
Separately, each plan has a monthly fair-use volume, sized well above what full-time development with agents consumes. Crossing it doesn't cut you off or start charging per token: your requests move to a lower-priority queue for the rest of the billing period, so you see latency rise rather than errors. The figures for your plan are in your dashboard.
Common errors
401 — authentication
Either no key reached the server or it isn't a valid llmax.ai key. Check
that your client sends Authorization: Bearer sk-… and that the environment variable is
actually populated in the shell that runs your code.
404 — wrong path
Almost always the base URL. It must end in /v1, and your client appends the rest. If you
set https://run.llmax.ai without /v1, or paste the full
/v1/chat/completions into a field that expects only the base, you get a 404.
429 — rate limited
Too many requests per minute, or too many in flight at once. Retry with exponential backoff. Note this is not what fair use looks like: passing the monthly volume lowers your priority, it doesn't return errors. If you're seeing 429s, it's the per-minute or concurrency ceiling, not volume.
What doesn't change
Your prompts, your tool definitions, your streaming code, your retry logic and your error handling. That's the point of being OpenAI-compatible: the switch is reversible, so trying it costs you an afternoon at most, and leaving would cost the same.
What you gain is on the other side of the wire — EU-only processing with zero logs and a bill that doesn't move with usage. If you want to see what that's worth at your volume, the cost calculator does the arithmetic.