On this page
XUSS AI API
A production OpenAI-compatible API on top of the XUSS hosting platform. Use any OpenAI SDK, editor, or chat client and point it at XUSS — no vendor lock-in, one key, pay-as-you-go from your XUSS balance.
- Base URL:
https://xuss.us/v1 - Auth:
Authorization: Bearer xsc-…(create a key in the panel) - Panel: https://xuss.us/panel/api
- Format: OpenAI Chat Completions (streaming, tools, vision, JSON mode)
Quickstart
1. Create an API key in the panel: API → API keys → Create key. Copy it — it is shown only once. Keys start with xsc-.
2. Call the API from any language — pick a tab:
curl https://xuss.us/v1/chat/completions \
-H "Authorization: Bearer xsc-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "xuss/kitsune",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(base_url="https://xuss.us/v1", api_key="xsc-YOUR_KEY")
r = client.chat.completions.create(
model="xuss/kitsune",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://xuss.us/v1", apiKey: "xsc-YOUR_KEY" });
const r = await client.chat.completions.create({
model: "xuss/kitsune",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(r.choices[0].message.content);package main
import (
"bytes"
"fmt"
"io"
"net/http"
)
func main() {
body := []byte(`{"model":"xuss/kitsune","messages":[{"role":"user","content":"Hello!"}]}`)
req, _ := http.NewRequest("POST", "https://xuss.us/v1/chat/completions", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer xsc-YOUR_KEY")
req.Header.Set("Content-Type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
out, _ := io.ReadAll(res.Body)
fmt.Println(string(out))
}<?php
$ch = curl_init('https://xuss.us/v1/chat/completions');
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer xsc-YOUR_KEY',
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode([
'model' => 'xuss/kitsune',
'messages' => [['role' => 'user', 'content' => 'Hello!']],
]),
]);
$res = json_decode(curl_exec($ch), true);
echo $res['choices'][0]['message']['content'], PHP_EOL;That's it — your XUSS balance is charged per token, no subscription required.
Authentication
Every request needs a bearer token:
Authorization: Bearer xsc-YOUR_KEY
Content-Type: application/jsonKeys are created and revoked in Panel → API. Rules:
- A key looks like
xsc-followed by a long random string. Only the first characters are stored for display; the full value is shown once at creation. - Up to 10 active keys per account.
- Each key can carry an optional expiry date and an optional spend limit (USD). Once a key expires or reaches its limit, requests with it fail with 401 (
key_expired) or 402 (key_spend_limit_reached) until you create a new key or raise the limit. - Creating a key in the panel is protected by Cloudflare Turnstile (a captcha), so keys cannot be minted by scripts.
- Revoking a key takes effect immediately.
- Keys inherit your account's access to the AI API. If the API is not enabled for your account, requests return 403.
Never expose a key in client-side code or public repos. If a key leaks, revoke it in the panel and create a new one.
Endpoints
| Method | Path | Description |
|---|---|---|
| GET | /v1/models | List the models available to your key, with pricing and capabilities |
| POST | /v1/chat/completions | Create a chat completion (streaming or not) |
Both accept Authorization: Bearer xsc-….
Models
GET /v1/models returns the models exposed on the API, including context window, vision support, and per-token pricing:
{
"object": "list",
"data": [
{
"id": "xuss/kitsune",
"object": "model",
"owned_by": "xuss",
"context_window": 1048576,
"vision": true,
"pricing": {
"prompt": 0.00000029,
"completion": 0.00000129,
"input_cache_read": 0.0000000099
}
}
]
}Pricing is expressed in USD per token. Multiply by 1,000,000 for the per-million rate shown in the panel.
Model routing. Behind a single model id, XUSS may route to different upstream capacity to keep the service fast and available. This is invisible to you: themodelyou requested is always what is reported back, and the model knows its own name. Prices shown in/v1/modelsare the prices you are billed.
Chat Completions
POST /v1/chat/completions
Request body
| Field | Type | Notes |
|---|---|---|
model | string | required — e.g. xuss/kitsune |
messages | array | required — see Messages |
stream | boolean | true = Server-Sent Events |
temperature | number | 0–2, default 0.2 |
max_tokens | integer | caps output length |
max_completion_tokens | integer | alias of max_tokens |
tools | array | function/tool definitions — see Tools |
tool_choice | string/object | passed through to the model |
response_format | object | {"type":"json_object"} for JSON mode |
top_p, stop, seed, presence_penalty, frequency_penalty, logit_bias, n, user, logprobs, top_logprobs | — | passed through |
messages follows the OpenAI schema (system, user, assistant, tool roles; multimodal content arrays).
Response
{
"id": "chatcmpl-3f9c1a...",
"object": "chat.completion",
"created": 1791000000,
"model": "xuss/kitsune",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Hello! How can I help?"},
"finish_reason": "stop",
"logprobs": null
}
],
"usage": {"prompt_tokens": 12, "completion_tokens": 7, "total_tokens": 19},
"system_fingerprint": "xuss"
}finish_reason is stop, length, or tool_calls.
Messages
A standard OpenAI messages array. A minimal single-turn request:
{"model": "xuss/kitsune", "messages": [{"role": "user", "content": "Hi"}]}Multi-turn conversation — send the full history each time:
{
"model": "xuss/kitsune",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is XUSS?"},
{"role": "assistant", "content": "XUSS is a hosting platform."},
{"role": "user", "content": "Does it have an AI API?"}
]
}Streaming
Set "stream": true to receive Server-Sent Events. Each line is data: {json}; the stream ends with data: [DONE].
The chunks look like this:
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":…,"model":"xuss/kitsune","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}],"system_fingerprint":"xuss"}
data: {"id":"chatcmpl-…","choices":[{"index":0,"delta":{"content":"Silent "},"finish_reason":null}]}
data: {"id":"chatcmpl-…","choices":[{"index":0,"delta":{"content":"racks hum"},"finish_reason":null}]}
data: {"id":"chatcmpl-…","choices":[],"usage":{"prompt_tokens":14,"completion_tokens":9,"total_tokens":23}}
data: [DONE]Read the stream in any language:
curl -N https://xuss.us/v1/chat/completions \
-H "Authorization: Bearer xsc-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"xuss/kitsune","stream":true,
"messages":[{"role":"user","content":"Write a haiku about servers"}]}'stream = client.chat.completions.create(
model="xuss/kitsune",
messages=[{"role": "user", "content": "Write a haiku about servers"}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content if chunk.choices else None
if delta:
print(delta, end="", flush=True)const stream = await client.chat.completions.create({
model: "xuss/kitsune",
messages: [{ role: "user", content: "Write a haiku about servers" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}package main
import (
"bufio"
"bytes"
"fmt"
"net/http"
"strings"
)
func main() {
body := []byte(`{"model":"xuss/kitsune","stream":true,"messages":[{"role":"user","content":"Write a haiku about servers"}]}`)
req, _ := http.NewRequest("POST", "https://xuss.us/v1/chat/completions", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer xsc-YOUR_KEY")
req.Header.Set("Content-Type", "application/json")
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
sc := bufio.NewScanner(res.Body)
for sc.Scan() {
line := sc.Text()
if !strings.HasPrefix(line, "data: ") || strings.HasSuffix(line, "[DONE]") {
continue
}
fmt.Println(line[6:]) // parse JSON, read choices[0].delta.content
}
}<?php
$ch = curl_init('https://xuss.us/v1/chat/completions');
curl_setopt_array($ch, [
CURLOPT_HTTPHEADER => [
'Authorization: Bearer xsc-YOUR_KEY',
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode([
'model' => 'xuss/kitsune',
'stream' => true,
'messages' => [['role' => 'user', 'content' => 'Write a haiku about servers']],
]),
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) {
foreach (explode("\n", $chunk) as $line) {
if (str_starts_with($line, 'data: ') && !str_contains($line, '[DONE]')) {
$j = json_decode(substr($line, 6), true);
echo $j['choices'][0]['delta']['content'] ?? '';
}
}
return strlen($chunk);
},
]);
curl_exec($ch);When the model supports reasoning, streamed chunks may additionally carry areasoning_contentfield indelta. It is ignored by standard SDKs and never counted as output you must handle.
System prompts and identity
Your system messages are honoured. The model also knows its own display name and, when asked which model it is, answers with that name only — it never reveals any upstream provider or route.
{
"model": "xuss/kitsune",
"messages": [
{"role": "system", "content": "You are a terse DevOps assistant. Answer in bullet points."},
{"role": "user", "content": "How do I restart a systemd service?"}
]
}Vision (images)
Pass images as OpenAI image_url content parts (data URLs or public http(s) URLs). Only models with "vision": true in /v1/models accept images; sending an image to a non-vision model returns a text note and the model answers on the text only.
r = client.chat.completions.create(
model="xuss/kitsune",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}},
],
}],
)Limits: up to 8 image parts, each up to ~6 MB; images do not count toward the text budget.
Function calling
Pass OpenAI-style tools. When the model decides to call a function, finish_reason is tool_calls and the message contains tool_calls.
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
r = client.chat.completions.create(
model="xuss/kitsune",
messages=[{"role": "user", "content": "Weather in Tashkent?"}],
tools=tools,
)
call = r.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)Then send the result back as a tool message:
messages = [
{"role": "user", "content": "Weather in Tashkent?"},
r.choices[0].message,
{"role": "tool", "tool_call_id": call.id, "content": '{"temp_c": 24, "sky": "clear"}'},
]
final = client.chat.completions.create(model="xuss/kitsune", messages=messages)Structured output (JSON mode)
Set response_format to {"type":"json_object"} and instruct the model to emit JSON. The reply's content is a JSON string.
r = client.chat.completions.create(
model="xuss/kitsune",
response_format={"type": "json_object"},
messages=[{"role": "user",
"content": "Return JSON with keys a and b, a=1, b=2"}],
)
import json
print(json.loads(r.choices[0].message.content)) # {'a': 1, 'b': 2}Always parse defensively and describe the exact shape you want in the prompt.
Billing and token accounting
- Charges are deducted from your XUSS balance (top up in the panel).
- Billed per token at the model's API price: input (cache miss), cache hit (cheaper when supported), and output.
- Every request's
usagereportsprompt_tokens,completion_tokens, andtotal_tokens. - The panel API page shows spending, requests, tokens, per-model and per-key breakdowns, and a recent-calls log, with selectable ranges (today, 7/30/90 days, this/last month, all time).
- Requests are rejected with 402 when your balance is insufficient.
Cost formula (per request):
cost = cache_miss_tokens × price_in
+ cache_hit_tokens × price_cache
+ completion_tokens × price_out(Prices are per token; divide by 1,000,000 for per-million.)
Rate limits and quotas
| Scope | Default |
|---|---|
| Per user | 300 requests / minute |
| Per API key | 600 requests / minute |
Exceeding a limit returns 429 with Retry-After semantics (back off and retry). Request body is capped at 25 MB.
Errors
Errors follow the OpenAI error format:
{
"error": {
"message": "Model 'xuss/foo' not found",
"type": "invalid_request_error",
"param": "model",
"code": "model_not_found"
}
}| Status | type / code | Meaning |
|---|---|---|
| 400 | invalid_request_error | Bad request (missing/invalid fields, unparseable body) |
| 401 | authentication_error · invalid_api_key, key_expired | Missing, invalid, or expired API key |
| 402 | insufficient_quota · key_spend_limit_reached | Balance empty or the key hit its spend limit |
| 403 | permission_error | API not enabled for your account, or account disabled |
| 404 | invalid_request_error · model_not_found | Unknown model |
| 413 | invalid_request_error | Request body too large (> 25 MB) |
| 422 | invalid_request_error | Invalid request payload (schema) |
| 429 | rate_limit_error · rate_limit_exceeded | Rate limit exceeded — slow down |
| 500 / 502 | api_error · upstream_error | Temporary upstream problem — retry with backoff |
| 503 | api_error · service_unavailable | The AI API is currently disabled |
param is set when the error concerns a specific request field. Standard SDKs read error.message (and error.code) directly.
Temporary failures are also delivered as a normal assistant message beginning "The model is temporarily unavailable…" when a stream has already started.
Using with OpenAI-compatible tools
Any client that supports a custom OpenAI base URL works. Set:
- Base URL:
https://xuss.us/v1 - API key: your
xsc-…key - Model: e.g.
xuss/kitsune
Examples:
# open-webui / LibreChat / Cursor / Cline / Continue / LangChain:
# set the OpenAI base URL to https://xuss.us/v1 and paste your keyEnvironment-variable style (many tools honour these):
export OPENAI_BASE_URL="https://xuss.us/v1"
export OPENAI_API_KEY="xsc-YOUR_KEY"LangChain (Python):
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="xuss/kitsune", base_url="https://xuss.us/v1",
api_key="xsc-YOUR_KEY")
print(llm.invoke("Hello").content)Advanced parameters
The following are accepted and passed through to the model when supported: top_p, stop, seed, presence_penalty, frequency_penalty, logit_bias, user, n, parallel_tool_calls, tool_choice.
FAQ
Do I need a separate subscription? No. You pay per token from your XUSS balance.
Which model ids do I use? Exactly the ids from GET /v1/models (e.g. xuss/kitsune). Display names like "Kitsune" are shown in the panel.
Can I send images? Yes, to models marked "vision": true.
Is my data used for training? No — requests are proxied to the model to produce your answer and are not used for training.
What happens if a provider is slow or down? The request is transparently retried on alternate capacity; you keep the same model id and pricing.
How do I rotate a key? Create a new key, switch your apps over, then revoke the old one in the panel.
Support
- Panel: https://xuss.us/panel/api
- Tickets & Telegram support from your panel dashboard
- Email: support@xuss.us