On this page
XUSS AI API
A production OpenAI-compatible API on top of the XUSS hosting platform. Use any OpenAI SDK, editor, or chat client and point it at XUSS — no vendor lock-in, one key, pay-as-you-go from your XUSS balance.
- Base URL:
https://xuss.us/v1 - Auth:
Authorization: Bearer xsk-…(create a key in the panel) - Panel: https://xuss.us/panel/api
- Format: OpenAI Chat Completions (streaming, tools, vision, JSON mode)
Quickstart
1. Create an API key in the panel: API → API keys → Create key. Copy it — it is shown only once. Keys start with xsk-.
2. Call the API from any language — pick a tab:
curl https://xuss.us/v1/chat/completions \
-H "Authorization: Bearer xsk-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "xuss/kitsune",
"messages": [{"role": "user", "content": "Hello!"}]
}'from openai import OpenAI
client = OpenAI(base_url="https://xuss.us/v1", api_key="xsk-YOUR_KEY")
r = client.chat.completions.create(
model="xuss/kitsune",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://xuss.us/v1", apiKey: "xsk-YOUR_KEY" });
const r = await client.chat.completions.create({
model: "xuss/kitsune",
messages: [{ role: "user", content: "Hello!" }],
});
console.log(r.choices[0].message.content);package main
import (
"bytes"
"fmt"
"io"
"net/http"
)
func main() {
body := []byte(`{"model":"xuss/kitsune","messages":[{"role":"user","content":"Hello!"}]}`)
req, _ := http.NewRequest("POST", "https://xuss.us/v1/chat/completions", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer xsk-YOUR_KEY")
req.Header.Set("Content-Type", "application/json")
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
out, _ := io.ReadAll(res.Body)
fmt.Println(string(out))
}<?php
$ch = curl_init('https://xuss.us/v1/chat/completions');
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer xsk-YOUR_KEY',
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode([
'model' => 'xuss/kitsune',
'messages' => [['role' => 'user', 'content' => 'Hello!']],
]),
]);
$res = json_decode(curl_exec($ch), true);
echo $res['choices'][0]['message']['content'], PHP_EOL;That's it — your XUSS balance is charged per token, no subscription required.
Authentication
Every request needs a bearer token:
Authorization: Bearer xsk-YOUR_KEY
Content-Type: application/jsonKeys are created and revoked in Panel → API. Rules:
- A key looks like
xsk-followed by a long random string. Only the first characters are stored for display; the full value is shown once at creation. - Up to 10 active keys per account.
- Each key can carry an optional expiry date and an optional spend limit (USD). Once a key expires or reaches its limit, requests with it fail with 401 (
key_expired) or 402 (key_spend_limit_reached) until you create a new key or raise the limit. - Creating a key in the panel is protected by Cloudflare Turnstile (a captcha), so keys cannot be minted by scripts.
- Revoking a key takes effect immediately.
- Keys inherit your account's access to the AI API. If the API is not enabled for your account, requests return 403.
Never expose a key in client-side code or public repos. If a key leaks, revoke it in the panel and create a new one.
Endpoints
| Method | Path | Description |
|---|---|---|
| GET | /v1/models | List the models available to your key, with pricing and capabilities |
| POST | /v1/chat/completions | Create a chat completion (streaming or not) |
| GET | /v1/tools | List the server-side tools the API can run for you |
| GET | /v1/skills | List the reference skills the assistant can load |
All accept Authorization: Bearer xsk-….
Models
GET /v1/models returns the models exposed on the API, including context window, vision support, and per-token pricing:
{
"object": "list",
"data": [
{
"id": "xuss/kitsune",
"object": "model",
"owned_by": "xuss",
"context_window": 1048576,
"vision": true,
"pricing": {
"prompt": 0.00000029,
"completion": 0.00000129,
"input_cache_read": 0.0000000099
}
}
]
}Pricing is expressed in USD per token. Multiply by 1,000,000 for the per-million rate shown in the panel.
Model routing. Behind a single model id, XUSS may route to different upstream capacity to keep the service fast and available. This is invisible to you: themodelyou requested is always what is reported back, and the model knows its own name. Prices shown in/v1/modelsare the prices you are billed.
Chat Completions
POST /v1/chat/completions
Request body
| Field | Type | Notes |
|---|---|---|
model | string | required — e.g. xuss/kitsune |
messages | array | required — see Messages |
stream | boolean | true = Server-Sent Events |
temperature | number | 0–2, default 0.2 |
max_tokens | integer | caps output length |
max_completion_tokens | integer | alias of max_tokens |
tools | array | function/tool definitions — see Tools |
tool_choice | string/object | passed through to the model |
response_format | object | {"type":"json_object"} for JSON mode |
top_p, stop, seed, presence_penalty, frequency_penalty, logit_bias, n, user, logprobs, top_logprobs | — | passed through |
messages follows the OpenAI schema (system, user, assistant, tool roles; multimodal content arrays).
Response
{
"id": "chatcmpl-3f9c1a...",
"object": "chat.completion",
"created": 1791000000,
"model": "xuss/kitsune",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "Hello! How can I help?"},
"finish_reason": "stop",
"logprobs": null
}
],
"usage": {"prompt_tokens": 12, "completion_tokens": 7, "total_tokens": 19},
"system_fingerprint": "xuss"
}finish_reason is stop, length, or tool_calls.
Messages
A standard OpenAI messages array. A minimal single-turn request:
{"model": "xuss/kitsune", "messages": [{"role": "user", "content": "Hi"}]}Multi-turn conversation — send the full history each time:
{
"model": "xuss/kitsune",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is XUSS?"},
{"role": "assistant", "content": "XUSS is a hosting platform."},
{"role": "user", "content": "Does it have an AI API?"}
]
}Streaming
Set "stream": true to receive Server-Sent Events. Each line is data: {json}; the stream ends with data: [DONE].
The chunks look like this:
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":…,"model":"xuss/kitsune","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}],"system_fingerprint":"xuss"}
data: {"id":"chatcmpl-…","choices":[{"index":0,"delta":{"content":"Silent "},"finish_reason":null}]}
data: {"id":"chatcmpl-…","choices":[{"index":0,"delta":{"content":"racks hum"},"finish_reason":null}]}
data: {"id":"chatcmpl-…","choices":[],"usage":{"prompt_tokens":14,"completion_tokens":9,"total_tokens":23}}
data: [DONE]Read the stream in any language:
curl -N https://xuss.us/v1/chat/completions \
-H "Authorization: Bearer xsk-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"xuss/kitsune","stream":true,
"messages":[{"role":"user","content":"Write a haiku about servers"}]}'stream = client.chat.completions.create(
model="xuss/kitsune",
messages=[{"role": "user", "content": "Write a haiku about servers"}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta.content if chunk.choices else None
if delta:
print(delta, end="", flush=True)const stream = await client.chat.completions.create({
model: "xuss/kitsune",
messages: [{ role: "user", content: "Write a haiku about servers" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content || "");
}package main
import (
"bufio"
"bytes"
"fmt"
"net/http"
"strings"
)
func main() {
body := []byte(`{"model":"xuss/kitsune","stream":true,"messages":[{"role":"user","content":"Write a haiku about servers"}]}`)
req, _ := http.NewRequest("POST", "https://xuss.us/v1/chat/completions", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer xsk-YOUR_KEY")
req.Header.Set("Content-Type", "application/json")
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
sc := bufio.NewScanner(res.Body)
for sc.Scan() {
line := sc.Text()
if !strings.HasPrefix(line, "data: ") || strings.HasSuffix(line, "[DONE]") {
continue
}
fmt.Println(line[6:]) // parse JSON, read choices[0].delta.content
}
}<?php
$ch = curl_init('https://xuss.us/v1/chat/completions');
curl_setopt_array($ch, [
CURLOPT_HTTPHEADER => [
'Authorization: Bearer xsk-YOUR_KEY',
'Content-Type: application/json',
],
CURLOPT_POSTFIELDS => json_encode([
'model' => 'xuss/kitsune',
'stream' => true,
'messages' => [['role' => 'user', 'content' => 'Write a haiku about servers']],
]),
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) {
foreach (explode("\n", $chunk) as $line) {
if (str_starts_with($line, 'data: ') && !str_contains($line, '[DONE]')) {
$j = json_decode(substr($line, 6), true);
echo $j['choices'][0]['delta']['content'] ?? '';
}
}
return strlen($chunk);
},
]);
curl_exec($ch);When the model supports reasoning, streamed chunks may additionally carry areasoning_contentfield indelta. It is ignored by standard SDKs and never counted as output you must handle.
System prompts and identity
Your system messages are honoured. The model also knows its own display name and, when asked which model it is, answers with that name only — it never reveals any upstream provider or route.
{
"model": "xuss/kitsune",
"messages": [
{"role": "system", "content": "You are a terse DevOps assistant. Answer in bullet points."},
{"role": "user", "content": "How do I restart a systemd service?"}
]
}Vision (images)
Pass images as OpenAI image_url content parts (data URLs or public http(s) URLs). Only models with "vision": true in /v1/models accept images; sending an image to a non-vision model returns a text note and the model answers on the text only.
r = client.chat.completions.create(
model="xuss/kitsune",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}},
],
}],
)Limits: up to 8 image parts, each up to ~6 MB; images do not count toward the text budget.
Function calling
Pass OpenAI-style tools. When the model decides to call a function, finish_reason is tool_calls and the message contains tool_calls.
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
r = client.chat.completions.create(
model="xuss/kitsune",
messages=[{"role": "user", "content": "Weather in Tashkent?"}],
tools=tools,
)
call = r.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)Then send the result back as a tool message:
messages = [
{"role": "user", "content": "Weather in Tashkent?"},
r.choices[0].message,
{"role": "tool", "tool_call_id": call.id, "content": '{"temp_c": 24, "sky": "clear"}'},
]
final = client.chat.completions.create(model="xuss/kitsune", messages=messages)Server tools (web + skills)
Opt in per request. Add "xuss_tools": true (all of them) or an array of the ones you want. XUSS then runs those tools for you on the server — the model calls them, the server executes them and continues, and you get the final answer in the same response. No client-side loop needed.
r = client.chat.completions.create(
model="xuss/kitsune",
xuss_tools=True, # or ["web_search", "read_skill"]
messages=[{"role": "user",
"content": "Search the web for the latest Python release and summarise it."}],
)
print(r.choices[0].message.content){
"model": "xuss/kitsune",
"xuss_tools": true,
"messages": [{"role": "user", "content": "Open https://example.com and give me the page title."}]
}| Tool | What it does |
|---|---|
web_search | Web search (SearXNG metasearch): titles, URLs and snippets |
fetch_url | Fetch an HTTP(S) URL server-side and return its text (HTML is converted to text) |
browser_render | Open a page in a real headless browser: HTTP status, title, visible text, console errors, failed requests, layout metrics, optional JS eval |
read_skill | Load one of the reference skills below (full how-to doc) into the model's context |
These are read-only: no access to files, databases, hostings or your account. Tools that change things (file management, DNS, databases, and so on) stay in your own hands and are intentionally not exposed.
- Non-streaming: the server runs the tool loop (up to 6 rounds) and returns the final answer;
usage(and billing) covers every round. - Streaming (
"stream": true): the tools run server-side, then the final answer is streamed as SSE as usual. - If you also send your own
tools, the model can call both — XUSS executes its own and hands any your tool call back throughtool_callsfor you to run (standard behaviour). - Discover them at runtime:
GET /v1/tools(definitions) andGET /v1/skills(the list below).
Available skills
read_skill loads the full reference document for any of these into context — just tell the model which one (e.g. "use the product-design skill").
| Skill id | Covers |
|---|---|
telegram-bots | complete Telegram Bot API 10.3: every method/type, aiogram 3 examples, payments/Stars, webhooks, premium emoji/gifts, rich messages & streaming drafts |
product-design | websites/pages/galleries/UI that look intentionally designed: tokens, anti-slop rules, layouts, photo-wallpaper galleries |
cybersecurity | write & audit secure code: injection/XSS/CSRF/IDOR, secrets management, bot security, incident response |
telegram-miniapps | Telegram WebApps: initData auth, theme vars, MainButton/BackButton, Stars, deployment |
shop-bot | complete Telegram store bot: catalog, cart, orders, admin panel, delivery, payment flow |
payments | Click/Payme/Paylov/Uzum + Crypto Pay + Telegram Stars: invoices, webhook verification, idempotency |
php-web | PHP sites & WordPress: structure, PDO, auth/CSRF, templates, security, deployment |
ai-integration | LLMs inside user apps: chat/streaming, RAG, prompting, cost limits, API key security |
python-backend | production Python: bots, FastAPI/Flask, asyncio discipline, DB access, run services, error handling |
node-backend | Node/TypeScript backends & bots: Express/Fastify/Telegraf, env, process management, errors |
databases | MySQL/PostgreSQL/SQLite: schema design, indexes, migrations, transactions, backups |
rest-api | API design: auth, validation, pagination, one error shape, rate limits, webhook signing |
deployment-ops | deploy projects on this hosting: domains/DNS/SSL, reverse proxy, ports, cron, backups, logs |
git-github | git workflows, deploy keys, auto-deploy webhooks, secret hygiene, rollback |
scraping-automation | ethical scrapers & watchers: structured sources, backoff, dedupe, scheduling, alerts |
seo | technical SEO: titles, structured data, sitemaps, hreflang, indexing (Google & Yandex) |
media-pipeline | images/video pipelines: resizing, WebP, thumbnails, compression, ffmpeg previews |
i18n-localization | uz/ru/en + RTL: string dicts, number/date/plural formats, bot language, hreflang |
testing-quality | run/verify/debug discipline, unit tests, linting, review habits before saying done |
analytics-monitoring | uptime checks, error tracking, daily stats, privacy-safe analytics and alerts |
legal-templates | privacy/terms/refund pages + consent + data-deletion flows (practical baseline) |
react-best-practices | React/Next.js performance: waterfalls, bundles, rendering, hydration, rerenders |
mobile-design | native mobile UX: platform patterns, touch psychology, mobile performance |
senior-frontend | senior frontend engineering: architecture, components, performance reviews |
senior-backend | senior backend engineering: API/DB design, scaling, code reviews |
senior-security | security architecture, threat modeling, crypto implementation, audits |
ui-design-system | design tokens, components, handoff; generate a token system from a brand color |
tgbot-clone | clone a Telegram bot's features safely: probe, map features, implement and test |
product-layers | layered product/UX method: needs → strategy → conceptual model → surface |
find-skills | find & install more reusable skills for agent projects |
Structured output (JSON mode)
Set response_format to {"type":"json_object"} and instruct the model to emit JSON. The reply's content is a JSON string.
r = client.chat.completions.create(
model="xuss/kitsune",
response_format={"type": "json_object"},
messages=[{"role": "user",
"content": "Return JSON with keys a and b, a=1, b=2"}],
)
import json
print(json.loads(r.choices[0].message.content)) # {'a': 1, 'b': 2}Always parse defensively and describe the exact shape you want in the prompt.
Billing and token accounting
- Charges are deducted from your XUSS balance (top up in the panel).
- Billed per token at the model's API price: input (cache miss), cache hit (cheaper when supported), and output.
- Every request's
usagereportsprompt_tokens,completion_tokens, andtotal_tokens. - The panel API page shows spending, requests, tokens, per-model and per-key breakdowns, and a recent-calls log, with selectable ranges (today, 7/30/90 days, this/last month, all time).
- Requests are rejected with 402 when your balance is insufficient.
Cost formula (per request):
cost = cache_miss_tokens × price_in
+ cache_hit_tokens × price_cache
+ completion_tokens × price_out(Prices are per token; divide by 1,000,000 for per-million.)
Rate limits and quotas
| Scope | Default |
|---|---|
| Per user | 300 requests / minute |
| Per API key | 600 requests / minute |
Exceeding a limit returns 429 with Retry-After semantics (back off and retry). Request body is capped at 25 MB.
Errors
Errors follow the OpenAI error format:
{
"error": {
"message": "Model 'xuss/foo' not found",
"type": "invalid_request_error",
"param": "model",
"code": "model_not_found"
}
}| Status | type / code | Meaning |
|---|---|---|
| 400 | invalid_request_error | Bad request (missing/invalid fields, unparseable body) |
| 401 | authentication_error · invalid_api_key, key_expired | Missing, invalid, or expired API key |
| 402 | insufficient_quota · key_spend_limit_reached | Balance empty or the key hit its spend limit |
| 403 | permission_error | API not enabled for your account, or account disabled |
| 404 | invalid_request_error · model_not_found | Unknown model |
| 413 | invalid_request_error | Request body too large (> 25 MB) |
| 422 | invalid_request_error | Invalid request payload (schema) |
| 429 | rate_limit_error · rate_limit_exceeded | Rate limit exceeded — slow down |
| 500 / 502 | api_error · upstream_error | Temporary upstream problem — retry with backoff |
| 503 | api_error · service_unavailable | The AI API is currently disabled |
param is set when the error concerns a specific request field. Standard SDKs read error.message (and error.code) directly.
Temporary failures are also delivered as a normal assistant message beginning "The model is temporarily unavailable…" when a stream has already started.
Using with OpenAI-compatible tools
Any client that supports a custom OpenAI base URL works. Set:
- Base URL:
https://xuss.us/v1 - API key: your
xsk-…key - Model: e.g.
xuss/kitsune
Examples:
# open-webui / LibreChat / Cursor / Cline / Continue / LangChain:
# set the OpenAI base URL to https://xuss.us/v1 and paste your keyEnvironment-variable style (many tools honour these):
export OPENAI_BASE_URL="https://xuss.us/v1"
export OPENAI_API_KEY="xsk-YOUR_KEY"LangChain (Python):
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="xuss/kitsune", base_url="https://xuss.us/v1",
api_key="xsk-YOUR_KEY")
print(llm.invoke("Hello").content)Advanced parameters
The following are accepted and passed through to the model when supported: top_p, stop, seed, presence_penalty, frequency_penalty, logit_bias, user, n, parallel_tool_calls, tool_choice.
FAQ
Do I need a separate subscription? No. You pay per token from your XUSS balance.
Which model ids do I use? Exactly the ids from GET /v1/models (e.g. xuss/kitsune). Display names like "Kitsune" are shown in the panel.
Can I send images? Yes, to models marked "vision": true.
Is my data used for training? No — requests are proxied to the model to produce your answer and are not used for training.
What happens if a provider is slow or down? The request is transparently retried on alternate capacity; you keep the same model id and pricing.
How do I rotate a key? Create a new key, switch your apps over, then revoke the old one in the panel.
Support
- Panel: https://xuss.us/panel/api
- Tickets & Telegram support from your panel dashboard
- Email: support@xuss.us