XUSS. / AI API Docs Panel
On this page

XUSS AI API

A production OpenAI-compatible API on top of the XUSS hosting platform. Use any OpenAI SDK, editor, or chat client and point it at XUSS — no vendor lock-in, one key, pay-as-you-go from your XUSS balance.

Quickstart

1. Create an API key in the panel: API → API keys → Create key. Copy it — it is shown only once. Keys start with xsk-.

2. Call the API from any language — pick a tab:

curl https://xuss.us/v1/chat/completions \
  -H "Authorization: Bearer xsk-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "xuss/kitsune",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'
from openai import OpenAI

client = OpenAI(base_url="https://xuss.us/v1", api_key="xsk-YOUR_KEY")

r = client.chat.completions.create(
    model="xuss/kitsune",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://xuss.us/v1", apiKey: "xsk-YOUR_KEY" });

const r = await client.chat.completions.create({
  model: "xuss/kitsune",
  messages: [{ role: "user", content: "Hello!" }],
});
console.log(r.choices[0].message.content);
package main

import (
	"bytes"
	"fmt"
	"io"
	"net/http"
)

func main() {
	body := []byte(`{"model":"xuss/kitsune","messages":[{"role":"user","content":"Hello!"}]}`)
	req, _ := http.NewRequest("POST", "https://xuss.us/v1/chat/completions", bytes.NewReader(body))
	req.Header.Set("Authorization", "Bearer xsk-YOUR_KEY")
	req.Header.Set("Content-Type", "application/json")

	res, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer res.Body.Close()
	out, _ := io.ReadAll(res.Body)
	fmt.Println(string(out))
}
<?php
$ch = curl_init('https://xuss.us/v1/chat/completions');
curl_setopt_array($ch, [
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer xsk-YOUR_KEY',
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode([
        'model' => 'xuss/kitsune',
        'messages' => [['role' => 'user', 'content' => 'Hello!']],
    ]),
]);
$res = json_decode(curl_exec($ch), true);
echo $res['choices'][0]['message']['content'], PHP_EOL;

That's it — your XUSS balance is charged per token, no subscription required.

Authentication

Every request needs a bearer token:

Authorization: Bearer xsk-YOUR_KEY
Content-Type: application/json

Keys are created and revoked in Panel → API. Rules:

Never expose a key in client-side code or public repos. If a key leaks, revoke it in the panel and create a new one.

Endpoints

MethodPathDescription
GET/v1/modelsList the models available to your key, with pricing and capabilities
POST/v1/chat/completionsCreate a chat completion (streaming or not)
GET/v1/toolsList the server-side tools the API can run for you
GET/v1/skillsList the reference skills the assistant can load

All accept Authorization: Bearer xsk-….

Models

…

GET /v1/models returns the models exposed on the API, including context window, vision support, and per-token pricing:

{
  "object": "list",
  "data": [
    {
      "id": "xuss/kitsune",
      "object": "model",
      "owned_by": "xuss",
      "context_window": 1048576,
      "vision": true,
      "pricing": {
        "prompt": 0.00000029,
        "completion": 0.00000129,
        "input_cache_read": 0.0000000099
      }
    }
  ]
}

Pricing is expressed in USD per token. Multiply by 1,000,000 for the per-million rate shown in the panel.

Model routing. Behind a single model id, XUSS may route to different upstream capacity to keep the service fast and available. This is invisible to you: the model you requested is always what is reported back, and the model knows its own name. Prices shown in /v1/models are the prices you are billed.

Chat Completions

POST /v1/chat/completions

Request body

FieldTypeNotes
modelstringrequired — e.g. xuss/kitsune
messagesarrayrequired — see Messages
streambooleantrue = Server-Sent Events
temperaturenumber0–2, default 0.2
max_tokensintegercaps output length
max_completion_tokensintegeralias of max_tokens
toolsarrayfunction/tool definitions — see Tools
tool_choicestring/objectpassed through to the model
response_formatobject{"type":"json_object"} for JSON mode
top_p, stop, seed, presence_penalty, frequency_penalty, logit_bias, n, user, logprobs, top_logprobs—passed through

messages follows the OpenAI schema (system, user, assistant, tool roles; multimodal content arrays).

Response

{
  "id": "chatcmpl-3f9c1a...",
  "object": "chat.completion",
  "created": 1791000000,
  "model": "xuss/kitsune",
  "choices": [
    {
      "index": 0,
      "message": {"role": "assistant", "content": "Hello! How can I help?"},
      "finish_reason": "stop",
      "logprobs": null
    }
  ],
  "usage": {"prompt_tokens": 12, "completion_tokens": 7, "total_tokens": 19},
  "system_fingerprint": "xuss"
}

finish_reason is stop, length, or tool_calls.

Messages

A standard OpenAI messages array. A minimal single-turn request:

{"model": "xuss/kitsune", "messages": [{"role": "user", "content": "Hi"}]}

Multi-turn conversation — send the full history each time:

{
  "model": "xuss/kitsune",
  "messages": [
    {"role": "system", "content": "You are a concise assistant."},
    {"role": "user", "content": "What is XUSS?"},
    {"role": "assistant", "content": "XUSS is a hosting platform."},
    {"role": "user", "content": "Does it have an AI API?"}
  ]
}

Streaming

Set "stream": true to receive Server-Sent Events. Each line is data: {json}; the stream ends with data: [DONE].

The chunks look like this:

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","created":…,"model":"xuss/kitsune","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}],"system_fingerprint":"xuss"}

data: {"id":"chatcmpl-…","choices":[{"index":0,"delta":{"content":"Silent "},"finish_reason":null}]}

data: {"id":"chatcmpl-…","choices":[{"index":0,"delta":{"content":"racks hum"},"finish_reason":null}]}

data: {"id":"chatcmpl-…","choices":[],"usage":{"prompt_tokens":14,"completion_tokens":9,"total_tokens":23}}

data: [DONE]

Read the stream in any language:

curl -N https://xuss.us/v1/chat/completions \
  -H "Authorization: Bearer xsk-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"xuss/kitsune","stream":true,
       "messages":[{"role":"user","content":"Write a haiku about servers"}]}'
stream = client.chat.completions.create(
    model="xuss/kitsune",
    messages=[{"role": "user", "content": "Write a haiku about servers"}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta.content if chunk.choices else None
    if delta:
        print(delta, end="", flush=True)
const stream = await client.chat.completions.create({
  model: "xuss/kitsune",
  messages: [{ role: "user", content: "Write a haiku about servers" }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content || "");
}
package main

import (
	"bufio"
	"bytes"
	"fmt"
	"net/http"
	"strings"
)

func main() {
	body := []byte(`{"model":"xuss/kitsune","stream":true,"messages":[{"role":"user","content":"Write a haiku about servers"}]}`)
	req, _ := http.NewRequest("POST", "https://xuss.us/v1/chat/completions", bytes.NewReader(body))
	req.Header.Set("Authorization", "Bearer xsk-YOUR_KEY")
	req.Header.Set("Content-Type", "application/json")

	res, _ := http.DefaultClient.Do(req)
	defer res.Body.Close()

	sc := bufio.NewScanner(res.Body)
	for sc.Scan() {
		line := sc.Text()
		if !strings.HasPrefix(line, "data: ") || strings.HasSuffix(line, "[DONE]") {
			continue
		}
		fmt.Println(line[6:]) // parse JSON, read choices[0].delta.content
	}
}
<?php
$ch = curl_init('https://xuss.us/v1/chat/completions');
curl_setopt_array($ch, [
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer xsk-YOUR_KEY',
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode([
        'model' => 'xuss/kitsune',
        'stream' => true,
        'messages' => [['role' => 'user', 'content' => 'Write a haiku about servers']],
    ]),
    CURLOPT_WRITEFUNCTION => function ($ch, $chunk) {
        foreach (explode("\n", $chunk) as $line) {
            if (str_starts_with($line, 'data: ') && !str_contains($line, '[DONE]')) {
                $j = json_decode(substr($line, 6), true);
                echo $j['choices'][0]['delta']['content'] ?? '';
            }
        }
        return strlen($chunk);
    },
]);
curl_exec($ch);
When the model supports reasoning, streamed chunks may additionally carry a reasoning_content field in delta. It is ignored by standard SDKs and never counted as output you must handle.

System prompts and identity

Your system messages are honoured. The model also knows its own display name and, when asked which model it is, answers with that name only — it never reveals any upstream provider or route.

{
  "model": "xuss/kitsune",
  "messages": [
    {"role": "system", "content": "You are a terse DevOps assistant. Answer in bullet points."},
    {"role": "user", "content": "How do I restart a systemd service?"}
  ]
}

Vision (images)

Pass images as OpenAI image_url content parts (data URLs or public http(s) URLs). Only models with "vision": true in /v1/models accept images; sending an image to a non-vision model returns a text note and the model answers on the text only.

r = client.chat.completions.create(
    model="xuss/kitsune",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url", "image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}},
        ],
    }],
)

Limits: up to 8 image parts, each up to ~6 MB; images do not count toward the text budget.

Function calling

Pass OpenAI-style tools. When the model decides to call a function, finish_reason is tool_calls and the message contains tool_calls.

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

r = client.chat.completions.create(
    model="xuss/kitsune",
    messages=[{"role": "user", "content": "Weather in Tashkent?"}],
    tools=tools,
)

call = r.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)

Then send the result back as a tool message:

messages = [
    {"role": "user", "content": "Weather in Tashkent?"},
    r.choices[0].message,
    {"role": "tool", "tool_call_id": call.id, "content": '{"temp_c": 24, "sky": "clear"}'},
]
final = client.chat.completions.create(model="xuss/kitsune", messages=messages)

Server tools (web + skills)

Opt in per request. Add "xuss_tools": true (all of them) or an array of the ones you want. XUSS then runs those tools for you on the server — the model calls them, the server executes them and continues, and you get the final answer in the same response. No client-side loop needed.

r = client.chat.completions.create(
    model="xuss/kitsune",
    xuss_tools=True,                       # or ["web_search", "read_skill"]
    messages=[{"role": "user",
               "content": "Search the web for the latest Python release and summarise it."}],
)
print(r.choices[0].message.content)
{
  "model": "xuss/kitsune",
  "xuss_tools": true,
  "messages": [{"role": "user", "content": "Open https://example.com and give me the page title."}]
}
ToolWhat it does
web_searchWeb search (SearXNG metasearch): titles, URLs and snippets
fetch_urlFetch an HTTP(S) URL server-side and return its text (HTML is converted to text)
browser_renderOpen a page in a real headless browser: HTTP status, title, visible text, console errors, failed requests, layout metrics, optional JS eval
read_skillLoad one of the reference skills below (full how-to doc) into the model's context

These are read-only: no access to files, databases, hostings or your account. Tools that change things (file management, DNS, databases, and so on) stay in your own hands and are intentionally not exposed.

Available skills

read_skill loads the full reference document for any of these into context — just tell the model which one (e.g. "use the product-design skill").

Skill idCovers
telegram-botscomplete Telegram Bot API 10.3: every method/type, aiogram 3 examples, payments/Stars, webhooks, premium emoji/gifts, rich messages & streaming drafts
product-designwebsites/pages/galleries/UI that look intentionally designed: tokens, anti-slop rules, layouts, photo-wallpaper galleries
cybersecuritywrite & audit secure code: injection/XSS/CSRF/IDOR, secrets management, bot security, incident response
telegram-miniappsTelegram WebApps: initData auth, theme vars, MainButton/BackButton, Stars, deployment
shop-botcomplete Telegram store bot: catalog, cart, orders, admin panel, delivery, payment flow
paymentsClick/Payme/Paylov/Uzum + Crypto Pay + Telegram Stars: invoices, webhook verification, idempotency
php-webPHP sites & WordPress: structure, PDO, auth/CSRF, templates, security, deployment
ai-integrationLLMs inside user apps: chat/streaming, RAG, prompting, cost limits, API key security
python-backendproduction Python: bots, FastAPI/Flask, asyncio discipline, DB access, run services, error handling
node-backendNode/TypeScript backends & bots: Express/Fastify/Telegraf, env, process management, errors
databasesMySQL/PostgreSQL/SQLite: schema design, indexes, migrations, transactions, backups
rest-apiAPI design: auth, validation, pagination, one error shape, rate limits, webhook signing
deployment-opsdeploy projects on this hosting: domains/DNS/SSL, reverse proxy, ports, cron, backups, logs
git-githubgit workflows, deploy keys, auto-deploy webhooks, secret hygiene, rollback
scraping-automationethical scrapers & watchers: structured sources, backoff, dedupe, scheduling, alerts
seotechnical SEO: titles, structured data, sitemaps, hreflang, indexing (Google & Yandex)
media-pipelineimages/video pipelines: resizing, WebP, thumbnails, compression, ffmpeg previews
i18n-localizationuz/ru/en + RTL: string dicts, number/date/plural formats, bot language, hreflang
testing-qualityrun/verify/debug discipline, unit tests, linting, review habits before saying done
analytics-monitoringuptime checks, error tracking, daily stats, privacy-safe analytics and alerts
legal-templatesprivacy/terms/refund pages + consent + data-deletion flows (practical baseline)
react-best-practicesReact/Next.js performance: waterfalls, bundles, rendering, hydration, rerenders
mobile-designnative mobile UX: platform patterns, touch psychology, mobile performance
senior-frontendsenior frontend engineering: architecture, components, performance reviews
senior-backendsenior backend engineering: API/DB design, scaling, code reviews
senior-securitysecurity architecture, threat modeling, crypto implementation, audits
ui-design-systemdesign tokens, components, handoff; generate a token system from a brand color
tgbot-cloneclone a Telegram bot's features safely: probe, map features, implement and test
product-layerslayered product/UX method: needs → strategy → conceptual model → surface
find-skillsfind & install more reusable skills for agent projects

Structured output (JSON mode)

Set response_format to {"type":"json_object"} and instruct the model to emit JSON. The reply's content is a JSON string.

r = client.chat.completions.create(
    model="xuss/kitsune",
    response_format={"type": "json_object"},
    messages=[{"role": "user",
               "content": "Return JSON with keys a and b, a=1, b=2"}],
)
import json
print(json.loads(r.choices[0].message.content))  # {'a': 1, 'b': 2}
Always parse defensively and describe the exact shape you want in the prompt.

Billing and token accounting

Cost formula (per request):

cost = cache_miss_tokens × price_in
     + cache_hit_tokens  × price_cache
     + completion_tokens × price_out

(Prices are per token; divide by 1,000,000 for per-million.)

Rate limits and quotas

ScopeDefault
Per user300 requests / minute
Per API key600 requests / minute

Exceeding a limit returns 429 with Retry-After semantics (back off and retry). Request body is capped at 25 MB.

Errors

Errors follow the OpenAI error format:

{
  "error": {
    "message": "Model 'xuss/foo' not found",
    "type": "invalid_request_error",
    "param": "model",
    "code": "model_not_found"
  }
}
Statustype / codeMeaning
400invalid_request_errorBad request (missing/invalid fields, unparseable body)
401authentication_error · invalid_api_key, key_expiredMissing, invalid, or expired API key
402insufficient_quota · key_spend_limit_reachedBalance empty or the key hit its spend limit
403permission_errorAPI not enabled for your account, or account disabled
404invalid_request_error · model_not_foundUnknown model
413invalid_request_errorRequest body too large (> 25 MB)
422invalid_request_errorInvalid request payload (schema)
429rate_limit_error · rate_limit_exceededRate limit exceeded — slow down
500 / 502api_error · upstream_errorTemporary upstream problem — retry with backoff
503api_error · service_unavailableThe AI API is currently disabled

param is set when the error concerns a specific request field. Standard SDKs read error.message (and error.code) directly.

Temporary failures are also delivered as a normal assistant message beginning "The model is temporarily unavailable…" when a stream has already started.

Using with OpenAI-compatible tools

Any client that supports a custom OpenAI base URL works. Set:

Examples:

# open-webui / LibreChat / Cursor / Cline / Continue / LangChain:
#   set the OpenAI base URL to https://xuss.us/v1 and paste your key

Environment-variable style (many tools honour these):

export OPENAI_BASE_URL="https://xuss.us/v1"
export OPENAI_API_KEY="xsk-YOUR_KEY"

LangChain (Python):

from langchain_openai import ChatOpenAI
llm = ChatOpenAI(model="xuss/kitsune", base_url="https://xuss.us/v1",
                 api_key="xsk-YOUR_KEY")
print(llm.invoke("Hello").content)

Advanced parameters

The following are accepted and passed through to the model when supported: top_p, stop, seed, presence_penalty, frequency_penalty, logit_bias, user, n, parallel_tool_calls, tool_choice.

FAQ

Do I need a separate subscription? No. You pay per token from your XUSS balance.

Which model ids do I use? Exactly the ids from GET /v1/models (e.g. xuss/kitsune). Display names like "Kitsune" are shown in the panel.

Can I send images? Yes, to models marked "vision": true.

Is my data used for training? No — requests are proxied to the model to produce your answer and are not used for training.

What happens if a provider is slow or down? The request is transparently retried on alternate capacity; you keep the same model id and pricing.

How do I rotate a key? Create a new key, switch your apps over, then revoke the old one in the panel.

Support