1. Documentation
  2. Developer docs

Call it from your own code.

An OpenAI-compatible API on the same balance — and the same keys — as the console. Every call goes to TokenFusion's endpoint with a TokenFusion key, whichever model it is for — one endpoint, one key format, one balance. Every failure this API can return is listed below with what it means and what to do about it.

Quickstart
  1. Top up your balance — from $1, in whole dollars.
  2. Issue a key on API keys, scoped to the model you want. It is shown once.
  3. Call TokenFusion's endpoint with it:
curl https://tokenfusion.io/v1/chat/completions \ -H "Authorization: Bearer $TOKENFUSION_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "kimi-k3", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 200}'

Every key issued on the API keys page is a TokenFusion key (tf-…) and is used at this endpoint, for every model on your rate card. The API keys page shows this call in cURL, Python, JavaScript and Java, for the key you pick.

Authentication

Pass your key as a Bearer token on every request. Every key is issued for one model, you hold one active key per model at a time (a key in its last week may be replaced early), and a key may carry an expiry; all of this is enforced on the server, so a key that has expired or that names a different model is refused before anything is charged. The console uses your keys too: it lists the models you hold a live key for and makes each call with that key, so a key's usage covers both its console and its API calls.

Authorization: Bearer tf-…

We store a hash of your key, never the key itself. If you lose it, revoke it and issue another — we cannot recover it.

POST /v1/chat/completions
FieldTypeNotes
modelstring A model id from your rate card (GET /v1/models lists them). Required.
messagesarray The usual role/content objects. Must not be empty. content may be a string or an array of parts; on a model marked Vision on your rate card a part may be an image — {"type": "image_url", "image_url": {"url": "data:image/png;base64,…"}} (PNG, JPEG, WebP or GIF). Every other model accepts only text parts, and refuses any other part — an image in any form, a file, audio — with 400 invalid_request: “This model does not accept images.” Each image is held for a fixed 2,500 tokens and charged at the model's published selling rate on the usage reported.
max_tokensinteger Up to 32,768. Sizes the balance reserved before the call runs: its worst case at this many output tokens — if your available balance cannot cover that, the call is refused with 402 and nothing is charged. Leave it out and the model's default output is used, or less when your available balance cannot pay for that much, so the call still runs. With a team key it is never sized to the team's money: leave it out and the model's default output is asked for. Reasoning models spend part of it thinking before they answer.
streamboolean With true the answer arrives as server-sent events in the usual chunk format, ending with data: [DONE]. Ask for stream_options.include_usage to get the usage chunk.

A reasoning model returns its thinking beside the answer in message.reasoning_content (in a stream, delta.reasoning_content); those tokens are output and are charged as output. When the answer stops at max_tokens, finish_reason is "length": send the conversation back with the partial answer and ask the model to continue, or raise max_tokens.

The response carries the usual usage block, plus a tokenfusion object with what the call cost and your balance after it, both in US dollars as plain decimal strings: charged_usd (to six decimal places) and balance_usd. A team key's answers carry no tokenfusion object.

POST /v1/images/generations

For a model marked Creates images on your rate card, with a key issued for it (the chat endpoint refuses such a model, and this one refuses a text model, with 400 invalid_request).

FieldTypeNotes
modelstringAn image model id. Required.
promptstringWhat to draw, up to 4,000 characters. Required.
ninteger1 to 4 images (default 1).
sizestringOne of the model's sizes, listed on your rate card and in GET /v1/models (tokenfusion.sizes); default the first, usually 1024x1024.
response_formatstringurl (default) or b64_json.

The answer is {"created", "data": [{"url"} or {"b64_json"}], "tokenfusion": {"charged_usd", "balance_usd"}} (a team key's answer has no tokenfusion). A url is a TokenFusion link that works without a key for 7 days, then expires — download what you want to keep. Before the call, n times the price per image is held (or, for a model priced on image tokens, the most those images can cost); you are charged per image delivered — or, for a model priced on image tokens, on the image tokens of the images delivered — never more than was held, and nothing when none could be delivered. Images you make count toward the storage your account keeps.

This endpoint is stateless. An image model is given one prompt and remembers nothing of your earlier calls, so each call stands alone: to change an image you have already made, send the whole description again with the change in it — a bare “make them taller” draws only that. The console works differently on purpose: a follow-up message there is treated as a change to the description your last image was made from, and the reply shows the exact prompt that was sent.

POST /v1/audio/speech

For a model marked Speech on your rate card, with a key issued for it. The answer is the audio itself.

FieldTypeNotes
modelstringA speech model id. Required.
inputstringThe text to speak, up to 4,096 characters. Required.
voicestringOne of the model's voices when it lists them (GET /v1/models, tokenfusion.voices; default the first).
response_formatstringmp3 (default), opus, aac, flac, wav or pcm.
speednumberOptional: from a quarter of normal speed up to four times it.

Priced per 1,000 characters of input. Before the call the price of the whole text is held; you are charged for the speech delivered, never more than was held, and nothing when none was. The charge and your balance are in the X-TokenFusion-Charged-USD and X-TokenFusion-Balance-USD headers.

POST /v1/videos · GET /v1/videos/{id} · GET /v1/videos/{id}/content

For a model marked Video. A video takes a while to make, so it is a job: start it, ask for it until its status is completed (every few seconds), then download it.

FieldTypeNotes
modelstringA video model id. Required.
promptstringWhat to film, up to 4,000 characters. Required.
secondsstringHow long, e.g. "5" — one of the model's lengths when it lists them (tokenfusion.seconds), else 1 to 20.
sizestringe.g. "1280x720" — one of the model's sizes when it lists them (tokenfusion.sizes).

Each answer is a video object: {"id": "video_…", "object": "video", "model", "status" (queued, in_progress, completed or failed), "progress", "seconds", "size", "created_at", "completed_at", "error"} — and, once completed, "tokenfusion": {"charged_usd"}. Priced per second of video: when the job starts, its length at the price per second is held; when it completes you are charged for the video delivered, never more than was held. A job that fails, or does not finish within three hours, is charged nothing. The finished video is kept for 7 days — download what you want to keep. It counts toward the storage your account keeps.

GET /v1/models

Returns exactly the models this key can reach — a key scoped to one model lists that one only, so what you discover is what you can actually call. A key that names no model (issued before keys named one) reaches nothing: it is refused with 403 model_not_authorized, as the call endpoints refuse it. Each entry's tokenfusion object carries its price in US dollars per million tokens (price_usd_in_per_m, price_usd_out_per_m) and its serving region (region) — a team key's entries carry no price. A model that creates images, video or speech adds:

FieldTypeNotes
typestring"image", "video" or "audio" (absent for a text model).
endpointstringWhere it is called: /v1/images/generations, /v1/videos or /v1/audio/speech.
price_usd_per_second · price_usd_per_1k_charsstringA video model's price per second of video; a speech model's per 1,000 characters.
seconds · voicesarray or nullA video model's lengths, a speech model's voices, when it lists them.
price_usd_per_imagestring or nullUS dollars per image delivered; null when the model is priced on image tokens instead (then the two per-million prices apply).
sizesarray of stringsThe sizes it can be asked for, the default first (e.g. "1024x1024").
How you are charged
  • Everything is in US dollars: your balance, the rates (published in dollars per million tokens) and every charge.
  • A call starts only when at least $1.00 is available — your balance less what calls already in flight hold. Keep that much in your balance to make calls.
  • Before a call runs, its worst case is held against your balance. The hold is released the moment the real token counts arrive.
  • You are charged on the token counts reported for the call at the model's rate — or, where the model serving it reports its own charge for the call, on that charge, converted to dollars. If a call comes back with neither, it is charged on an estimate — the prompt plus the text actually returned — and your usage page marks it as estimated.
  • A call that is refused costs nothing; a call that fails after the model started answering may still be charged for the part it served.
  • Your balance can reach zero but cannot go negative.
Team keys
  • A member of a team creates team keys on the API keys page, for the models the team admin has chosen. The team pays for their calls; the member's own balance is kept, untouched, until they leave.
  • A team key's answers carry no money: no tokenfusion object, no price in GET /v1/models, and refusals with no amounts — 402 insufficient_balance (the team's balance is too low) and 402 team_limit_reached (the member's monthly limit).
  • max_tokens is never sized to the team's money: leave it out and the model's default output is asked for, so an answer that ends with finish_reason: "length" says nothing about the team's balance or limit.
  • While in a team, the member's own keys are paused (403 key_paused); a key of a membership that has ended is revoked (403 key_revoked); a model the team admin takes off the team refuses its keys (403 model_not_enabled) until it is put back.
Error reference 18 documented failures
HTTPtypeWhat it means What to do
401 invalid_api_key No key, or a key we do not recognise
The Authorization header was missing, malformed, or names a key that has been revoked or has passed its expiry date.
Send `Authorization: Bearer tf-…`. If the key has expired, issue a new one on the API keys page — an expired key cannot be revived.
403 account_not_active The account is suspended
The key is valid but the account behind it is not active, so nothing may be spent on it.
Contact support. No charge is made while an account is suspended.
403 model_not_authorized This key is issued for a different model
Keys are issued per model. A key scoped to one model cannot call another, which is what stops a leaked key from spending across your whole catalogue.
Call the model this key was issued for, or issue a key for the model you want.
404 model_not_found Unknown model
The `model` field does not match any model on your rate card.
GET /v1/models returns exactly the models this key can reach.
409 model_unpriced The model has no published rate
The model is listed but TokenFusion has not set a price for it, so it cannot trade. We refuse rather than guess at a price.
Choose a priced model. The rate card marks anything unpriced.
402 insufficient_balance Not enough balance
The available balance — your balance minus anything held by calls already in flight — is under the $1.00 needed to start a call, or cannot cover the worst case of this one (its maximum output). The error carries `available_usd`, and `required_usd` when the call's maximum was the problem.
Top up, or lower `max_tokens`. A balance can reach zero but never goes negative.
402 team_limit_reached A team member's monthly limit is reached
The key is a team key and its holder has reached the monthly limit their team admin set ("You've reached your team's monthly limit. Ask your team admin."). No amounts are given. Nothing was charged.
Ask your team admin to raise or remove the limit, or wait for the next month (UTC).
403 key_paused Your own key is paused while you're in a team
While you are a member of a team, the keys you created for yourself are paused (not revoked): your team's keys are used instead. Nothing was charged.
Use a team key from the API keys page. Your own keys work again when you leave the team.
403 key_revoked The key's team membership has ended
The key was created under a team its holder is no longer in (they left, were removed, or the team was archived), so it has been revoked. Nothing was charged.
Use your own keys again, or a key of the team you are in now.
403 model_not_enabled The model isn't enabled for your team
The team admin has taken this model off the team. The key is not revoked: it works again if the model is put back. Nothing was charged.
Ask your team admin, or use a key for one of your team's models.
403 team_unavailable Your team can't make calls right now
The team's account is not active, so it cannot pay for calls. Nothing was charged.
Ask your team admin.
503 no_supplier_connection No live supplier for this model
The model is priced, but the supplier that serves it has no live connection right now. Nothing was called and nothing was charged.
Retry later, or use a model whose card shows Available. This platform refuses in this situation rather than returning a fabricated answer.
400 invalid_request The request body was not usable
`messages` was missing or empty, a parameter had the wrong type, or `n` asked for more than one completion — then the message names the problem — or the model itself refused the request as invalid, which comes back with a general message: the model's own words are not passed on.
Send a non-empty `messages` array and fix the parameter the error names; for a general message, check the request against what the model accepts.
429 rate_limited Too many requests
Calls are arriving faster than the supplier accepts for this account. The call was not run and nothing was charged.
Retry after a short wait (honour Retry-After when it is sent), with backoff rather than more concurrency.
502 upstream_error The supplier failed the call
The call reached the supplier and the supplier failed it, or the connection to the supplier broke before its answer arrived. A failure before any output costs nothing; a stream that breaks part-way is charged for what arrived, and its error says so.
Retry. If it persists, the supplier is having an incident.
503 temporarily_unavailable The model is temporarily unavailable
The model is on sale, but it cannot be served for a moment — its access is being set up, or the capacity behind it is being topped up. Nothing was charged, and we have been told.
Retry after a short wait (honour Retry-After when it is sent).
502 upstream_unavailable The model cannot be reached
The supplier refused TokenFusion's own access to the model. Nothing was charged, and we have been told.
Retry later.
504 upstream_timeout The model did not answer in time
The supplier did not answer before the time limit. Nothing was charged unless part of a streamed answer arrived — then that part is.
Retry, or ask for a shorter answer (a lower `max_tokens`).

Every error is returned as {"error": {"message": …, "type": …}}. This table is checked against the running application by the test suite, so it cannot drift away from what the API actually does.