SonicBitSign inGet 4 GB free
Reference
For developersUpdated 2026-09-30

SonicBit AI Model API

An OpenAI-compatible API for DeepSeek, GLM and Kimi models, sold as prepaid token packs. Your key works anywhere the OpenAI API does: point any OpenAI SDK or app at our base URL, use your sk-sb- key and pick a model. DeepSeek-style clients work too.

Base URL
https://api.sonicb.it/v1
API key
sk-sb-…comes with your pack
Default model
deepseek-v4-flash

Get an API key

Keys come with a prepaid token pack, from $3 for 50M tokens. Message us on Telegram or contact support to buy one, and your key is delivered with the purchase. See packs & pricing

Your first request

Send a chat message and print the reply. The OpenAI SDKs need only the base URL and your key.

Keep your key private: anyone who has it can spend your tokens. Store it in an environment variable, for example SONICBIT_API_KEY as the samples below do, rather than in your code.
curl https://api.sonicb.it/v1/chat/completions \
  -H "Authorization: Bearer $SONICBIT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "messages": [{"role": "user", "content": "Hello! What can you do?"}]
  }'

A system message works the usual way: put your instructions first, as {"role": "system", "content": "…"}.

Streaming

Add stream: true to receive the reply as it is written. Add stream_options: {"include_usage": true} to get the token count in the last chunk.

stream = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Write a haiku about the sea."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Use it in an app

Most chat and coding apps that support an OpenAI-compatible API (a “custom provider” or “OpenAI-compatible” setting) work with your key:

  • Base URL: https://api.sonicb.it/v1
  • API key: your sk-sb- key
  • Model: a name from the list below, for example deepseek-v4-flash

Apps set up for DeepSeek can keep the model name deepseek-chat. Apps that show a balance read it from the endpoints under Check your balance.

Models

Every model costs the same per token, streams, and supports function calling. GET /v1/models lists them.

ModelGood forContextVisionThinking
deepseek-v4-flashFast everyday chat and coding. The default.1MNoNo
deepseek-chatThe same model as deepseek-v4-flash, for apps preset to DeepSeek.1MNoNo
deepseek-v4.1-flashFast, and reads images.1MYesNo
deepseek-v4-proStronger reasoning and coding.512KNoNo
glm-5.3Works through a problem step by step before answering.512KNoYes
kimi-k2.7Step-by-step answers, and reads images.128KYesYes
kimi-k3The most capable. Slower on long answers.256KYesYes

Context is the longest prompt we have tested end to end for each model, counted in tokens (1M = one million). Very long prompts take a while before the answer starts: about 30 seconds for a million tokens on deepseek-v4-flash.

Images

Send pictures to deepseek-v4.1-flash, kimi-k2.7 or kimi-k3 as parts of a user message, next to your text.

Python
import base64

with open("photo.jpg", "rb") as f:
    photo = base64.b64encode(f.read()).decode()

reply = client.chat.completions.create(
    model="kimi-k2.7",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this picture?"},
            {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{photo}"}},
        ],
    }],
    max_tokens=1500,
)
print(reply.choices[0].message.content)
  • An image can be a data:image/…;base64, URL (PNG, JPEG, WebP or GIF) or an https:// link.
  • Up to 8 images per request, in user messages only.
  • A request can be up to 4 MB, which is about 3 MB of images once encoded. Shrink large photos, or send https links instead.
  • Other models answer an image with 400 unsupported_content.

Thinking models

glm-5.3, kimi-k2.7 and kimi-k3 reason before they answer. The reasoning arrives in reasoning_content, separately from the answer in content, the same way DeepSeek’s reasoning model returns it.

Python
reply = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Is 391 a prime number?"}],
    max_tokens=2000,
)
message = reply.choices[0].message
print("Reasoning:", message.reasoning_content)
print("Answer:", message.content)
  • When streaming, the reasoning comes as delta.reasoning_content before the answer.
  • Reasoning tokens count toward your usage and toward max_tokens. Give these models room, for example max_tokens of 2000 or more, or the answer can be cut short.
  • You don’t need to send the reasoning back in later messages.

Function calling

Describe your functions in tools. When the model wants one, the reply carries tool_calls with the function name and its arguments as JSON text.

Python
tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]
messages = [{"role": "user", "content": "What's the weather in Kuala Lumpur?"}]

reply = client.chat.completions.create(model="deepseek-v4-flash", messages=messages, tools=tools)
call = reply.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)  # get_weather {"city": "Kuala Lumpur"}

# Run the function yourself, then send its result back for the final answer:
messages.append(reply.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": '{"temp_c": 31, "sky": "sunny"}'})
final = client.chat.completions.create(model="deepseek-v4-flash", messages=messages, tools=tools)
print(final.choices[0].message.content)
To make sure the model calls a function, name it: tool_choice={"type": "function", "function": {"name": "get_weather"}}. This works on every model. "auto" (the default), "none" and "required" work too.

Packs & pricing

Keys run on prepaid token packs, priced in US dollars. Each request uses its prompt tokens plus the tokens it generates, reasoning included, and every model costs the same per token.

PackPriceTokensRequests / minAt the same timeValid for
50M$350,000,00010230 days
100M$5100,000,00015330 days
200M$9200,000,00020330 days
500M$20500,000,00030530 days
  • Each pack lasts 30 days from the day you buy it. Tokens from your oldest pack are used first, and tokens left in a pack expire with it.
  • Buying another pack adds its tokens to your balance and raises your limits to that pack’s, if they are higher.
  • All keys on your account share one balance and one set of limits.

How to buy

Packs are sold by hand for now. Message SonicBit on Telegram or contact support with the pack you want; your key is delivered with the purchase.

Check your balance

curl
curl https://api.sonicb.it/user/balance \
  -H "Authorization: Bearer $SONICBIT_API_KEY"

The reply has the DeepSeek balance format: balance_infos[0].total_tokens is the number of tokens you have left. Beside it, next_expiry_at (Unix time, in seconds) and next_expiry_tokens tell you when your next pack expires and how many tokens expire with it.

Apps that read OpenAI-style billing can use /v1/dashboard/billing/subscription and /v1/dashboard/billing/usage.

Errors

Errors come back in the OpenAI format, {"error": {"message": …, "type": …, "code": …}}. The code tells you what happened.

StatusCodeWhat to do
400invalid_request_errorFix the request; the message names the field.
400unsupported_contentThis model doesn't take images. Use an image model.
401invalid_api_keyCheck the key: it starts with sk-sb-.
402insufficient_balanceYour tokens are used up or expired. Buy a pack.
402key_spend_cap_reachedThis key reached the spending cap set for it. Use another key on the account.
403key_disabledaccount_disabledContact support.
404model_not_foundCheck the model name against GET /v1/models.
413payload_too_largeThe request is over 4 MB. Shrink images or send https links.
429rate_limit_exceededkey_rate_limit_exceededconcurrency_limitSlow down. Wait for the Retry-After header, or send fewer requests at once.
502upstream_errorA temporary failure. Retry.
503server_busyupstream_unavailableTemporarily busy. Retry after the Retry-After header.

The OpenAI SDKs retry busy and temporary errors for you. A request that fails before the model starts answering is not charged.

Was this helpful?Still stuck? Contact support