SonicBit AI Model API
An OpenAI-compatible API for DeepSeek, GLM and Kimi models, sold as prepaid token packs. Your key works anywhere the OpenAI API does: point any OpenAI SDK or app at our base URL, use your sk-sb- key and pick a model. DeepSeek-style clients work too.
- Base URL
https://api.sonicb.it/v1- API key
sk-sb-…comes with your pack- Default model
deepseek-v4-flash
Get an API key
Keys come with a prepaid token pack, from $3 for 50M tokens. Message us on Telegram or contact support to buy one, and your key is delivered with the purchase. See packs & pricing
Your first request
Send a chat message and print the reply. The OpenAI SDKs need only the base URL and your key.
SONICBIT_API_KEY as the samples below do, rather than in your code.curl https://api.sonicb.it/v1/chat/completions \
-H "Authorization: Bearer $SONICBIT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-flash",
"messages": [{"role": "user", "content": "Hello! What can you do?"}]
}'# pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.sonicb.it/v1",
api_key=os.environ["SONICBIT_API_KEY"],
)
reply = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Hello! What can you do?"}],
)
print(reply.choices[0].message.content)// npm install openai
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://api.sonicb.it/v1",
apiKey: process.env.SONICBIT_API_KEY,
});
const reply = await client.chat.completions.create({
model: "deepseek-v4-flash",
messages: [{ role: "user", content: "Hello! What can you do?" }],
});
console.log(reply.choices[0].message.content);A system message works the usual way: put your instructions first, as {"role": "system", "content": "…"}.
Streaming
Add stream: true to receive the reply as it is written. Add stream_options: {"include_usage": true} to get the token count in the last chunk.
stream = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "Write a haiku about the sea."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)const stream = await client.chat.completions.create({
model: "deepseek-v4-flash",
messages: [{ role: "user", content: "Write a haiku about the sea." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Use it in an app
Most chat and coding apps that support an OpenAI-compatible API (a “custom provider” or “OpenAI-compatible” setting) work with your key:
- Base URL:
https://api.sonicb.it/v1 - API key: your
sk-sb-key - Model: a name from the list below, for example
deepseek-v4-flash
Apps set up for DeepSeek can keep the model name deepseek-chat. Apps that show a balance read it from the endpoints under Check your balance.
Models
Every model costs the same per token, streams, and supports function calling. GET /v1/models lists them.
| Model | Good for | Context | Vision | Thinking |
|---|---|---|---|---|
deepseek-v4-flash | Fast everyday chat and coding. The default. | 1M | No | No |
deepseek-chat | The same model as deepseek-v4-flash, for apps preset to DeepSeek. | 1M | No | No |
deepseek-v4.1-flash | Fast, and reads images. | 1M | Yes | No |
deepseek-v4-pro | Stronger reasoning and coding. | 512K | No | No |
glm-5.3 | Works through a problem step by step before answering. | 512K | No | Yes |
kimi-k2.7 | Step-by-step answers, and reads images. | 128K | Yes | Yes |
kimi-k3 | The most capable. Slower on long answers. | 256K | Yes | Yes |
Context is the longest prompt we have tested end to end for each model, counted in tokens (1M = one million). Very long prompts take a while before the answer starts: about 30 seconds for a million tokens on deepseek-v4-flash.
Images
Send pictures to deepseek-v4.1-flash, kimi-k2.7 or kimi-k3 as parts of a user message, next to your text.
import base64
with open("photo.jpg", "rb") as f:
photo = base64.b64encode(f.read()).decode()
reply = client.chat.completions.create(
model="kimi-k2.7",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this picture?"},
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{photo}"}},
],
}],
max_tokens=1500,
)
print(reply.choices[0].message.content)- An image can be a
data:image/…;base64,URL (PNG, JPEG, WebP or GIF) or anhttps://link. - Up to 8 images per request, in user messages only.
- A request can be up to 4 MB, which is about 3 MB of images once encoded. Shrink large photos, or send https links instead.
- Other models answer an image with 400
unsupported_content.
Thinking models
glm-5.3, kimi-k2.7 and kimi-k3 reason before they answer. The reasoning arrives in reasoning_content, separately from the answer in content, the same way DeepSeek’s reasoning model returns it.
reply = client.chat.completions.create(
model="glm-5.3",
messages=[{"role": "user", "content": "Is 391 a prime number?"}],
max_tokens=2000,
)
message = reply.choices[0].message
print("Reasoning:", message.reasoning_content)
print("Answer:", message.content)- When streaming, the reasoning comes as
delta.reasoning_contentbefore the answer. - Reasoning tokens count toward your usage and toward
max_tokens. Give these models room, for examplemax_tokensof 2000 or more, or the answer can be cut short. - You don’t need to send the reasoning back in later messages.
Function calling
Describe your functions in tools. When the model wants one, the reply carries tool_calls with the function name and its arguments as JSON text.
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
messages = [{"role": "user", "content": "What's the weather in Kuala Lumpur?"}]
reply = client.chat.completions.create(model="deepseek-v4-flash", messages=messages, tools=tools)
call = reply.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments) # get_weather {"city": "Kuala Lumpur"}
# Run the function yourself, then send its result back for the final answer:
messages.append(reply.choices[0].message)
messages.append({"role": "tool", "tool_call_id": call.id, "content": '{"temp_c": 31, "sky": "sunny"}'})
final = client.chat.completions.create(model="deepseek-v4-flash", messages=messages, tools=tools)
print(final.choices[0].message.content)tool_choice={"type": "function", "function": {"name": "get_weather"}}. This works on every model. "auto" (the default), "none" and "required" work too.Packs & pricing
Keys run on prepaid token packs, priced in US dollars. Each request uses its prompt tokens plus the tokens it generates, reasoning included, and every model costs the same per token.
| Pack | Price | Tokens | Requests / min | At the same time | Valid for |
|---|---|---|---|---|---|
| 50M | $3 | 50,000,000 | 10 | 2 | 30 days |
| 100M | $5 | 100,000,000 | 15 | 3 | 30 days |
| 200M | $9 | 200,000,000 | 20 | 3 | 30 days |
| 500M | $20 | 500,000,000 | 30 | 5 | 30 days |
- Each pack lasts 30 days from the day you buy it. Tokens from your oldest pack are used first, and tokens left in a pack expire with it.
- Buying another pack adds its tokens to your balance and raises your limits to that pack’s, if they are higher.
- All keys on your account share one balance and one set of limits.
How to buy
Packs are sold by hand for now. Message SonicBit on Telegram or contact support with the pack you want; your key is delivered with the purchase.
Check your balance
curl https://api.sonicb.it/user/balance \
-H "Authorization: Bearer $SONICBIT_API_KEY"The reply has the DeepSeek balance format: balance_infos[0].total_tokens is the number of tokens you have left. Beside it, next_expiry_at (Unix time, in seconds) and next_expiry_tokens tell you when your next pack expires and how many tokens expire with it.
Apps that read OpenAI-style billing can use /v1/dashboard/billing/subscription and /v1/dashboard/billing/usage.
Errors
Errors come back in the OpenAI format, {"error": {"message": …, "type": …, "code": …}}. The code tells you what happened.
| Status | Code | What to do |
|---|---|---|
| 400 | invalid_request_error | Fix the request; the message names the field. |
| 400 | unsupported_content | This model doesn't take images. Use an image model. |
| 401 | invalid_api_key | Check the key: it starts with sk-sb-. |
| 402 | insufficient_balance | Your tokens are used up or expired. Buy a pack. |
| 402 | key_spend_cap_reached | This key reached the spending cap set for it. Use another key on the account. |
| 403 | key_disabledaccount_disabled | Contact support. |
| 404 | model_not_found | Check the model name against GET /v1/models. |
| 413 | payload_too_large | The request is over 4 MB. Shrink images or send https links. |
| 429 | rate_limit_exceededkey_rate_limit_exceededconcurrency_limit | Slow down. Wait for the Retry-After header, or send fewer requests at once. |
| 502 | upstream_error | A temporary failure. Retry. |
| 503 | server_busyupstream_unavailable | Temporarily busy. Retry after the Retry-After header. |
The OpenAI SDKs retry busy and temporary errors for you. A request that fails before the model starts answering is not charged.

