Appearance
Text
Send chat messages through RemoteGPU's Text API to generate a reply in your application or script. Include a model ID in every request to choose which model answers.
Send requests
Quickstart chat completion
Create a key with Token Factory access using API keys, and add credits to your wallet. Set REMOTEGPU_API_KEY in the environment where your client runs. This example uses GPT-6 Luna with an explicit output limit.
bash
curl -X POST "https://inference.remotegpu.ai/v1/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $REMOTEGPU_API_KEY" \
-d '{
"model": "gpt-6-luna",
"messages": [
{
"role": "user",
"content": "Write a one-sentence greeting."
}
],
"max_completion_tokens": 256
}'Replace the example message with your prompt. Without stream: true, the API waits for the model to finish and returns a chat.completion object. Read the reply from choices[0].message.content.
json
{
"id": "chatcmpl-example",
"object": "chat.completion",
"created": 1779174000,
"model": "gpt-6-luna",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 32,
"completion_tokens": 48,
"total_tokens": 80,
"prompt_tokens_details": {
"cached_tokens": 0,
"cache_write_tokens": 0
}
}
}The response above is illustrative. Token counts and cache use depend on the request. Read Billing before using the token counts to calculate a charge.
How requests work
Authentication
Send a Token Factory key in the Authorization header as a Bearer token on every inference request. You can create or replace a key in API keys.
If the key is missing or invalid, the API returns 401. If the key is valid but does not allow inference APIs, the API returns 403.
Connect an OpenAI-compatible client
Use https://inference.remotegpu.ai/v1 as the base URL. The Token Factory documentation path does not change this API address.
| Endpoint | Purpose |
|---|---|
GET /v1/models | List the available text model IDs; requires your API key |
POST /v1/chat/completions | Generate a reply, with an optional SSE stream |
GET /v1/inference/models | Read model parameters and serving status; public, no key required |
With the OpenAI Python SDK, set your RemoteGPU key and base URL:
python
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["REMOTEGPU_API_KEY"],
base_url="https://inference.remotegpu.ai/v1",
max_retries=0,
)
completion = client.chat.completions.create(
model="gpt-6-luna",
messages=[{"role": "user", "content": "Write a one-sentence greeting."}],
max_completion_tokens=256,
)
print(completion.choices[0].message.content)The example disables the SDK's automatic retries with max_retries=0. Choose your own retry policy after considering the possibility of duplicate requests and charges.
These models use the Chat Completions format, including Claude and Gemini. This integration does not expose the native Anthropic Messages API, Google GenerateContent API, or OpenAI Responses API. OpenAI-compatible clients must send only fields supported by the selected model.
Choose a model first
Call GET /v1/models to choose an available text model ID. Read the public catalog at GET /v1/inference/models for the selected model's parameters, defaults, and limits. The model table and parameter tables below are generated from that catalog. Check the live catalog when you send requests because availability can change after this page is published.
| State | Description |
|---|---|
ready | Ready to serve requests |
starting | The model is starting |
sleeping | The model needs to start before serving requests |
A request to a model in sleeping state can take longer while the model starts. Models that are not available are omitted from the catalog.
OpenAI has deprecated gpt-5.1 and scheduled its API shutdown for April 1, 2027. It recommends gpt-6-sol as the replacement. Plan your migration before that date and check the live RemoteGPU catalog for availability. See OpenAI's deprecation schedule.
Message content and supported features
Send text in messages[].content as a string or a list of text parts:
json
{
"role": "user",
"content": "Write a short summary."
}json
{
"role": "user",
"content": [{ "type": "text", "text": "Write a short summary." }]
}The listed models accept text input. Image input, tool calls, structured-output controls, and reasoning controls are not enabled for these models. A feature supported by a model's official API is not necessarily available through RemoteGPU. Unsupported fields return an error before the request is sent to the model.
GPT-6's model-specific fields are messages, max_completion_tokens, and stream. Requests also require model and can use the endpoint-level stream_options.include_usage setting. Do not send max_tokens, temperature, top_p, or reasoning_effort for GPT-6.
Response and timeouts
Non-streaming responses
A successful non-streaming request returns 200 OK with a JSON response. Read the generated text from choices[0].message.content and token counts from usage. The output limit includes the model's completion tokens, which can include reasoning tokens as well as visible text. A small limit can end a response before it contains visible text.
Streaming responses
Set stream: true to receive server-sent events (SSE) with content type text/event-stream. With curl, use -N to display events without buffering:
bash
curl -N -X POST "https://inference.remotegpu.ai/v1/chat/completions" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $REMOTEGPU_API_KEY" \
-d '{
"model": "gpt-6-luna",
"messages": [
{
"role": "user",
"content": "Write a one-sentence greeting."
}
],
"max_completion_tokens": 256,
"stream": true,
"stream_options": {
"include_usage": true
}
}'Each data: event contains a JSON chunk. Append text from choices[].delta.content when it is present. The final usage chunk can have an empty choices array, so handle usage separately from generated text. The stream ends with data: [DONE].
Token usage is needed to finalize billing, and the API requests it for streamed responses. You can send stream_options: {"include_usage": true} to receive the final usage chunk in your client. If you omit it or set include_usage to false, the API uses that chunk for billing but does not forward it to your client.
Keep your connection open until [DONE]. If the connection closes early, you might not receive complete usage. After response headers have been sent, a stream failure cannot change the HTTP status; an initial 200 alone does not mean the generation completed.
Timeouts and retries
Set client timeouts long enough for the output you request. A request that times out before response headers are sent can return 504. A streaming request can end after its headers have been sent, so also check whether the stream completed.
RemoteGPU does not automatically resubmit a request after sending it to the model. A client retry creates another request and can incur another charge.
Common status codes
| Status code | Description |
|---|---|
400 | Unsupported field or value for the selected model or endpoint |
401 | Missing, invalid, revoked, or expired API key |
402 | Insufficient available wallet credit for the reservation |
403 | API key is valid but not authorized for inference APIs |
404 | The selected model does not exist or is not available |
422 | Request validation failed, such as a missing model field |
429 | Model capacity or a rate limit was reached; retry later |
502 | The model provider returned an error or an invalid response |
503 | The model or its billing configuration cannot serve the request |
504 | The request timed out before a response was available |
Errors include an OpenAI-style error object. Check error.message, error.code, and error.param when present to identify what to change.
Billing
Wallet reservations
Before sending a request to the model, RemoteGPU reserves wallet credit for the model's maximum input allowance and your output limit. The input allowance uses the highest applicable input or cache-write rate. For models with long-context pricing, the reservation also covers the applicable long-context rates.
Send max_completion_tokens to limit the output portion of the reservation. For models that accept max_tokens, you can use that field instead, but do not send both. If you omit the output limit, the reservation uses the model's maximum output allowance. With n, it covers that limit for each completion. The input portion still covers the model's maximum input allowance, so even a short prompt can require a substantial reservation.
This reservation reduces available wallet credit while the request is in progress. When complete usage is available, RemoteGPU charges for actual usage and releases unused reserved credit. The reservation is not the final charge, and RemoteGPU does not insert an output limit into your request when you omit one. If usage is incomplete, including an interrupted stream, the reservation can remain held until billing is reconciled. Contact support if a hold remains after a failed request.
Token and cache charges
See Text pricing for customer rates in USD per 1 million tokens. Input, cached input, cache writes, and output can have different rates. Use only the categories that apply to the selected model.
usage.prompt_tokens includes cached input and cache writes. Those quantities are charged in their respective categories rather than charged again as ordinary input. usage.completion_tokens is the output quantity used for billing, including reasoning tokens when the model reports them as completion tokens.
| Usage field | Meaning |
|---|---|
prompt_tokens | Total input tokens, including cache reads and writes |
completion_tokens | Total completion tokens |
total_tokens | Aggregate token count, not a separate charge |
prompt_tokens_details.cached_tokens | Input tokens read from cache |
prompt_tokens_details.cache_write_tokens | Input tokens written to cache, including GPT-6 cache writes |
claude_cache_creation_5_m_tokens | Claude cache writes with a five-minute lifetime |
claude_cache_creation_1_h_tokens | Claude cache writes with a one-hour lifetime |
Cache fields depend on the model. Claude's five-minute and one-hour writes are priced separately. Cache use comes from reported usage; the Text API does not expose cache-management controls for these models.
For GPT-6, requests with more than 272,000 input tokens use long-context rates for the entire request, including input, cached input, cache writes, and output. At or below that threshold, standard-context rates apply. The total input count determines the tier, before subtracting cache reads or writes.
Rates are refreshed from official price sources. Each accepted request uses the pricing version selected for that request. If valid pricing is unavailable, the API rejects the request instead of using a hard-coded price.
Reference
Text models
This table lists models published in the public catalog. A missing message count limit means the catalog does not publish a count cap; token and request size limits still apply. The maximum output is a limit, not an automatically inserted default.
| Model | Max messages | Max output tokens | Streaming |
|---|---|---|---|
claude-opus-4-8 | 1,000 | 128,000 | Yes |
claude-sonnet-4-6 | 1,000 | 128,000 | Yes |
gemini-3.5-flash | 1,000 | 65,536 | Yes |
gpt-5 | 1,000 | 128,000 | Yes |
gpt-5-mini | 1,000 | 128,000 | Yes |
gpt-5.1 | 1,000 | 128,000 | Yes |
gpt-6-astra | Not published | 128,000 | Yes |
gpt-6-luna | Not published | 128,000 | Yes |
gpt-6-sol | Not published | 128,000 | Yes |
Model parameters
Every request requires model and messages. Expand a model below for its accepted fields and limits. Numeric limits are inclusive. For messages and stop, the limits count list items. A missing default means the catalog does not specify one; it does not mean the field is required.
The catalog type chat_messages means an array of message objects. string_list accepts a string or an array of strings for stop.
| Field | Purpose |
|---|---|
max_completion_tokens or max_tokens | Limit completion tokens; send only one of these fields |
temperature | Control sampling randomness within the model's accepted range |
top_p | Restrict sampling to tokens within a cumulative probability |
n | Request this many completions, each with its own output allowance |
stop | Stop generation at a matching string |
frequency_penalty | Adjust sampling according to how often a token has appeared |
presence_penalty | Adjust sampling according to whether a token has appeared |
This list explains field meanings. It does not mean every model accepts all of them; check the model's table before adding optional fields.
stream defaults to false at the API endpoint. stream_options.include_usage controls whether the final usage chunk is returned to your streaming client. These endpoint behaviors apply in addition to the model-specific fields below.
claude-opus-4-8
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
max_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | 1 to 1,000 items | Not specified |
n | integer | No | 1 to 1 | Not specified |
stop | string_list | No | At most 4 items | Not specified |
stream | boolean | No | Not published | Not specified |
claude-sonnet-4-6
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
max_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | 1 to 1,000 items | Not specified |
n | integer | No | 1 to 1 | Not specified |
stop | string_list | No | At most 4 items | Not specified |
stream | boolean | No | Not published | Not specified |
temperature | number | No | 0 to 1 | Not specified |
gemini-3.5-flash
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
max_completion_tokens | integer | No | 1 to 65,536 | Not specified |
max_tokens | integer | No | 1 to 65,536 | Not specified |
messages | chat_messages | Yes | 1 to 1,000 items | Not specified |
n | integer | No | 1 to 1 | Not specified |
stream | boolean | No | Not published | Not specified |
temperature | number | No | 0 to 2 | Not specified |
top_p | number | No | 0 to 1 | Not specified |
gpt-5
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
max_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | 1 to 1,000 items | Not specified |
n | integer | No | 1 to 4 | Not specified |
stream | boolean | No | Not published | Not specified |
temperature | number | No | 0 to 2 | Not specified |
top_p | number | No | 0 to 1 | Not specified |
gpt-5-mini
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
frequency_penalty | number | No | -2 to 2 | Not specified |
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
max_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | 1 to 1,000 items | Not specified |
n | integer | No | 1 to 4 | Not specified |
presence_penalty | number | No | -2 to 2 | Not specified |
stop | string_list | No | At most 4 items | Not specified |
stream | boolean | No | Not published | Not specified |
temperature | number | No | 0 to 2 | Not specified |
top_p | number | No | 0 to 1 | Not specified |
gpt-5.1
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
frequency_penalty | number | No | -2 to 2 | Not specified |
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
max_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | 1 to 1,000 items | Not specified |
n | integer | No | 1 to 4 | Not specified |
presence_penalty | number | No | -2 to 2 | Not specified |
stream | boolean | No | Not published | Not specified |
temperature | number | No | 0 to 2 | Not specified |
top_p | number | No | 0 to 1 | Not specified |
gpt-6-astra
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | At least 1 item; no maximum published | Not specified |
stream | boolean | No | Not published | Not specified |
gpt-6-luna
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | At least 1 item; no maximum published | Not specified |
stream | boolean | No | Not published | Not specified |
gpt-6-sol
| Field | Type | Required | Limits | Catalog default |
|---|---|---|---|---|
max_completion_tokens | integer | No | 1 to 128,000 | Not specified |
messages | chat_messages | Yes | At least 1 item; no maximum published | Not specified |
stream | boolean | No | Not published | Not specified |
Model catalog endpoint
Read model parameters and serving status without an API key:
bash
curl "https://inference.remotegpu.ai/v1/inference/models"The response includes both Text and Image models. For Text, read:
text[].model: the identifier to send inPOST /v1/chat/completions.text[].parameters: accepted fields, policies, defaults, and limits.text[].runtime.state: whether the model is ready, starting, or sleeping.text[].recent_summary_stats: recent request wait, generation, and total times.
Use this catalog to populate model selectors and validate optional fields. Use the authenticated GET /v1/models endpoint when an OpenAI-compatible client only needs model IDs.
Read next
- Read API keys to create or rotate Token Factory keys.
- Read Token Factory overview for Text, Image, and BytePlus API guides.