Skip to content

Text ​

Send chat messages through RemoteGPU's Text API to generate a reply in your application or script. Include a model ID in every request to choose which model answers.

Send requests ​

Quickstart chat completion ​

Create a key with Token Factory access using API keys, and add credits to your wallet. Set REMOTEGPU_API_KEY in the environment where your client runs. This example uses GPT-6 Luna with an explicit output limit.

bash
curl -X POST "https://inference.remotegpu.ai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $REMOTEGPU_API_KEY" \
  -d '{
    "model": "gpt-6-luna",
    "messages": [
      {
        "role": "user",
        "content": "Write a one-sentence greeting."
      }
    ],
    "max_completion_tokens": 256
  }'

Replace the example message with your prompt. Without stream: true, the API waits for the model to finish and returns a chat.completion object. Read the reply from choices[0].message.content.

json
{
  "id": "chatcmpl-example",
  "object": "chat.completion",
  "created": 1779174000,
  "model": "gpt-6-luna",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! How can I help?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 32,
    "completion_tokens": 48,
    "total_tokens": 80,
    "prompt_tokens_details": {
      "cached_tokens": 0,
      "cache_write_tokens": 0
    }
  }
}

The response above is illustrative. Token counts and cache use depend on the request. Read Billing before using the token counts to calculate a charge.

How requests work ​

Authentication ​

Send a Token Factory key in the Authorization header as a Bearer token on every inference request. You can create or replace a key in API keys.

If the key is missing or invalid, the API returns 401. If the key is valid but does not allow inference APIs, the API returns 403.

Connect an OpenAI-compatible client ​

Use https://inference.remotegpu.ai/v1 as the base URL. The Token Factory documentation path does not change this API address.

EndpointPurpose
GET /v1/modelsList the available text model IDs; requires your API key
POST /v1/chat/completionsGenerate a reply, with an optional SSE stream
GET /v1/inference/modelsRead model parameters and serving status; public, no key required

With the OpenAI Python SDK, set your RemoteGPU key and base URL:

python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["REMOTEGPU_API_KEY"],
    base_url="https://inference.remotegpu.ai/v1",
    max_retries=0,
)

completion = client.chat.completions.create(
    model="gpt-6-luna",
    messages=[{"role": "user", "content": "Write a one-sentence greeting."}],
    max_completion_tokens=256,
)
print(completion.choices[0].message.content)

The example disables the SDK's automatic retries with max_retries=0. Choose your own retry policy after considering the possibility of duplicate requests and charges.

These models use the Chat Completions format, including Claude and Gemini. This integration does not expose the native Anthropic Messages API, Google GenerateContent API, or OpenAI Responses API. OpenAI-compatible clients must send only fields supported by the selected model.

Choose a model first ​

Call GET /v1/models to choose an available text model ID. Read the public catalog at GET /v1/inference/models for the selected model's parameters, defaults, and limits. The model table and parameter tables below are generated from that catalog. Check the live catalog when you send requests because availability can change after this page is published.

StateDescription
readyReady to serve requests
startingThe model is starting
sleepingThe model needs to start before serving requests

A request to a model in sleeping state can take longer while the model starts. Models that are not available are omitted from the catalog.

OpenAI has deprecated gpt-5.1 and scheduled its API shutdown for April 1, 2027. It recommends gpt-6-sol as the replacement. Plan your migration before that date and check the live RemoteGPU catalog for availability. See OpenAI's deprecation schedule.

Message content and supported features ​

Send text in messages[].content as a string or a list of text parts:

json
{
  "role": "user",
  "content": "Write a short summary."
}
json
{
  "role": "user",
  "content": [{ "type": "text", "text": "Write a short summary." }]
}

The listed models accept text input. Image input, tool calls, structured-output controls, and reasoning controls are not enabled for these models. A feature supported by a model's official API is not necessarily available through RemoteGPU. Unsupported fields return an error before the request is sent to the model.

GPT-6's model-specific fields are messages, max_completion_tokens, and stream. Requests also require model and can use the endpoint-level stream_options.include_usage setting. Do not send max_tokens, temperature, top_p, or reasoning_effort for GPT-6.

Response and timeouts ​

Non-streaming responses ​

A successful non-streaming request returns 200 OK with a JSON response. Read the generated text from choices[0].message.content and token counts from usage. The output limit includes the model's completion tokens, which can include reasoning tokens as well as visible text. A small limit can end a response before it contains visible text.

Streaming responses ​

Set stream: true to receive server-sent events (SSE) with content type text/event-stream. With curl, use -N to display events without buffering:

bash
curl -N -X POST "https://inference.remotegpu.ai/v1/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $REMOTEGPU_API_KEY" \
  -d '{
    "model": "gpt-6-luna",
    "messages": [
      {
        "role": "user",
        "content": "Write a one-sentence greeting."
      }
    ],
    "max_completion_tokens": 256,
    "stream": true,
    "stream_options": {
      "include_usage": true
    }
  }'

Each data: event contains a JSON chunk. Append text from choices[].delta.content when it is present. The final usage chunk can have an empty choices array, so handle usage separately from generated text. The stream ends with data: [DONE].

Token usage is needed to finalize billing, and the API requests it for streamed responses. You can send stream_options: {"include_usage": true} to receive the final usage chunk in your client. If you omit it or set include_usage to false, the API uses that chunk for billing but does not forward it to your client.

Keep your connection open until [DONE]. If the connection closes early, you might not receive complete usage. After response headers have been sent, a stream failure cannot change the HTTP status; an initial 200 alone does not mean the generation completed.

Timeouts and retries ​

Set client timeouts long enough for the output you request. A request that times out before response headers are sent can return 504. A streaming request can end after its headers have been sent, so also check whether the stream completed.

RemoteGPU does not automatically resubmit a request after sending it to the model. A client retry creates another request and can incur another charge.

Common status codes ​

Status codeDescription
400Unsupported field or value for the selected model or endpoint
401Missing, invalid, revoked, or expired API key
402Insufficient available wallet credit for the reservation
403API key is valid but not authorized for inference APIs
404The selected model does not exist or is not available
422Request validation failed, such as a missing model field
429Model capacity or a rate limit was reached; retry later
502The model provider returned an error or an invalid response
503The model or its billing configuration cannot serve the request
504The request timed out before a response was available

Errors include an OpenAI-style error object. Check error.message, error.code, and error.param when present to identify what to change.

Billing ​

Wallet reservations ​

Before sending a request to the model, RemoteGPU reserves wallet credit for the model's maximum input allowance and your output limit. The input allowance uses the highest applicable input or cache-write rate. For models with long-context pricing, the reservation also covers the applicable long-context rates.

Send max_completion_tokens to limit the output portion of the reservation. For models that accept max_tokens, you can use that field instead, but do not send both. If you omit the output limit, the reservation uses the model's maximum output allowance. With n, it covers that limit for each completion. The input portion still covers the model's maximum input allowance, so even a short prompt can require a substantial reservation.

This reservation reduces available wallet credit while the request is in progress. When complete usage is available, RemoteGPU charges for actual usage and releases unused reserved credit. The reservation is not the final charge, and RemoteGPU does not insert an output limit into your request when you omit one. If usage is incomplete, including an interrupted stream, the reservation can remain held until billing is reconciled. Contact support if a hold remains after a failed request.

Token and cache charges ​

See Text pricing for customer rates in USD per 1 million tokens. Input, cached input, cache writes, and output can have different rates. Use only the categories that apply to the selected model.

usage.prompt_tokens includes cached input and cache writes. Those quantities are charged in their respective categories rather than charged again as ordinary input. usage.completion_tokens is the output quantity used for billing, including reasoning tokens when the model reports them as completion tokens.

Usage fieldMeaning
prompt_tokensTotal input tokens, including cache reads and writes
completion_tokensTotal completion tokens
total_tokensAggregate token count, not a separate charge
prompt_tokens_details.cached_tokensInput tokens read from cache
prompt_tokens_details.cache_write_tokensInput tokens written to cache, including GPT-6 cache writes
claude_cache_creation_5_m_tokensClaude cache writes with a five-minute lifetime
claude_cache_creation_1_h_tokensClaude cache writes with a one-hour lifetime

Cache fields depend on the model. Claude's five-minute and one-hour writes are priced separately. Cache use comes from reported usage; the Text API does not expose cache-management controls for these models.

For GPT-6, requests with more than 272,000 input tokens use long-context rates for the entire request, including input, cached input, cache writes, and output. At or below that threshold, standard-context rates apply. The total input count determines the tier, before subtracting cache reads or writes.

Rates are refreshed from official price sources. Each accepted request uses the pricing version selected for that request. If valid pricing is unavailable, the API rejects the request instead of using a hard-coded price.

Reference ​

Text models ​

This table lists models published in the public catalog. A missing message count limit means the catalog does not publish a count cap; token and request size limits still apply. The maximum output is a limit, not an automatically inserted default.

ModelMax messagesMax output tokensStreaming
claude-opus-4-81,000128,000Yes
claude-sonnet-4-61,000128,000Yes
gemini-3.5-flash1,00065,536Yes
gpt-51,000128,000Yes
gpt-5-mini1,000128,000Yes
gpt-5.11,000128,000Yes
gpt-6-astraNot published128,000Yes
gpt-6-lunaNot published128,000Yes
gpt-6-solNot published128,000Yes

Model parameters ​

Every request requires model and messages. Expand a model below for its accepted fields and limits. Numeric limits are inclusive. For messages and stop, the limits count list items. A missing default means the catalog does not specify one; it does not mean the field is required.

The catalog type chat_messages means an array of message objects. string_list accepts a string or an array of strings for stop.

FieldPurpose
max_completion_tokens or max_tokensLimit completion tokens; send only one of these fields
temperatureControl sampling randomness within the model's accepted range
top_pRestrict sampling to tokens within a cumulative probability
nRequest this many completions, each with its own output allowance
stopStop generation at a matching string
frequency_penaltyAdjust sampling according to how often a token has appeared
presence_penaltyAdjust sampling according to whether a token has appeared

This list explains field meanings. It does not mean every model accepts all of them; check the model's table before adding optional fields.

stream defaults to false at the API endpoint. stream_options.include_usage controls whether the final usage chunk is returned to your streaming client. These endpoint behaviors apply in addition to the model-specific fields below.

claude-opus-4-8
FieldTypeRequiredLimitsCatalog default
max_completion_tokensintegerNo1 to 128,000Not specified
max_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYes1 to 1,000 itemsNot specified
nintegerNo1 to 1Not specified
stopstring_listNoAt most 4 itemsNot specified
streambooleanNoNot publishedNot specified
claude-sonnet-4-6
FieldTypeRequiredLimitsCatalog default
max_completion_tokensintegerNo1 to 128,000Not specified
max_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYes1 to 1,000 itemsNot specified
nintegerNo1 to 1Not specified
stopstring_listNoAt most 4 itemsNot specified
streambooleanNoNot publishedNot specified
temperaturenumberNo0 to 1Not specified
gemini-3.5-flash
FieldTypeRequiredLimitsCatalog default
max_completion_tokensintegerNo1 to 65,536Not specified
max_tokensintegerNo1 to 65,536Not specified
messageschat_messagesYes1 to 1,000 itemsNot specified
nintegerNo1 to 1Not specified
streambooleanNoNot publishedNot specified
temperaturenumberNo0 to 2Not specified
top_pnumberNo0 to 1Not specified
gpt-5
FieldTypeRequiredLimitsCatalog default
max_completion_tokensintegerNo1 to 128,000Not specified
max_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYes1 to 1,000 itemsNot specified
nintegerNo1 to 4Not specified
streambooleanNoNot publishedNot specified
temperaturenumberNo0 to 2Not specified
top_pnumberNo0 to 1Not specified
gpt-5-mini
FieldTypeRequiredLimitsCatalog default
frequency_penaltynumberNo-2 to 2Not specified
max_completion_tokensintegerNo1 to 128,000Not specified
max_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYes1 to 1,000 itemsNot specified
nintegerNo1 to 4Not specified
presence_penaltynumberNo-2 to 2Not specified
stopstring_listNoAt most 4 itemsNot specified
streambooleanNoNot publishedNot specified
temperaturenumberNo0 to 2Not specified
top_pnumberNo0 to 1Not specified
gpt-5.1
FieldTypeRequiredLimitsCatalog default
frequency_penaltynumberNo-2 to 2Not specified
max_completion_tokensintegerNo1 to 128,000Not specified
max_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYes1 to 1,000 itemsNot specified
nintegerNo1 to 4Not specified
presence_penaltynumberNo-2 to 2Not specified
streambooleanNoNot publishedNot specified
temperaturenumberNo0 to 2Not specified
top_pnumberNo0 to 1Not specified
gpt-6-astra
FieldTypeRequiredLimitsCatalog default
max_completion_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYesAt least 1 item; no maximum publishedNot specified
streambooleanNoNot publishedNot specified
gpt-6-luna
FieldTypeRequiredLimitsCatalog default
max_completion_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYesAt least 1 item; no maximum publishedNot specified
streambooleanNoNot publishedNot specified
gpt-6-sol
FieldTypeRequiredLimitsCatalog default
max_completion_tokensintegerNo1 to 128,000Not specified
messageschat_messagesYesAt least 1 item; no maximum publishedNot specified
streambooleanNoNot publishedNot specified

Model catalog endpoint ​

Read model parameters and serving status without an API key:

bash
curl "https://inference.remotegpu.ai/v1/inference/models"

The response includes both Text and Image models. For Text, read:

  • text[].model: the identifier to send in POST /v1/chat/completions.
  • text[].parameters: accepted fields, policies, defaults, and limits.
  • text[].runtime.state: whether the model is ready, starting, or sleeping.
  • text[].recent_summary_stats: recent request wait, generation, and total times.

Use this catalog to populate model selectors and validate optional fields. Use the authenticated GET /v1/models endpoint when an OpenAI-compatible client only needs model IDs.

RemoteGPU customer documentation