This API is OpenAI-compatible, but key differences exist in parameters, features, and behaviors. Requests process only the parameters listed on this page. Any other OpenAI parameters are accepted and ignored.
Key differences:
Unsupported parameters: previous_response_id, conversation, background, store: true, and non-text content parts (input_image, input_file) are rejected — this API is stateless and text-only. See the conversation guide.
from openai import OpenAIclient = OpenAI( api_key="UPSTAGE_API_KEY", base_url="https://api.upstage.ai/v1",)response = client.responses.create( model="solar-pro4", input="Hi, how are you?",)print(response.output_text)
Set stream to true to receive the response as Server-Sent Events. Each event carries its type on the event: line and a JSON payload on the data: line; every payload repeats the type and carries a monotonically increasing sequence_number. There is no [DONE] sentinel — the stream ends with response.completed, response.incomplete, or response.failed.
Errors use the OpenAI error envelope. param identifies the field associated with the failure, and code classifies it; either may be null. Some validation errors use the corresponding Chat Completions field name, such as max_tokens or response_format.
{ "error": { "message": "the 'previous_response_id' parameter is not supported: this API is stateless and does not store responses. Send the full conversation history in 'input' instead", "type": "invalid_request_error", "param": "previous_response_id", "code": "invalid_request_body" }}
400 A parameter is invalid, the context length is exceeded, or a stateless-incompatible parameter (previous_response_id, conversation, background, store: true) was sent. Depending on the error, code may be invalid_request_body, context_length_exceeded, another error code, or null.
401 The API key is missing or invalid.
404model_not_found The model does not exist or does not support the Responses API.
408 The request timed out.
429 Rate limit exceeded.
5xx An upstream or internal error. Retry with backoff.
The model alias or snapshot to use. The Responses API is available on Solar Pro 4 and Solar Mini 4.
Value in: "solar-pro4" | "solar-pro4-260806" | "solar-mini4" | "solar-mini4-260922"
input
Required
Text | Item list
The conversation to send to the model. A string is a single user message; an array carries the full conversation as typed items — this API is stateless, so replay the previous response's output items before the new message (see the conversation guide). Non-text content parts are rejected.
instructionsstring
A system prompt inserted ahead of the conversation. Equivalent to a leading system message, and not carried over between requests.
textobject
Output-format configuration. format selects plain text or structured output; the applied value is echoed back on the response so you can confirm it was honored.
{"type": "text"} — the default, free-form text.
{"type": "json_object"} — JSON mode. The conversation must mention JSON, otherwise the request is rejected. Guarantees valid JSON but not a particular shape.
{"type": "json_schema", "name": ..., "schema": ...} — structured outputs, constrained to the JSON Schema you supply. For compatibility with OpenAI strict schemas, we recommend setting additionalProperties to false on every object and listing all properties in required; make a value optional by adding null to its type. Learn more in the Structured outputs guide.
Unlike Chat Completions, the schema fields sit flat next to type rather than nested under a json_schema key. verbosity is accepted and ignored.
formatPlain text | JSON mode | Structured outputs
type
Required
string
Value: "text"
toolsarray<object>
A list of tools the model may call. Only function tools are supported.
Unlike Chat Completions, a tool is declared flat — name, description, and parameters sit next to type rather than nested under a function key. Learn more in the Tool calling guide.
tool_choicestring | object
Controls which (if any) tool is called by the model.
none means the model will not call any tool and instead generates a message.
auto means the model can pick between generating a message or calling one or more tools.
required means the model must call one or more tools.
{"type": "function", "name": "my_function"} forces the model to call that tool. Note the flat shape — Chat Completions nests the name under a function key.
none is the default when no tools are present, auto when tools are present. Object forms other than function (allowed_tools, mcp, custom, …) are rejected.
parallel_tool_callsboolean
Whether to allow parallel function calling during tool use. When enabled, the model may generate multiple tool calls in a single response, allowing independent calls to run concurrently. This setting does not guarantee more than one tool call; process all function_call items in output.
Default: true
reasoningobject
Controls how much reasoning the model does before answering. Solar Pro 4 and Solar Mini 4 accept every effort level; none and minimal turn reasoning off.
When reasoning runs, the visible reasoning text is returned as a reasoning item at the head of output, and reasoning tokens are counted in usage.output_tokens_details.reasoning_tokens. Note that max_output_tokens covers reasoning tokens as well, so a small budget can be consumed entirely by reasoning. Learn more in the Reasoning guide.
summary is accepted and ignored — Solar returns raw reasoning text rather than a summary.
max_output_tokensinteger
An upper bound on the tokens generated for this response, including reasoning tokens. The maximum is 131,072 tokens. The sum of input tokens and this value must not exceed the model's context length. When the limit is reached the response comes back with statusincomplete.
Maximum: 131072
temperaturenumber
An optional parameter to set the sampling temperature. The value should lie between 0 and 2. Higher values like 0.8 result in a more random output, whereas lower values such as 0.2 enhance focus and determinism in the output.
Defaults to 1.0.
Default: 1Minimum: 0Maximum: 2Format: "float"
top_pnumber
An optional parameter to trigger nucleus sampling. The tokens with top_p probability mass will be considered, which means, setting this value to 0.1 will consider tokens comprising the top 10% probability.
Default: 1Minimum: 0Maximum: 1Format: "float"
streamboolean
Whether to stream the response as named Server-Sent Events, with an event: name and a JSON data: payload. The stream ends with response.completed, response.incomplete, or response.failed; no [DONE] marker is sent.
Default: false
prompt_cache_keystring
An optional parameter that specifies a unique key for identifying and caching the prompt. Use a distinct key for each conversational context to improve cache utilization.
Default: null
frequency_penaltynumber
Reduces the chance of repeating the same token in proportion to how often it has already appeared. Between -2.0 and 2.0.
This is not a standard OpenAI Responses parameter. The Python SDK passes it using extra_body={"frequency_penalty": 0.8}; the Node.js SDK and curl send it as a top-level parameter. presence_penalty is accepted but has no effect.