File Deleted
View Diff
1#### Inference API
2
3# Chat
4
5## POST /v1/chat/completions
6
7Create a chat response from text/image chat prompts. This is the endpoint for making requests to chat and image understanding models.
8
9### Request Body
10
11* `deferred` (boolean | null) — If set to \`true\`, the request returns a \`request\_id\`. You can then get the deferred response by GET \`/v1/chat/deferred-completion/\{request\_id}\`.
12
13* `frequency_penalty` (number | null) — (Not supported by reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
14
15* `logit_bias` (object | null) — (Unsupported) A JSON object that maps tokens (specified by their token ID in the tokenizer) to an associated bias value from -100 to 100. Mathematically, the bias is added to the logits generated by the model prior to sampling. The exact effect will vary per model, but values between -1 and 1 should decrease or increase likelihood of selection; values like -100 or 100 should result in a ban or exclusive selection of the relevant token.
16
17* `logprobs` (boolean | null) — Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.
18
19* `max_completion_tokens` (integer | null) — An upper bound for the number of tokens that can be generated for a completion, only applies to visible output tokens (i.e. does not apply to tokens used for reasoning or function calls). Defaults to 128,000 when unset; set a larger value to allow longer generations.
20
21* `max_tokens` (integer | null) — \\\[DEPRECATED\\] The maximum number of tokens that can be generated in the chat completion. Deprecated in favor of \`max\_completion\_tokens\`.
22
23* `messages` (array\<object | object | object | object | object>) — A list of messages that make up the chat conversation. Different models support different message types, such as image and text.
24
25* `model` (string) — Model name for the model to use. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.
26
27* `n` (integer | null) — How many chat completion choices to generate for each input message. Note that you will be charged based on the number of generated tokens across all of the choices. Keep n as 1 to minimize costs.
28
29* `parallel_tool_calls` (boolean | null) — If set to false, the model can perform maximum one tool call.
30
31* `presence_penalty` (number | null) — (Not supported by \`grok-3\` and reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
32
33* `prompt_cache_key` (string | null) — A stable cache key for best-effort sticky routing / prompt-cache hits
34 across requests sharing a prompt prefix. Plumbed to \`x-grok-conv-id\`,
35 same as on \`/v1/responses\`.
36
37* `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).
38
39* `response_format` (object | object | object)
40
41* `search_parameters` (object)
42
43 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See
44 \<https://en.wikipedia.org/wiki/ISO\_8601>.
45
46 * `max_search_results` (integer | null) — Maximum number of search results to use.
47
48 * `mode` (string | null) — Choose the mode to query realtime data:
49 \* \`off\`: no search performed and no external will be considered.
50 \* \`on\` (default): the model will search in every sources for relevant data.
51 \* \`auto\`: the model choose whether to search data or not and where to search the data.
52
53 * `return_citations` (boolean | null) — Whether to return citations in the response or not.
54
55 * `sources` (array | null) — List of sources to search in. If no sources specified, the model will look over the web and X by default.
56
57 * `to_date` (string | null) — Date up to which to consider the results in ISO-8601 YYYY-MM-DD. See
58 \<https://en.wikipedia.org/wiki/ISO\_8601>.
59
60* `seed` (integer | null) — If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same \`seed\` and parameters should return the same result. Determinism is not guaranteed, and you should refer to the \`system\_fingerprint\` response parameter to monitor changes in the backend.
61
62* `service_tier` ("default" | "priority")
63
64* `stop` (array | null) — (Not supported by reasoning models) Up to 4 sequences where the API will stop generating further tokens.
65
66* `stream` (boolean | null) — If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a \`data: \[DONE]\` message.
67
68* `stream_options` (object)
69
70 * `include_usage` (boolean, required) — Set an additional chunk to be streamed before the \`data: \[DONE]\` message. The other chunks will return \`null\` in \`usage\` field.
71
72* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
73
74* `tool_choice` (string | object)
75
76* `tools` (array | null) — A list of tools the model may call in JSON-schema. Currently, only functions are supported as a tool. Use this to provide a list of functions the model may generate JSON inputs for. A max of 128 functions are supported.
77
78* `top_logprobs` (integer | null) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position, each with an associated log probability. logprobs must be set to true if this parameter is used. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.
79
80* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.
81
82* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.
83
84* `web_search_options` (object)
85
86 * `filters` (object) — Only included for compatibility.
87
88 * `search_context_size` (string | null) — This field included for compatibility reason with OpenAI's API. It is mapped to \`max\_search\`.
89
90 * `user_location` (object) — Only included for compatibility.
91
92### Response Body
93
94* `choices` (array\<object>, required) — A list of response choices from the model. The length corresponds to the \`n\` in request body (default to 1).
95
96 * `finish_reason` (string | null) — Finish reason. \`"stop"\` means the inference has reached a model-defined or user-supplied stop sequence in \`stop\`. \`"length"\` means the inference result has reached models' maximum allowed token length or user defined value in \`max\_tokens\`. \`"end\_turn"\` or \`null\` in streaming mode when the chunk is not the last.
97
98 * `index` (integer, required) — Index of the choice within the response choices, starting from 0.
99
100 * `logprobs` (object)
101
102 * `content` (array | null) — An array the log probabilities of each output token returned.
103
104 * `message` (object, required)
105
106 * `content` (string | null) — The content of the message.
107
108 * `reasoning_content` (string | null) — The reasoning trace generated by the model.
109
110 * `refusal` (string | null) — The reason given by model if the model is unable to generate a response. null if model is able to generate.
111
112 * `role` (string, required) — The role that the message belongs to, the response from model is always \`"assistant"\`.
113
114 * `tool_calls` (array | null) — A list of tool calls asked by model for user to perform.
115
116* `citations` (array | null) — List of all the external pages used by the model to answer.
117
118* `created` (integer, required) — The chat completion creation time in Unix timestamp.
119
120* `id` (string, required) — A unique ID for the chat response.
121
122* `model` (string, required) — Model ID used to create chat completion.
123
124* `object` (string, required) — The object type, which is always \`"chat.completion"\`.
125
126* `output_files` (array | null) — Files generated during the response (e.g., by the code execution tool).
127 Only populated when \`code\_execution\_files\_output\` is included.
128
129* `service_tier` ("default" | "priority", required) — Processing tier for a request. Determines scheduling priority and billing.
130
131* `system_fingerprint` (string | null) — System fingerprint, used to indicate xAI system configuration changes.
132
133* `usage` (object)
134
135 * `completion_tokens` (integer, required) — Total completion token used.
136
137 * `completion_tokens_details` (object, required) — Details of completion usage.
138
139 * `accepted_prediction_tokens` (integer, required) — The number of tokens in the prediction that appeared in the completion.
140
141 * `audio_tokens` (integer, required) — Audio input tokens generated by the model.
142
143 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.
144
145 * `rejected_prediction_tokens` (integer, required) — The number of tokens in the prediction that did not appear in the completion.
146
147 * `cost_in_usd_ticks` (integer, required) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:
148 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000
149 which means there is 10'000'000'000 ticks in one \*dollar\*.
150
151 * `num_sources_used` (integer, required) — Number of individual live search source used.
152
153 * `prompt_tokens` (integer, required) — Total prompt token used.
154
155 * `prompt_tokens_details` (object, required) — Details of prompt usage.
156
157 * `audio_tokens` (integer, required) — Audio prompt token used.
158
159 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.
160
161 * `image_tokens` (integer, required) — Image prompt token used.
162
163 * `text_tokens` (integer, required) — Total text prompt token used (cached + non-cached text tokens).
164
165 * `total_tokens` (integer, required) — Total token used, the sum of prompt token and completion token amount.
166
167\*\*Request example:\*\*
168
169```json
170{
171 "messages": [
172 {
173 "role": "system",
174 "content": "You are a helpful assistant that can answer questions and help with tasks."
175 },
176 {
177 "role": "user",
178 "content": "What is 101*3?"
179 }
180 ],
181 "model": "latest"
182}
183```
184
185\*\*Response example:\*\*
186
187```json
188{
189 "id": "a3d1008e-4544-40d4-d075-11527e794e4a",
190 "object": "chat.completion",
191 "created": 1752854522,
192 "model": "latest",
193 "choices": [
194 {
195 "index": 0,
196 "message": {
197 "role": "assistant",
198 "content": "101 multiplied by 3 is 303.",
199 "refusal": null
200 },
201 "finish_reason": "stop"
202 }
203 ],
204 "usage": {
205 "prompt_tokens": 32,
206 "completion_tokens": 9,
207 "total_tokens": 135,
208 "prompt_tokens_details": {
209 "text_tokens": 32,
210 "audio_tokens": 0,
211 "image_tokens": 0,
212 "cached_tokens": 6
213 },
214 "completion_tokens_details": {
215 "reasoning_tokens": 94,
216 "audio_tokens": 0,
217 "accepted_prediction_tokens": 0,
218 "rejected_prediction_tokens": 0
219 },
220 "num_sources_used": 0
221 },
222 "system_fingerprint": "fp_3a7881249c"
223}
224```
225
226***
227
228## POST /v1/responses
229
230Generates a response based on text or image prompts. The response ID can be used to retrieve the response later or to continue the conversation without repeating prior context. New responses will be stored for 30 days and then permanently deleted.
231
232### Request Body
233
234* `background` (boolean | null) — (Unsupported) Whether to process the response asynchronously in the background.
235
236* `context_management` (array | null) — Optional context-management directives (e.g. compaction). Parsed but not yet executed.
237
238* `include` (array | null) — What additional output data to include in the response. Supported values include
239 \`reasoning.encrypted\_content\` (encrypted reasoning tokens) and tool-output options.
240 OpenAI's \`message.output\_text.logprobs\` is accepted for compatibility but silently ignored.
241
242* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.
243
244* `instructions` (string | null) — An alternate way to specify the system prompt. Note that this cannot be used alongside \`previous\_response\_id\`, where the system prompt of the previous message will be used.
245
246* `logprobs` (boolean | null) — Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.
247
248* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens. Defaults to 128,000 when unset; set a larger value to allow longer generations.
249
250* `max_turns` (integer | null) — Maximum number of agentic tool calling turns allowed for this request.
251 If not set, defaults to the server's global cap.
252 This parameter will be ignored for any non-agentic requests.
253
254* `metadata` (object) — Not supported. Only maintained for compatibility reasons.
255
256* `min_p` (number | null) — Min-p sampling: tokens whose probability is below \`min\_p\` times the probability of the most likely token are excluded from sampling. Disabled when unset.
257
258* `model` (string) — Model name for the model to use. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.
259
260* `parallel_tool_calls` (boolean | null) — Whether to allow the model to run parallel tool calls.
261
262* `previous_response_id` (string | null) — The ID of the previous response from the model.
263
264* `prompt_cache_key` (string | null) — Plumbed to x-grok-conv-id for Open Responses compatibility, used for routing.
265
266* `reasoning` (object)
267
268 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).
269
270 * `generate_summary` (string | null) — Only included for compatibility.
271
272 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.
273
274* `reasoning_effort` (string | null) — reasoning\_effort alternative to reasoning configuration. This is a non-standard field meant to ease user experience. We only look at this if the reasoning field is unset.
275
276* `search_parameters` (object)
277
278 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See
279 \<https://en.wikipedia.org/wiki/ISO\_8601>.
280
281 * `max_search_results` (integer | null) — Maximum number of search results to use.
282
283 * `mode` (string | null) — Choose the mode to query realtime data:
284 \* \`off\`: no search performed and no external will be considered.
285 \* \`on\` (default): the model will search in every sources for relevant data.
286 \* \`auto\`: the model choose whether to search data or not and where to search the data.
287
288 * `return_citations` (boolean | null) — Whether to return citations in the response or not.
289
290 * `sources` (array | null) — List of sources to search in. If no sources specified, the model will look over the web and X by default.
291
292 * `to_date` (string | null) — Date up to which to consider the results in ISO-8601 YYYY-MM-DD. See
293 \<https://en.wikipedia.org/wiki/ISO\_8601>.
294
295* `service_tier` ("default" | "priority")
296
297* `store` (boolean | null) — Whether to store the input message(s) and model response for later retrieval.
298
299* `stream` (boolean | null) — If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a \`data: \[DONE]\` message.
300
301* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
302
303* `text` (object)
304
305 * `format` (object | object | object)
306
307* `tool_choice` (string | object)
308
309* `tools` (array | null) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.\`web\_search\_preview\` tool, if specified, will be overridden by \`search\_parameters\`.
310
311* `top_k` (integer | null) — Top-k sampling: only the \`top\_k\` most probable tokens are considered at each sampling step. Disabled when unset.
312
313* `top_logprobs` (integer | null) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position, each with an associated log probability. logprobs must be set to true if this parameter is used. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.
314
315* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.
316
317* `truncation` (string | null) — Not supported. Only maintained for compatibility reasons.
318
319* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.
320
321### Response Body
322
323* `background` (boolean, required) — OpenResponses compatibility fields.
324 Not used at the moment. Just for OpenResponses compatibility.
325 Whether to process the response asynchronously in the background.
326
327* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.
328
329* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.
330
331* `error` (object) — An error object returned when the model fails to generate a response.
332
333* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
334
335* `id` (string, required) — Unique ID of the response.
336
337* `incomplete_details` (object | object | object)
338
339* `instructions` (string | null) — A system (or developer) message inserted into the model's context.
340
341* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.
342
343* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.
344
345* `metadata` (object, required) — Only included for compatibility.
346
347* `model` (string, required) — Model name used to generate the response.
348
349* `object` (string, required) — The object type of this resource. Always set to \`response\`.
350
351* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.
352
353* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.
354
355* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
356
357* `previous_response_id` (string | null) — The ID of the previous response from the model.
358
359* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.
360
361* `reasoning` (object)
362
363 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).
364
365 * `generate_summary` (string | null) — Only included for compatibility.
366
367 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.
368
369* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.
370
371* `service_tier` ("default" | "priority", required)
372
373* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.
374
375* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.
376
377* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
378
379* `text` (object, required)
380
381 * `format` (object | object | object)
382
383* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.
384
385 * `name` (string, required) — Name of the function to use.
386
387 * `type` (string, required) — Type is always \`"function"\`.
388
389* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.
390
391* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.
392
393* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.
394
395* `truncation` (string, required) — The truncation strategy to use for the model response.
396
397* `usage` (object)
398
399 * `context_details` (object)
400
401 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from
402 \`SamplingUsage.context\_prompt\_tokens\`).
403
404 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from
405 \`SamplingUsage.context\_output\_tokens\`).
406
407 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.
408
409 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:
410 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000
411 which means there is 10'000'000'000 ticks in one \*dollar\*.
412
413 * `input_tokens` (integer, required) — Number of input tokens used.
414
415 * `input_tokens_details` (object, required)
416
417 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.
418
419 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.
420
421 * `num_sources_used` (integer, required) — Number of sources used (for live search).
422
423 * `output_tokens` (integer, required) — Number of output tokens used.
424
425 * `output_tokens_details` (object, required)
426
427 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.
428
429 * `server_side_tool_usage_details` (object)
430
431 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.
432
433 * `document_search_calls` (integer, required) — Number of document search calls.
434
435 * `file_search_calls` (integer, required) — Number of file search calls.
436
437 * `image_generation_calls` (integer, required) — Number of image generation calls.
438
439 * `mcp_calls` (integer, required) — Number of MCP calls.
440
441 * `web_search_calls` (integer, required) — Number of web search calls.
442
443 * `x_search_calls` (integer, required) — Number of X search calls.
444
445 * `total_tokens` (integer, required) — Total tokens used.
446
447* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.
448
449### Code Examples
450
451```bash
452curl -s https://api.x.ai/v1/responses \
453 -H "Content-Type: application/json" \
454 -H "Authorization: Bearer $XAI_API_KEY" \
455 -d '{
456 "model": "grok-4.6",
457 "input": "What is the meaning of life?"
458 }'
459```
460
461```javascriptAISDK
462import { xai } from "@ai-sdk/xai";
463import { generateText } from "ai";
464
465const result = await generateText({
466 model: xai.responses("grok-4.6"),
467 prompt: "What is the meaning of life?",
468});
469
470console.log(JSON.stringify(result, null, 2));
471```
472
473```pythonOpenAISDK
474import os
475
476from openai import OpenAI
477
478client = OpenAI(
479 api_key=os.environ["XAI_API_KEY"],
480 base_url="https://api.x.ai/v1",
481)
482
483response = client.responses.create(
484 model="grok-4.6",
485 input="What is the meaning of life?",
486)
487
488print(response.model_dump_json(indent=2))
489```
490
491```javascriptOpenAISDK
492import OpenAI from "openai";
493
494const client = new OpenAI({
495 apiKey: process.env.XAI_API_KEY,
496 baseURL: "https://api.x.ai/v1",
497});
498
499const response = await client.responses.create({
500 model: "grok-4.6",
501 input: "What is the meaning of life?",
502});
503
504console.log(JSON.stringify(response, null, 2));
505```
506
507\*\*Response example:\*\*
508
509```json
510{
511 "created_at": 1754475266,
512 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",
513 "max_output_tokens": null,
514 "model": "latest",
515 "object": "response",
516 "output": [
517 {
518 "content": [
519 {
520 "type": "output_text",
521 "text": "101 multiplied by 3 is 303.",
522 "logprobs": null,
523 "annotations": []
524 }
525 ],
526 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",
527 "role": "assistant",
528 "type": "message",
529 "status": "completed"
530 }
531 ],
532 "parallel_tool_calls": true,
533 "previous_response_id": null,
534 "reasoning": null,
535 "temperature": null,
536 "text": {
537 "format": {
538 "type": "text"
539 }
540 },
541 "tool_choice": "auto",
542 "tools": [],
543 "top_p": null,
544 "usage": {
545 "input_tokens": 32,
546 "input_tokens_details": {
547 "cached_tokens": 8
548 },
549 "output_tokens": 9,
550 "output_tokens_details": {
551 "reasoning_tokens": 110
552 },
553 "total_tokens": 151,
554 "num_sources_used": 0,
555 "num_server_side_tools_used": 0
556 },
557 "user": null,
558 "incomplete_details": null,
559 "status": "completed",
560 "store": true
561}
562```
563
564***
565
566## POST /v1/responses/compact
567
568Compacts a full Responses API input window into a shorter canonical window.
569
570### Request Body
571
572* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.
573
574* `model` (string, required) — Model to use for compaction summarization (required).
575
576### Response Body
577
578* `created_at` (integer, required) — Unix timestamp (in seconds) when the compacted conversation was created.
579
580* `id` (string, required) — Unique ID for this compaction (e.g. \`cmp\_\<uuid>\`).
581
582* `model` (string, required) — Model used for the compaction summary.
583
584* `object` (string, required) — Always \`"response.compaction"\`.
585
586* `output` (array\<object | object | object | object | object>, required) — Compacted output containing a single compaction item.
587 Pass this verbatim as input to the next \`/v1/responses\` call.
588
589* `usage` (object)
590
591 * `dropped_message_count` (integer, required) — Number of messages dropped/summarized during compaction.
592
593 * `input_tokens` (integer, required) — Number of tokens in the original (pre-compaction) input.
594
595 * `input_tokens_details` (object, required) — Breakdown of input tokens for a compaction call.
596
597 * `cached_tokens` (integer, required) — Number of input tokens that were served from the prompt cache.
598
599 * `output_tokens` (integer, required) — Number of tokens in the compacted output.
600
601 * `output_tokens_details` (object, required) — Breakdown of output tokens for a compaction call.
602
603 * `reasoning_tokens` (integer, required) — Number of reasoning tokens generated during compaction.
604
605 * `total_tokens` (integer, required) — Total number of tokens used (input + output, including reasoning).
606
607### Code Examples
608
609```bash
610curl -s https://api.x.ai/v1/responses/compact \
611 -H "Content-Type: application/json" \
612 -H "Authorization: Bearer $XAI_API_KEY" \
613 -d '{
614 "model": "grok-4.6",
615 "input": [
616 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},
617 {"role": "user", "content": "What is the Higgs boson and why is it important?"},
618 {"role": "assistant", "content": "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."},
619 {"role": "user", "content": "How does the Higgs mechanism actually work?"},
620 {"role": "assistant", "content": "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless."}
621 ]
622 }'
623```
624
625```pythonOpenAISDK
626import os
627
628from openai import OpenAI
629
630client = OpenAI(
631 api_key=os.environ["XAI_API_KEY"],
632 base_url="https://api.x.ai/v1",
633)
634
635compacted = client.responses.compact(
636 model="grok-4.6",
637 input=[
638 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},
639 {"role": "user", "content": "What is the Higgs boson and why is it important?"},
640 {
641 "role": "assistant",
642 "content": (
643 "The Higgs boson is an elementary particle in the Standard Model, predicted by "
644 "Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation "
645 "of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."
646 ),
647 },
648 {"role": "user", "content": "How does the Higgs mechanism actually work?"},
649 {
650 "role": "assistant",
651 "content": (
652 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, "
653 "and particles acquire mass in proportion to how strongly they couple to it. Photons "
654 "do not couple, which is why they remain massless."
655 ),
656 },
657 ],
658)
659
660print(compacted.model_dump_json(indent=2))
661```
662
663```javascriptOpenAISDK
664import OpenAI from "openai";
665
666const client = new OpenAI({
667 apiKey: process.env.XAI_API_KEY,
668 baseURL: "https://api.x.ai/v1",
669});
670
671const compacted = await client.responses.compact({
672 model: "grok-4.6",
673 input: [
674 { role: "system", content: "You are a concise and knowledgeable science tutor." },
675 { role: "user", content: "What is the Higgs boson and why is it important?" },
676 {
677 role: "assistant",
678 content:
679 "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism.",
680 },
681 { role: "user", content: "How does the Higgs mechanism actually work?" },
682 {
683 role: "assistant",
684 content:
685 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless.",
686 },
687 ],
688});
689
690console.log(JSON.stringify(compacted, null, 2));
691```
692
693\*\*Response example:\*\*
694
695```json
696{}
697```
698
699***
700
701## GET /v1/responses/\{response\_id}
702
703Retrieve a previously generated response.
704
705### Path Parameters
706
707* `response_id` (string, required) — The response id returned by a previous create response request.
708
709### Response Body
710
711* `background` (boolean, required) — OpenResponses compatibility fields.
712 Not used at the moment. Just for OpenResponses compatibility.
713 Whether to process the response asynchronously in the background.
714
715* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.
716
717* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.
718
719* `error` (object) — An error object returned when the model fails to generate a response.
720
721* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
722
723* `id` (string, required) — Unique ID of the response.
724
725* `incomplete_details` (object | object | object)
726
727* `instructions` (string | null) — A system (or developer) message inserted into the model's context.
728
729* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.
730
731* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.
732
733* `metadata` (object, required) — Only included for compatibility.
734
735* `model` (string, required) — Model name used to generate the response.
736
737* `object` (string, required) — The object type of this resource. Always set to \`response\`.
738
739* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.
740
741* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.
742
743* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
744
745* `previous_response_id` (string | null) — The ID of the previous response from the model.
746
747* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.
748
749* `reasoning` (object)
750
751 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).
752
753 * `generate_summary` (string | null) — Only included for compatibility.
754
755 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.
756
757* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.
758
759* `service_tier` ("default" | "priority", required)
760
761* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.
762
763* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.
764
765* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
766
767* `text` (object, required)
768
769 * `format` (object | object | object)
770
771* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.
772
773 * `name` (string, required) — Name of the function to use.
774
775 * `type` (string, required) — Type is always \`"function"\`.
776
777* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.
778
779* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.
780
781* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.
782
783* `truncation` (string, required) — The truncation strategy to use for the model response.
784
785* `usage` (object)
786
787 * `context_details` (object)
788
789 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from
790 \`SamplingUsage.context\_prompt\_tokens\`).
791
792 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from
793 \`SamplingUsage.context\_output\_tokens\`).
794
795 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.
796
797 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:
798 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000
799 which means there is 10'000'000'000 ticks in one \*dollar\*.
800
801 * `input_tokens` (integer, required) — Number of input tokens used.
802
803 * `input_tokens_details` (object, required)
804
805 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.
806
807 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.
808
809 * `num_sources_used` (integer, required) — Number of sources used (for live search).
810
811 * `output_tokens` (integer, required) — Number of output tokens used.
812
813 * `output_tokens_details` (object, required)
814
815 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.
816
817 * `server_side_tool_usage_details` (object)
818
819 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.
820
821 * `document_search_calls` (integer, required) — Number of document search calls.
822
823 * `file_search_calls` (integer, required) — Number of file search calls.
824
825 * `image_generation_calls` (integer, required) — Number of image generation calls.
826
827 * `mcp_calls` (integer, required) — Number of MCP calls.
828
829 * `web_search_calls` (integer, required) — Number of web search calls.
830
831 * `x_search_calls` (integer, required) — Number of X search calls.
832
833 * `total_tokens` (integer, required) — Total tokens used.
834
835* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.
836
837\*\*Response example:\*\*
838
839```json
840{
841 "created_at": 1754475266,
842 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",
843 "max_output_tokens": null,
844 "model": "latest",
845 "object": "response",
846 "output": [
847 {
848 "content": [
849 {
850 "type": "output_text",
851 "text": "101 multiplied by 3 is 303.",
852 "logprobs": null,
853 "annotations": []
854 }
855 ],
856 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",
857 "role": "assistant",
858 "type": "message",
859 "status": "completed"
860 },
861 {
862 "id": "",
863 "summary": [
864 {
865 "text": "First, the user asked: \"What is 101*3?\"\n\nThis is a simple multiplication: 101 multiplied by 3.\n\nCalculating: 100 * 3 = 300, and 1 * 3 = 3, so 300 + 3 = 303.\n\nI should respond helpfully and directly, as per my system prompt: \"You are a helpful assistant that can answer questions and help with tasks.\"\n\nKeep the response concise and accurate. No need for extra fluff unless it adds value.\n\nFinal answer: 303.",
866 "type": "summary_text"
867 }
868 ],
869 "type": "reasoning",
870 "status": "completed"
871 }
872 ],
873 "parallel_tool_calls": true,
874 "previous_response_id": null,
875 "reasoning": null,
876 "temperature": null,
877 "text": {
878 "format": {
879 "type": "text"
880 }
881 },
882 "tool_choice": "auto",
883 "tools": [],
884 "top_p": null,
885 "usage": {
886 "prompt_tokens": 32,
887 "completion_tokens": 9,
888 "total_tokens": 151,
889 "prompt_tokens_details": {
890 "text_tokens": 32,
891 "audio_tokens": 0,
892 "image_tokens": 0,
893 "cached_tokens": 8
894 },
895 "completion_tokens_details": {
896 "reasoning_tokens": 110,
897 "audio_tokens": 0,
898 "accepted_prediction_tokens": 0,
899 "rejected_prediction_tokens": 0
900 },
901 "num_sources_used": 0
902 },
903 "user": null,
904 "incomplete_details": null,
905 "status": "completed",
906 "store": true
907}
908```
909
910***
911
912## DELETE /v1/responses/\{response\_id}
913
914Delete a previously generated response.
915
916### Path Parameters
917
918* `response_id` (string, required) — The response id returned by a previous create response request.
919
920### Response Body
921
922* `deleted` (boolean, required) — Whether the response was successfully deleted.
923
924* `id` (string, required) — The response\_id to be deleted.
925
926* `object` (string, required) — The deleted object type, which is always \`response\`.
927
928\*\*Response example:\*\*
929
930```json
931{
932 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",
933 "object": "response",
934 "deleted": true
935}
936```
937
938***
939
940## GET /v1/chat/deferred-completion/\{request\_id}
941
942Tries to fetch a result for a previously-started deferred completion. Returns \`200 Success\` with the response body, if the request has been completed. Returns \`202 Accepted\` when the request is pending processing.
943
944### Path Parameters
945
946* `request_id` (string, required) — The deferred request id returned by a previous deferred chat request.
947
948### Response Body
949
950* `choices` (array\<object>, required) — A list of response choices from the model. The length corresponds to the \`n\` in request body (default to 1).
951
952 * `finish_reason` (string | null) — Finish reason. \`"stop"\` means the inference has reached a model-defined or user-supplied stop sequence in \`stop\`. \`"length"\` means the inference result has reached models' maximum allowed token length or user defined value in \`max\_tokens\`. \`"end\_turn"\` or \`null\` in streaming mode when the chunk is not the last.
953
954 * `index` (integer, required) — Index of the choice within the response choices, starting from 0.
955
956 * `logprobs` (object)
957
958 * `content` (array | null) — An array the log probabilities of each output token returned.
959
960 * `message` (object, required)
961
962 * `content` (string | null) — The content of the message.
963
964 * `reasoning_content` (string | null) — The reasoning trace generated by the model.
965
966 * `refusal` (string | null) — The reason given by model if the model is unable to generate a response. null if model is able to generate.
967
968 * `role` (string, required) — The role that the message belongs to, the response from model is always \`"assistant"\`.
969
970 * `tool_calls` (array | null) — A list of tool calls asked by model for user to perform.
971
972* `citations` (array | null) — List of all the external pages used by the model to answer.
973
974* `created` (integer, required) — The chat completion creation time in Unix timestamp.
975
976* `id` (string, required) — A unique ID for the chat response.
977
978* `model` (string, required) — Model ID used to create chat completion.
979
980* `object` (string, required) — The object type, which is always \`"chat.completion"\`.
981
982* `output_files` (array | null) — Files generated during the response (e.g., by the code execution tool).
983 Only populated when \`code\_execution\_files\_output\` is included.
984
985* `service_tier` ("default" | "priority", required) — Processing tier for a request. Determines scheduling priority and billing.
986
987* `system_fingerprint` (string | null) — System fingerprint, used to indicate xAI system configuration changes.
988
989* `usage` (object)
990
991 * `completion_tokens` (integer, required) — Total completion token used.
992
993 * `completion_tokens_details` (object, required) — Details of completion usage.
994
995 * `accepted_prediction_tokens` (integer, required) — The number of tokens in the prediction that appeared in the completion.
996
997 * `audio_tokens` (integer, required) — Audio input tokens generated by the model.
998
999 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.
1000
1001 * `rejected_prediction_tokens` (integer, required) — The number of tokens in the prediction that did not appear in the completion.
1002
1003 * `cost_in_usd_ticks` (integer, required) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:
1004 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000
1005 which means there is 10'000'000'000 ticks in one \*dollar\*.
1006
1007 * `num_sources_used` (integer, required) — Number of individual live search source used.
1008
1009 * `prompt_tokens` (integer, required) — Total prompt token used.
1010
1011 * `prompt_tokens_details` (object, required) — Details of prompt usage.
1012
1013 * `audio_tokens` (integer, required) — Audio prompt token used.
1014
1015 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.
1016
1017 * `image_tokens` (integer, required) — Image prompt token used.
1018
1019 * `text_tokens` (integer, required) — Total text prompt token used (cached + non-cached text tokens).
1020
1021 * `total_tokens` (integer, required) — Total token used, the sum of prompt token and completion token amount.
1022
1023\*\*Response example:\*\*
1024
1025```json
1026{
1027 "id": "335b92e4-afa5-48e7-b99c-b9a4eabc1c8e",
1028 "object": "chat.completion",
1029 "created": 1743770624,
1030 "model": "latest",
1031 "choices": [
1032 {
1033 "index": 0,
1034 "message": {
1035 "role": "assistant",
1036 "content": "101 multiplied by 3 is 303.",
1037 "refusal": null
1038 },
1039 "finish_reason": "stop"
1040 }
1041 ],
1042 "usage": {
1043 "prompt_tokens": 31,
1044 "completion_tokens": 11,
1045 "total_tokens": 42,
1046 "prompt_tokens_details": {
1047 "text_tokens": 31,
1048 "audio_tokens": 0,
1049 "image_tokens": 0,
1050 "cached_tokens": 0
1051 },
1052 "completion_tokens_details": {
1053 "reasoning_tokens": 0,
1054 "audio_tokens": 0,
1055 "accepted_prediction_tokens": 0,
1056 "rejected_prediction_tokens": 0
1057 }
1058 },
1059 "system_fingerprint": "fp_156d35dcaa"
1060}
1061```