SpyBara
Go Premium

Documentation 2026-09-16 19:01 UTC to 2026-09-17 10:04 UTC

3 files changed +5 −1,064. View all changes and history on the product overview
2026
Wed 30 23:57 Tue 29 23:59 Mon 28 23:57 Sun 27 22:59 Sat 26 23:59 Fri 25 23:01 Thu 24 23:59 Wed 23 23:59 Tue 22 23:58 Mon 21 23:00 Sun 20 23:01 Sat 19 23:59 Fri 18 23:59 Thu 17 10:04 Wed 16 19:01 Tue 15 17:00 Mon 14 06:00 Sun 13 05:00 Fri 11 21:00 Tue 8 21:00 Mon 7 22:57 Thu 3 16:59 Wed 2 22:03

rate-limits.md +2 −2

Details

40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

43| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |43| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

44| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |45| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |46| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48 48 

49**Voice & Audio**49**Voice & Audio**

50 50 

rest-api-reference/inference/chat.md +0 −1061 deleted

File Deleted View Diff

1#### Inference API

2 

3# Chat

4 

5## POST /v1/chat/completions

6 

7Create a chat response from text/image chat prompts. This is the endpoint for making requests to chat and image understanding models.

8 

9### Request Body

10 

11* `deferred` (boolean | null) — If set to \`true\`, the request returns a \`request\_id\`. You can then get the deferred response by GET \`/v1/chat/deferred-completion/\{request\_id}\`.

12 

13* `frequency_penalty` (number | null) — (Not supported by reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.

14 

15* `logit_bias` (object | null) — (Unsupported) A JSON object that maps tokens (specified by their token ID in the tokenizer) to an associated bias value from -100 to 100. Mathematically, the bias is added to the logits generated by the model prior to sampling. The exact effect will vary per model, but values between -1 and 1 should decrease or increase likelihood of selection; values like -100 or 100 should result in a ban or exclusive selection of the relevant token.

16 

17* `logprobs` (boolean | null) — Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

18 

19* `max_completion_tokens` (integer | null) — An upper bound for the number of tokens that can be generated for a completion, only applies to visible output tokens (i.e. does not apply to tokens used for reasoning or function calls). Defaults to 128,000 when unset; set a larger value to allow longer generations.

20 

21* `max_tokens` (integer | null) — \\\[DEPRECATED\\] The maximum number of tokens that can be generated in the chat completion. Deprecated in favor of \`max\_completion\_tokens\`.

22 

23* `messages` (array\<object | object | object | object | object>) — A list of messages that make up the chat conversation. Different models support different message types, such as image and text.

24 

25* `model` (string) — Model name for the model to use. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

26 

27* `n` (integer | null) — How many chat completion choices to generate for each input message. Note that you will be charged based on the number of generated tokens across all of the choices. Keep n as 1 to minimize costs.

28 

29* `parallel_tool_calls` (boolean | null) — If set to false, the model can perform maximum one tool call.

30 

31* `presence_penalty` (number | null) — (Not supported by \`grok-3\` and reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

32 

33* `prompt_cache_key` (string | null) — A stable cache key for best-effort sticky routing / prompt-cache hits

34 across requests sharing a prompt prefix. Plumbed to \`x-grok-conv-id\`,

35 same as on \`/v1/responses\`.

36 

37* `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).

38 

39* `response_format` (object | object | object)

40 

41* `search_parameters` (object)

42 

43 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See

44 \<https://en.wikipedia.org/wiki/ISO\_8601>.

45 

46 * `max_search_results` (integer | null) — Maximum number of search results to use.

47 

48 * `mode` (string | null) — Choose the mode to query realtime data:

49 \* \`off\`: no search performed and no external will be considered.

50 \* \`on\` (default): the model will search in every sources for relevant data.

51 \* \`auto\`: the model choose whether to search data or not and where to search the data.

52 

53 * `return_citations` (boolean | null) — Whether to return citations in the response or not.

54 

55 * `sources` (array | null) — List of sources to search in. If no sources specified, the model will look over the web and X by default.

56 

57 * `to_date` (string | null) — Date up to which to consider the results in ISO-8601 YYYY-MM-DD. See

58 \<https://en.wikipedia.org/wiki/ISO\_8601>.

59 

60* `seed` (integer | null) — If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same \`seed\` and parameters should return the same result. Determinism is not guaranteed, and you should refer to the \`system\_fingerprint\` response parameter to monitor changes in the backend.

61 

62* `service_tier` ("default" | "priority")

63 

64* `stop` (array | null) — (Not supported by reasoning models) Up to 4 sequences where the API will stop generating further tokens.

65 

66* `stream` (boolean | null) — If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a \`data: \[DONE]\` message.

67 

68* `stream_options` (object)

69 

70 * `include_usage` (boolean, required) — Set an additional chunk to be streamed before the \`data: \[DONE]\` message. The other chunks will return \`null\` in \`usage\` field.

71 

72* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

73 

74* `tool_choice` (string | object)

75 

76* `tools` (array | null) — A list of tools the model may call in JSON-schema. Currently, only functions are supported as a tool. Use this to provide a list of functions the model may generate JSON inputs for. A max of 128 functions are supported.

77 

78* `top_logprobs` (integer | null) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position, each with an associated log probability. logprobs must be set to true if this parameter is used. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

79 

80* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

81 

82* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

83 

84* `web_search_options` (object)

85 

86 * `filters` (object) — Only included for compatibility.

87 

88 * `search_context_size` (string | null) — This field included for compatibility reason with OpenAI's API. It is mapped to \`max\_search\`.

89 

90 * `user_location` (object) — Only included for compatibility.

91 

92### Response Body

93 

94* `choices` (array\<object>, required) — A list of response choices from the model. The length corresponds to the \`n\` in request body (default to 1).

95 

96 * `finish_reason` (string | null) — Finish reason. \`"stop"\` means the inference has reached a model-defined or user-supplied stop sequence in \`stop\`. \`"length"\` means the inference result has reached models' maximum allowed token length or user defined value in \`max\_tokens\`. \`"end\_turn"\` or \`null\` in streaming mode when the chunk is not the last.

97 

98 * `index` (integer, required) — Index of the choice within the response choices, starting from 0.

99 

100 * `logprobs` (object)

101 

102 * `content` (array | null) — An array the log probabilities of each output token returned.

103 

104 * `message` (object, required)

105 

106 * `content` (string | null) — The content of the message.

107 

108 * `reasoning_content` (string | null) — The reasoning trace generated by the model.

109 

110 * `refusal` (string | null) — The reason given by model if the model is unable to generate a response. null if model is able to generate.

111 

112 * `role` (string, required) — The role that the message belongs to, the response from model is always \`"assistant"\`.

113 

114 * `tool_calls` (array | null) — A list of tool calls asked by model for user to perform.

115 

116* `citations` (array | null) — List of all the external pages used by the model to answer.

117 

118* `created` (integer, required) — The chat completion creation time in Unix timestamp.

119 

120* `id` (string, required) — A unique ID for the chat response.

121 

122* `model` (string, required) — Model ID used to create chat completion.

123 

124* `object` (string, required) — The object type, which is always \`"chat.completion"\`.

125 

126* `output_files` (array | null) — Files generated during the response (e.g., by the code execution tool).

127 Only populated when \`code\_execution\_files\_output\` is included.

128 

129* `service_tier` ("default" | "priority", required) — Processing tier for a request. Determines scheduling priority and billing.

130 

131* `system_fingerprint` (string | null) — System fingerprint, used to indicate xAI system configuration changes.

132 

133* `usage` (object)

134 

135 * `completion_tokens` (integer, required) — Total completion token used.

136 

137 * `completion_tokens_details` (object, required) — Details of completion usage.

138 

139 * `accepted_prediction_tokens` (integer, required) — The number of tokens in the prediction that appeared in the completion.

140 

141 * `audio_tokens` (integer, required) — Audio input tokens generated by the model.

142 

143 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

144 

145 * `rejected_prediction_tokens` (integer, required) — The number of tokens in the prediction that did not appear in the completion.

146 

147 * `cost_in_usd_ticks` (integer, required) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

148 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

149 which means there is 10'000'000'000 ticks in one \*dollar\*.

150 

151 * `num_sources_used` (integer, required) — Number of individual live search source used.

152 

153 * `prompt_tokens` (integer, required) — Total prompt token used.

154 

155 * `prompt_tokens_details` (object, required) — Details of prompt usage.

156 

157 * `audio_tokens` (integer, required) — Audio prompt token used.

158 

159 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

160 

161 * `image_tokens` (integer, required) — Image prompt token used.

162 

163 * `text_tokens` (integer, required) — Total text prompt token used (cached + non-cached text tokens).

164 

165 * `total_tokens` (integer, required) — Total token used, the sum of prompt token and completion token amount.

166 

167\*\*Request example:\*\*

168 

169```json

170{

171 "messages": [

172 {

173 "role": "system",

174 "content": "You are a helpful assistant that can answer questions and help with tasks."

175 },

176 {

177 "role": "user",

178 "content": "What is 101*3?"

179 }

180 ],

181 "model": "latest"

182}

183```

184 

185\*\*Response example:\*\*

186 

187```json

188{

189 "id": "a3d1008e-4544-40d4-d075-11527e794e4a",

190 "object": "chat.completion",

191 "created": 1752854522,

192 "model": "latest",

193 "choices": [

194 {

195 "index": 0,

196 "message": {

197 "role": "assistant",

198 "content": "101 multiplied by 3 is 303.",

199 "refusal": null

200 },

201 "finish_reason": "stop"

202 }

203 ],

204 "usage": {

205 "prompt_tokens": 32,

206 "completion_tokens": 9,

207 "total_tokens": 135,

208 "prompt_tokens_details": {

209 "text_tokens": 32,

210 "audio_tokens": 0,

211 "image_tokens": 0,

212 "cached_tokens": 6

213 },

214 "completion_tokens_details": {

215 "reasoning_tokens": 94,

216 "audio_tokens": 0,

217 "accepted_prediction_tokens": 0,

218 "rejected_prediction_tokens": 0

219 },

220 "num_sources_used": 0

221 },

222 "system_fingerprint": "fp_3a7881249c"

223}

224```

225 

226***

227 

228## POST /v1/responses

229 

230Generates a response based on text or image prompts. The response ID can be used to retrieve the response later or to continue the conversation without repeating prior context. New responses will be stored for 30 days and then permanently deleted.

231 

232### Request Body

233 

234* `background` (boolean | null) — (Unsupported) Whether to process the response asynchronously in the background.

235 

236* `context_management` (array | null) — Optional context-management directives (e.g. compaction). Parsed but not yet executed.

237 

238* `include` (array | null) — What additional output data to include in the response. Supported values include

239 \`reasoning.encrypted\_content\` (encrypted reasoning tokens) and tool-output options.

240 OpenAI's \`message.output\_text.logprobs\` is accepted for compatibility but silently ignored.

241 

242* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.

243 

244* `instructions` (string | null) — An alternate way to specify the system prompt. Note that this cannot be used alongside \`previous\_response\_id\`, where the system prompt of the previous message will be used.

245 

246* `logprobs` (boolean | null) — Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

247 

248* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens. Defaults to 128,000 when unset; set a larger value to allow longer generations.

249 

250* `max_turns` (integer | null) — Maximum number of agentic tool calling turns allowed for this request.

251 If not set, defaults to the server's global cap.

252 This parameter will be ignored for any non-agentic requests.

253 

254* `metadata` (object) — Not supported. Only maintained for compatibility reasons.

255 

256* `min_p` (number | null) — Min-p sampling: tokens whose probability is below \`min\_p\` times the probability of the most likely token are excluded from sampling. Disabled when unset.

257 

258* `model` (string) — Model name for the model to use. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

259 

260* `parallel_tool_calls` (boolean | null) — Whether to allow the model to run parallel tool calls.

261 

262* `previous_response_id` (string | null) — The ID of the previous response from the model.

263 

264* `prompt_cache_key` (string | null) — Plumbed to x-grok-conv-id for Open Responses compatibility, used for routing.

265 

266* `reasoning` (object)

267 

268 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).

269 

270 * `generate_summary` (string | null) — Only included for compatibility.

271 

272 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.

273 

274* `reasoning_effort` (string | null) — reasoning\_effort alternative to reasoning configuration. This is a non-standard field meant to ease user experience. We only look at this if the reasoning field is unset.

275 

276* `search_parameters` (object)

277 

278 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See

279 \<https://en.wikipedia.org/wiki/ISO\_8601>.

280 

281 * `max_search_results` (integer | null) — Maximum number of search results to use.

282 

283 * `mode` (string | null) — Choose the mode to query realtime data:

284 \* \`off\`: no search performed and no external will be considered.

285 \* \`on\` (default): the model will search in every sources for relevant data.

286 \* \`auto\`: the model choose whether to search data or not and where to search the data.

287 

288 * `return_citations` (boolean | null) — Whether to return citations in the response or not.

289 

290 * `sources` (array | null) — List of sources to search in. If no sources specified, the model will look over the web and X by default.

291 

292 * `to_date` (string | null) — Date up to which to consider the results in ISO-8601 YYYY-MM-DD. See

293 \<https://en.wikipedia.org/wiki/ISO\_8601>.

294 

295* `service_tier` ("default" | "priority")

296 

297* `store` (boolean | null) — Whether to store the input message(s) and model response for later retrieval.

298 

299* `stream` (boolean | null) — If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a \`data: \[DONE]\` message.

300 

301* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

302 

303* `text` (object)

304 

305 * `format` (object | object | object)

306 

307* `tool_choice` (string | object)

308 

309* `tools` (array | null) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.\`web\_search\_preview\` tool, if specified, will be overridden by \`search\_parameters\`.

310 

311* `top_k` (integer | null) — Top-k sampling: only the \`top\_k\` most probable tokens are considered at each sampling step. Disabled when unset.

312 

313* `top_logprobs` (integer | null) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position, each with an associated log probability. logprobs must be set to true if this parameter is used. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

314 

315* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

316 

317* `truncation` (string | null) — Not supported. Only maintained for compatibility reasons.

318 

319* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

320 

321### Response Body

322 

323* `background` (boolean, required) — OpenResponses compatibility fields.

324 Not used at the moment. Just for OpenResponses compatibility.

325 Whether to process the response asynchronously in the background.

326 

327* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.

328 

329* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.

330 

331* `error` (object) — An error object returned when the model fails to generate a response.

332 

333* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.

334 

335* `id` (string, required) — Unique ID of the response.

336 

337* `incomplete_details` (object | object | object)

338 

339* `instructions` (string | null) — A system (or developer) message inserted into the model's context.

340 

341* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.

342 

343* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.

344 

345* `metadata` (object, required) — Only included for compatibility.

346 

347* `model` (string, required) — Model name used to generate the response.

348 

349* `object` (string, required) — The object type of this resource. Always set to \`response\`.

350 

351* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.

352 

353* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.

354 

355* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

356 

357* `previous_response_id` (string | null) — The ID of the previous response from the model.

358 

359* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.

360 

361* `reasoning` (object)

362 

363 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).

364 

365 * `generate_summary` (string | null) — Only included for compatibility.

366 

367 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.

368 

369* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.

370 

371* `service_tier` ("default" | "priority", required)

372 

373* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.

374 

375* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.

376 

377* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

378 

379* `text` (object, required)

380 

381 * `format` (object | object | object)

382 

383* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.

384 

385 * `name` (string, required) — Name of the function to use.

386 

387 * `type` (string, required) — Type is always \`"function"\`.

388 

389* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.

390 

391* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.

392 

393* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

394 

395* `truncation` (string, required) — The truncation strategy to use for the model response.

396 

397* `usage` (object)

398 

399 * `context_details` (object)

400 

401 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from

402 \`SamplingUsage.context\_prompt\_tokens\`).

403 

404 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from

405 \`SamplingUsage.context\_output\_tokens\`).

406 

407 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.

408 

409 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

410 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

411 which means there is 10'000'000'000 ticks in one \*dollar\*.

412 

413 * `input_tokens` (integer, required) — Number of input tokens used.

414 

415 * `input_tokens_details` (object, required)

416 

417 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

418 

419 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.

420 

421 * `num_sources_used` (integer, required) — Number of sources used (for live search).

422 

423 * `output_tokens` (integer, required) — Number of output tokens used.

424 

425 * `output_tokens_details` (object, required)

426 

427 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

428 

429 * `server_side_tool_usage_details` (object)

430 

431 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.

432 

433 * `document_search_calls` (integer, required) — Number of document search calls.

434 

435 * `file_search_calls` (integer, required) — Number of file search calls.

436 

437 * `image_generation_calls` (integer, required) — Number of image generation calls.

438 

439 * `mcp_calls` (integer, required) — Number of MCP calls.

440 

441 * `web_search_calls` (integer, required) — Number of web search calls.

442 

443 * `x_search_calls` (integer, required) — Number of X search calls.

444 

445 * `total_tokens` (integer, required) — Total tokens used.

446 

447* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

448 

449### Code Examples

450 

451```bash

452curl -s https://api.x.ai/v1/responses \

453 -H "Content-Type: application/json" \

454 -H "Authorization: Bearer $XAI_API_KEY" \

455 -d '{

456 "model": "grok-4.6",

457 "input": "What is the meaning of life?"

458 }'

459```

460 

461```javascriptAISDK

462import { xai } from "@ai-sdk/xai";

463import { generateText } from "ai";

464 

465const result = await generateText({

466 model: xai.responses("grok-4.6"),

467 prompt: "What is the meaning of life?",

468});

469 

470console.log(JSON.stringify(result, null, 2));

471```

472 

473```pythonOpenAISDK

474import os

475 

476from openai import OpenAI

477 

478client = OpenAI(

479 api_key=os.environ["XAI_API_KEY"],

480 base_url="https://api.x.ai/v1",

481)

482 

483response = client.responses.create(

484 model="grok-4.6",

485 input="What is the meaning of life?",

486)

487 

488print(response.model_dump_json(indent=2))

489```

490 

491```javascriptOpenAISDK

492import OpenAI from "openai";

493 

494const client = new OpenAI({

495 apiKey: process.env.XAI_API_KEY,

496 baseURL: "https://api.x.ai/v1",

497});

498 

499const response = await client.responses.create({

500 model: "grok-4.6",

501 input: "What is the meaning of life?",

502});

503 

504console.log(JSON.stringify(response, null, 2));

505```

506 

507\*\*Response example:\*\*

508 

509```json

510{

511 "created_at": 1754475266,

512 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",

513 "max_output_tokens": null,

514 "model": "latest",

515 "object": "response",

516 "output": [

517 {

518 "content": [

519 {

520 "type": "output_text",

521 "text": "101 multiplied by 3 is 303.",

522 "logprobs": null,

523 "annotations": []

524 }

525 ],

526 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",

527 "role": "assistant",

528 "type": "message",

529 "status": "completed"

530 }

531 ],

532 "parallel_tool_calls": true,

533 "previous_response_id": null,

534 "reasoning": null,

535 "temperature": null,

536 "text": {

537 "format": {

538 "type": "text"

539 }

540 },

541 "tool_choice": "auto",

542 "tools": [],

543 "top_p": null,

544 "usage": {

545 "input_tokens": 32,

546 "input_tokens_details": {

547 "cached_tokens": 8

548 },

549 "output_tokens": 9,

550 "output_tokens_details": {

551 "reasoning_tokens": 110

552 },

553 "total_tokens": 151,

554 "num_sources_used": 0,

555 "num_server_side_tools_used": 0

556 },

557 "user": null,

558 "incomplete_details": null,

559 "status": "completed",

560 "store": true

561}

562```

563 

564***

565 

566## POST /v1/responses/compact

567 

568Compacts a full Responses API input window into a shorter canonical window.

569 

570### Request Body

571 

572* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.

573 

574* `model` (string, required) — Model to use for compaction summarization (required).

575 

576### Response Body

577 

578* `created_at` (integer, required) — Unix timestamp (in seconds) when the compacted conversation was created.

579 

580* `id` (string, required) — Unique ID for this compaction (e.g. \`cmp\_\<uuid>\`).

581 

582* `model` (string, required) — Model used for the compaction summary.

583 

584* `object` (string, required) — Always \`"response.compaction"\`.

585 

586* `output` (array\<object | object | object | object | object>, required) — Compacted output containing a single compaction item.

587 Pass this verbatim as input to the next \`/v1/responses\` call.

588 

589* `usage` (object)

590 

591 * `dropped_message_count` (integer, required) — Number of messages dropped/summarized during compaction.

592 

593 * `input_tokens` (integer, required) — Number of tokens in the original (pre-compaction) input.

594 

595 * `input_tokens_details` (object, required) — Breakdown of input tokens for a compaction call.

596 

597 * `cached_tokens` (integer, required) — Number of input tokens that were served from the prompt cache.

598 

599 * `output_tokens` (integer, required) — Number of tokens in the compacted output.

600 

601 * `output_tokens_details` (object, required) — Breakdown of output tokens for a compaction call.

602 

603 * `reasoning_tokens` (integer, required) — Number of reasoning tokens generated during compaction.

604 

605 * `total_tokens` (integer, required) — Total number of tokens used (input + output, including reasoning).

606 

607### Code Examples

608 

609```bash

610curl -s https://api.x.ai/v1/responses/compact \

611 -H "Content-Type: application/json" \

612 -H "Authorization: Bearer $XAI_API_KEY" \

613 -d '{

614 "model": "grok-4.6",

615 "input": [

616 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},

617 {"role": "user", "content": "What is the Higgs boson and why is it important?"},

618 {"role": "assistant", "content": "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."},

619 {"role": "user", "content": "How does the Higgs mechanism actually work?"},

620 {"role": "assistant", "content": "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless."}

621 ]

622 }'

623```

624 

625```pythonOpenAISDK

626import os

627 

628from openai import OpenAI

629 

630client = OpenAI(

631 api_key=os.environ["XAI_API_KEY"],

632 base_url="https://api.x.ai/v1",

633)

634 

635compacted = client.responses.compact(

636 model="grok-4.6",

637 input=[

638 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},

639 {"role": "user", "content": "What is the Higgs boson and why is it important?"},

640 {

641 "role": "assistant",

642 "content": (

643 "The Higgs boson is an elementary particle in the Standard Model, predicted by "

644 "Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation "

645 "of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."

646 ),

647 },

648 {"role": "user", "content": "How does the Higgs mechanism actually work?"},

649 {

650 "role": "assistant",

651 "content": (

652 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, "

653 "and particles acquire mass in proportion to how strongly they couple to it. Photons "

654 "do not couple, which is why they remain massless."

655 ),

656 },

657 ],

658)

659 

660print(compacted.model_dump_json(indent=2))

661```

662 

663```javascriptOpenAISDK

664import OpenAI from "openai";

665 

666const client = new OpenAI({

667 apiKey: process.env.XAI_API_KEY,

668 baseURL: "https://api.x.ai/v1",

669});

670 

671const compacted = await client.responses.compact({

672 model: "grok-4.6",

673 input: [

674 { role: "system", content: "You are a concise and knowledgeable science tutor." },

675 { role: "user", content: "What is the Higgs boson and why is it important?" },

676 {

677 role: "assistant",

678 content:

679 "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism.",

680 },

681 { role: "user", content: "How does the Higgs mechanism actually work?" },

682 {

683 role: "assistant",

684 content:

685 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless.",

686 },

687 ],

688});

689 

690console.log(JSON.stringify(compacted, null, 2));

691```

692 

693\*\*Response example:\*\*

694 

695```json

696{}

697```

698 

699***

700 

701## GET /v1/responses/\{response\_id}

702 

703Retrieve a previously generated response.

704 

705### Path Parameters

706 

707* `response_id` (string, required) — The response id returned by a previous create response request.

708 

709### Response Body

710 

711* `background` (boolean, required) — OpenResponses compatibility fields.

712 Not used at the moment. Just for OpenResponses compatibility.

713 Whether to process the response asynchronously in the background.

714 

715* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.

716 

717* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.

718 

719* `error` (object) — An error object returned when the model fails to generate a response.

720 

721* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.

722 

723* `id` (string, required) — Unique ID of the response.

724 

725* `incomplete_details` (object | object | object)

726 

727* `instructions` (string | null) — A system (or developer) message inserted into the model's context.

728 

729* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.

730 

731* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.

732 

733* `metadata` (object, required) — Only included for compatibility.

734 

735* `model` (string, required) — Model name used to generate the response.

736 

737* `object` (string, required) — The object type of this resource. Always set to \`response\`.

738 

739* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.

740 

741* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.

742 

743* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

744 

745* `previous_response_id` (string | null) — The ID of the previous response from the model.

746 

747* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.

748 

749* `reasoning` (object)

750 

751 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Only supported by \`grok-4.3\`. Possible values are \`none\` (disables reasoning completely), \`low\` (this is the default if not specified), \`medium\` and \`high\` (uses the most reasoning tokens).

752 

753 * `generate_summary` (string | null) — Only included for compatibility.

754 

755 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.

756 

757* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.

758 

759* `service_tier` ("default" | "priority", required)

760 

761* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.

762 

763* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.

764 

765* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

766 

767* `text` (object, required)

768 

769 * `format` (object | object | object)

770 

771* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.

772 

773 * `name` (string, required) — Name of the function to use.

774 

775 * `type` (string, required) — Type is always \`"function"\`.

776 

777* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.

778 

779* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.

780 

781* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

782 

783* `truncation` (string, required) — The truncation strategy to use for the model response.

784 

785* `usage` (object)

786 

787 * `context_details` (object)

788 

789 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from

790 \`SamplingUsage.context\_prompt\_tokens\`).

791 

792 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from

793 \`SamplingUsage.context\_output\_tokens\`).

794 

795 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.

796 

797 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

798 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

799 which means there is 10'000'000'000 ticks in one \*dollar\*.

800 

801 * `input_tokens` (integer, required) — Number of input tokens used.

802 

803 * `input_tokens_details` (object, required)

804 

805 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

806 

807 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.

808 

809 * `num_sources_used` (integer, required) — Number of sources used (for live search).

810 

811 * `output_tokens` (integer, required) — Number of output tokens used.

812 

813 * `output_tokens_details` (object, required)

814 

815 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

816 

817 * `server_side_tool_usage_details` (object)

818 

819 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.

820 

821 * `document_search_calls` (integer, required) — Number of document search calls.

822 

823 * `file_search_calls` (integer, required) — Number of file search calls.

824 

825 * `image_generation_calls` (integer, required) — Number of image generation calls.

826 

827 * `mcp_calls` (integer, required) — Number of MCP calls.

828 

829 * `web_search_calls` (integer, required) — Number of web search calls.

830 

831 * `x_search_calls` (integer, required) — Number of X search calls.

832 

833 * `total_tokens` (integer, required) — Total tokens used.

834 

835* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

836 

837\*\*Response example:\*\*

838 

839```json

840{

841 "created_at": 1754475266,

842 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",

843 "max_output_tokens": null,

844 "model": "latest",

845 "object": "response",

846 "output": [

847 {

848 "content": [

849 {

850 "type": "output_text",

851 "text": "101 multiplied by 3 is 303.",

852 "logprobs": null,

853 "annotations": []

854 }

855 ],

856 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",

857 "role": "assistant",

858 "type": "message",

859 "status": "completed"

860 },

861 {

862 "id": "",

863 "summary": [

864 {

865 "text": "First, the user asked: \"What is 101*3?\"\n\nThis is a simple multiplication: 101 multiplied by 3.\n\nCalculating: 100 * 3 = 300, and 1 * 3 = 3, so 300 + 3 = 303.\n\nI should respond helpfully and directly, as per my system prompt: \"You are a helpful assistant that can answer questions and help with tasks.\"\n\nKeep the response concise and accurate. No need for extra fluff unless it adds value.\n\nFinal answer: 303.",

866 "type": "summary_text"

867 }

868 ],

869 "type": "reasoning",

870 "status": "completed"

871 }

872 ],

873 "parallel_tool_calls": true,

874 "previous_response_id": null,

875 "reasoning": null,

876 "temperature": null,

877 "text": {

878 "format": {

879 "type": "text"

880 }

881 },

882 "tool_choice": "auto",

883 "tools": [],

884 "top_p": null,

885 "usage": {

886 "prompt_tokens": 32,

887 "completion_tokens": 9,

888 "total_tokens": 151,

889 "prompt_tokens_details": {

890 "text_tokens": 32,

891 "audio_tokens": 0,

892 "image_tokens": 0,

893 "cached_tokens": 8

894 },

895 "completion_tokens_details": {

896 "reasoning_tokens": 110,

897 "audio_tokens": 0,

898 "accepted_prediction_tokens": 0,

899 "rejected_prediction_tokens": 0

900 },

901 "num_sources_used": 0

902 },

903 "user": null,

904 "incomplete_details": null,

905 "status": "completed",

906 "store": true

907}

908```

909 

910***

911 

912## DELETE /v1/responses/\{response\_id}

913 

914Delete a previously generated response.

915 

916### Path Parameters

917 

918* `response_id` (string, required) — The response id returned by a previous create response request.

919 

920### Response Body

921 

922* `deleted` (boolean, required) — Whether the response was successfully deleted.

923 

924* `id` (string, required) — The response\_id to be deleted.

925 

926* `object` (string, required) — The deleted object type, which is always \`response\`.

927 

928\*\*Response example:\*\*

929 

930```json

931{

932 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",

933 "object": "response",

934 "deleted": true

935}

936```

937 

938***

939 

940## GET /v1/chat/deferred-completion/\{request\_id}

941 

942Tries to fetch a result for a previously-started deferred completion. Returns \`200 Success\` with the response body, if the request has been completed. Returns \`202 Accepted\` when the request is pending processing.

943 

944### Path Parameters

945 

946* `request_id` (string, required) — The deferred request id returned by a previous deferred chat request.

947 

948### Response Body

949 

950* `choices` (array\<object>, required) — A list of response choices from the model. The length corresponds to the \`n\` in request body (default to 1).

951 

952 * `finish_reason` (string | null) — Finish reason. \`"stop"\` means the inference has reached a model-defined or user-supplied stop sequence in \`stop\`. \`"length"\` means the inference result has reached models' maximum allowed token length or user defined value in \`max\_tokens\`. \`"end\_turn"\` or \`null\` in streaming mode when the chunk is not the last.

953 

954 * `index` (integer, required) — Index of the choice within the response choices, starting from 0.

955 

956 * `logprobs` (object)

957 

958 * `content` (array | null) — An array the log probabilities of each output token returned.

959 

960 * `message` (object, required)

961 

962 * `content` (string | null) — The content of the message.

963 

964 * `reasoning_content` (string | null) — The reasoning trace generated by the model.

965 

966 * `refusal` (string | null) — The reason given by model if the model is unable to generate a response. null if model is able to generate.

967 

968 * `role` (string, required) — The role that the message belongs to, the response from model is always \`"assistant"\`.

969 

970 * `tool_calls` (array | null) — A list of tool calls asked by model for user to perform.

971 

972* `citations` (array | null) — List of all the external pages used by the model to answer.

973 

974* `created` (integer, required) — The chat completion creation time in Unix timestamp.

975 

976* `id` (string, required) — A unique ID for the chat response.

977 

978* `model` (string, required) — Model ID used to create chat completion.

979 

980* `object` (string, required) — The object type, which is always \`"chat.completion"\`.

981 

982* `output_files` (array | null) — Files generated during the response (e.g., by the code execution tool).

983 Only populated when \`code\_execution\_files\_output\` is included.

984 

985* `service_tier` ("default" | "priority", required) — Processing tier for a request. Determines scheduling priority and billing.

986 

987* `system_fingerprint` (string | null) — System fingerprint, used to indicate xAI system configuration changes.

988 

989* `usage` (object)

990 

991 * `completion_tokens` (integer, required) — Total completion token used.

992 

993 * `completion_tokens_details` (object, required) — Details of completion usage.

994 

995 * `accepted_prediction_tokens` (integer, required) — The number of tokens in the prediction that appeared in the completion.

996 

997 * `audio_tokens` (integer, required) — Audio input tokens generated by the model.

998 

999 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

1000 

1001 * `rejected_prediction_tokens` (integer, required) — The number of tokens in the prediction that did not appear in the completion.

1002 

1003 * `cost_in_usd_ticks` (integer, required) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

1004 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

1005 which means there is 10'000'000'000 ticks in one \*dollar\*.

1006 

1007 * `num_sources_used` (integer, required) — Number of individual live search source used.

1008 

1009 * `prompt_tokens` (integer, required) — Total prompt token used.

1010 

1011 * `prompt_tokens_details` (object, required) — Details of prompt usage.

1012 

1013 * `audio_tokens` (integer, required) — Audio prompt token used.

1014 

1015 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

1016 

1017 * `image_tokens` (integer, required) — Image prompt token used.

1018 

1019 * `text_tokens` (integer, required) — Total text prompt token used (cached + non-cached text tokens).

1020 

1021 * `total_tokens` (integer, required) — Total token used, the sum of prompt token and completion token amount.

1022 

1023\*\*Response example:\*\*

1024 

1025```json

1026{

1027 "id": "335b92e4-afa5-48e7-b99c-b9a4eabc1c8e",

1028 "object": "chat.completion",

1029 "created": 1743770624,

1030 "model": "latest",

1031 "choices": [

1032 {

1033 "index": 0,

1034 "message": {

1035 "role": "assistant",

1036 "content": "101 multiplied by 3 is 303.",

1037 "refusal": null

1038 },

1039 "finish_reason": "stop"

1040 }

1041 ],

1042 "usage": {

1043 "prompt_tokens": 31,

1044 "completion_tokens": 11,

1045 "total_tokens": 42,

1046 "prompt_tokens_details": {

1047 "text_tokens": 31,

1048 "audio_tokens": 0,

1049 "image_tokens": 0,

1050 "cached_tokens": 0

1051 },

1052 "completion_tokens_details": {

1053 "reasoning_tokens": 0,

1054 "audio_tokens": 0,

1055 "accepted_prediction_tokens": 0,

1056 "rejected_prediction_tokens": 0

1057 }

1058 },

1059 "system_fingerprint": "fp_156d35dcaa"

1060}

1061```

Details

30 30 

31* `max_turns` (integer | null) — Maximum number of agentic tool calling turns allowed for this request.31* `max_turns` (integer | null) — Maximum number of agentic tool calling turns allowed for this request.

32 If not set, defaults to the server's global cap.32 If not set, defaults to the server's global cap.

33 This parameter will be ignored for any non-agentic requests.33 This parameter will be ignored for any non-agentic requests, and for

34 agentic SLOP requests that have neither a server-side tool nor a file

35 attachment.

34 36 

35* `metadata` (object) — Not supported. Only maintained for compatibility reasons.37* `metadata` (object) — Not supported. Only maintained for compatibility reasons.

36 38