1#### Inference API
2
3# Responses
4
5The Responses API is the primary interface for text generation, reasoning, and tool use. See the [Text Generation guide](/developers/model-capabilities/text/generate-text) for usage.
6
7***
8
9## POST /v1/responses
10
11Generates a response based on text or image prompts. The response ID can be used to retrieve the response later or to continue the conversation without repeating prior context. New responses will be stored for 30 days and then permanently deleted.
12
13### Request Body
14
15* `background` (boolean | null) — (Unsupported) Whether to process the response asynchronously in the background.
16
17* `context_management` (array | null) — Optional context-management directives (e.g. compaction). Parsed but not yet executed.
18
19* `include` (array | null) — What additional output data to include in the response. Supported values include
20 \`reasoning.encrypted\_content\` (encrypted reasoning tokens) and tool-output options.
21 OpenAI's \`message.output\_text.logprobs\` is accepted for compatibility but silently ignored.
22
23* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.
24
25* `instructions` (string | null) — An alternate way to specify the system prompt. Note that this cannot be used alongside \`previous\_response\_id\`, where the system prompt of the previous message will be used.
26
27* `logprobs` (boolean | null) — Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.
28
29* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens. Defaults to 128,000 when unset; set a larger value to allow longer generations.
30
31* `max_turns` (integer | null) — Maximum number of agentic tool calling turns allowed for this request.
32 If not set, defaults to the server's global cap.
33 This parameter will be ignored for any non-agentic requests.
34
35* `metadata` (object) — Not supported. Only maintained for compatibility reasons.
36
37* `min_p` (number | null) — Min-p sampling: tokens whose probability is below \`min\_p\` times the probability of the most likely token are excluded from sampling. Disabled when unset.
38
39* `model` (string) — Model name for the model to use. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.
40
41* `parallel_tool_calls` (boolean | null) — Whether to allow the model to run parallel tool calls.
42
43* `previous_response_id` (string | null) — The ID of the previous response from the model.
44
45* `prompt_cache_key` (string | null) — Plumbed to x-grok-conv-id for Open Responses compatibility, used for routing.
46
47* `reasoning` (object)
48
49 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.
50
51 * `generate_summary` (string | null) — Only included for compatibility.
52
53 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.
54
55* `reasoning_effort` (string | null) — reasoning\_effort alternative to reasoning configuration. This is a non-standard field meant to ease user experience. We only look at this if the reasoning field is unset.
56
57* `search_parameters` (object)
58
59 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See
60 \<https://en.wikipedia.org/wiki/ISO\_8601>.
61
62 * `max_search_results` (integer | null) — Maximum number of search results to use.
63
64 * `mode` (string | null) — Choose the mode to query realtime data:
65 \* \`off\`: no search performed and no external will be considered.
66 \* \`on\` (default): the model will search in every sources for relevant data.
67 \* \`auto\`: the model choose whether to search data or not and where to search the data.
68
69 * `return_citations` (boolean | null) — Whether to return citations in the response or not.
70
71 * `sources` (array | null) — List of sources to search in. If no sources specified, the model will look over the web and X by default.
72
73 * `to_date` (string | null) — Date up to which to consider the results in ISO-8601 YYYY-MM-DD. See
74 \<https://en.wikipedia.org/wiki/ISO\_8601>.
75
76* `service_tier` ("default" | "priority")
77
78* `store` (boolean | null) — Whether to store the input message(s) and model response for later retrieval.
79
80* `stream` (boolean | null) — If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a \`data: \[DONE]\` message.
81
82* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
83
84* `text` (object)
85
86 * `format` (object | object | object)
87
88* `tool_choice` (string | object)
89
90* `tools` (array | null) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.\`web\_search\_preview\` tool, if specified, will be overridden by \`search\_parameters\`.
91
92* `top_k` (integer | null) — Top-k sampling: only the \`top\_k\` most probable tokens are considered at each sampling step. Disabled when unset.
93
94* `top_logprobs` (integer | null) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position, each with an associated log probability. logprobs must be set to true if this parameter is used. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.
95
96* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.
97
98* `truncation` (string | null) — Not supported. Only maintained for compatibility reasons.
99
100* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.
101
102### Response Body
103
104* `background` (boolean, required) — OpenResponses compatibility fields.
105 Not used at the moment. Just for OpenResponses compatibility.
106 Whether to process the response asynchronously in the background.
107
108* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.
109
110* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.
111
112* `error` (object) — An error object returned when the model fails to generate a response.
113
114* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
115
116* `id` (string, required) — Unique ID of the response.
117
118* `incomplete_details` (object | object | object)
119
120* `instructions` (string | null) — A system (or developer) message inserted into the model's context.
121
122* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.
123
124* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.
125
126* `metadata` (object, required) — Only included for compatibility.
127
128* `model` (string, required) — Model name used to generate the response.
129
130* `object` (string, required) — The object type of this resource. Always set to \`response\`.
131
132* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.
133
134* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.
135
136* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
137
138* `previous_response_id` (string | null) — The ID of the previous response from the model.
139
140* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.
141
142* `reasoning` (object)
143
144 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.
145
146 * `generate_summary` (string | null) — Only included for compatibility.
147
148 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.
149
150* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.
151
152* `service_tier` ("default" | "priority", required)
153
154* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.
155
156* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.
157
158* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
159
160* `text` (object, required)
161
162 * `format` (object | object | object)
163
164* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.
165
166 * `name` (string, required) — Name of the function to use.
167
168 * `type` (string, required) — Type is always \`"function"\`.
169
170* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.
171
172* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.
173
174* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.
175
176* `truncation` (string, required) — The truncation strategy to use for the model response.
177
178* `usage` (object)
179
180 * `context_details` (object)
181
182 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from
183 \`SamplingUsage.context\_prompt\_tokens\`).
184
185 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from
186 \`SamplingUsage.context\_output\_tokens\`).
187
188 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.
189
190 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:
191 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000
192 which means there is 10'000'000'000 ticks in one \*dollar\*.
193
194 * `input_tokens` (integer, required) — Number of input tokens used.
195
196 * `input_tokens_details` (object, required)
197
198 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.
199
200 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.
201
202 * `num_sources_used` (integer, required) — Number of sources used (for live search).
203
204 * `output_tokens` (integer, required) — Number of output tokens used.
205
206 * `output_tokens_details` (object, required)
207
208 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.
209
210 * `server_side_tool_usage_details` (object)
211
212 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.
213
214 * `document_search_calls` (integer, required) — Number of document search calls.
215
216 * `file_search_calls` (integer, required) — Number of file search calls.
217
218 * `image_generation_calls` (integer, required) — Number of image generation calls.
219
220 * `mcp_calls` (integer, required) — Number of MCP calls.
221
222 * `web_search_calls` (integer, required) — Number of web search calls.
223
224 * `x_search_calls` (integer, required) — Number of X search calls.
225
226 * `total_tokens` (integer, required) — Total tokens used.
227
228* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.
229
230### Code Examples
231
232```bash
233curl -s https://api.x.ai/v1/responses \
234 -H "Content-Type: application/json" \
235 -H "Authorization: Bearer $XAI_API_KEY" \
236 -d '{
237 "model": "grok-4.6",
238 "input": "What is the meaning of life?"
239 }'
240```
241
242```javascriptAISDK
243import { xai } from "@ai-sdk/xai";
244import { generateText } from "ai";
245
246const result = await generateText({
247 model: xai.responses("grok-4.6"),
248 prompt: "What is the meaning of life?",
249});
250
251console.log(JSON.stringify(result, null, 2));
252```
253
254```pythonOpenAISDK
255import os
256
257from openai import OpenAI
258
259client = OpenAI(
260 api_key=os.environ["XAI_API_KEY"],
261 base_url="https://api.x.ai/v1",
262)
263
264response = client.responses.create(
265 model="grok-4.6",
266 input="What is the meaning of life?",
267)
268
269print(response.model_dump_json(indent=2))
270```
271
272```javascriptOpenAISDK
273import OpenAI from "openai";
274
275const client = new OpenAI({
276 apiKey: process.env.XAI_API_KEY,
277 baseURL: "https://api.x.ai/v1",
278});
279
280const response = await client.responses.create({
281 model: "grok-4.6",
282 input: "What is the meaning of life?",
283});
284
285console.log(JSON.stringify(response, null, 2));
286```
287
288\*\*Response example:\*\*
289
290```json
291{
292 "created_at": 1754475266,
293 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",
294 "max_output_tokens": null,
295 "model": "latest",
296 "object": "response",
297 "output": [
298 {
299 "content": [
300 {
301 "type": "output_text",
302 "text": "101 multiplied by 3 is 303.",
303 "logprobs": null,
304 "annotations": []
305 }
306 ],
307 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",
308 "role": "assistant",
309 "type": "message",
310 "status": "completed"
311 }
312 ],
313 "parallel_tool_calls": true,
314 "previous_response_id": null,
315 "reasoning": null,
316 "temperature": null,
317 "text": {
318 "format": {
319 "type": "text"
320 }
321 },
322 "tool_choice": "auto",
323 "tools": [],
324 "top_p": null,
325 "usage": {
326 "input_tokens": 32,
327 "input_tokens_details": {
328 "cached_tokens": 8
329 },
330 "output_tokens": 9,
331 "output_tokens_details": {
332 "reasoning_tokens": 110
333 },
334 "total_tokens": 151,
335 "num_sources_used": 0,
336 "num_server_side_tools_used": 0
337 },
338 "user": null,
339 "incomplete_details": null,
340 "status": "completed",
341 "store": true
342}
343```
344
345***
346
347## GET /v1/responses/\{response\_id}
348
349Retrieve a previously generated response.
350
351### Path Parameters
352
353* `response_id` (string, required) — The response id returned by a previous create response request.
354
355### Response Body
356
357* `background` (boolean, required) — OpenResponses compatibility fields.
358 Not used at the moment. Just for OpenResponses compatibility.
359 Whether to process the response asynchronously in the background.
360
361* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.
362
363* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.
364
365* `error` (object) — An error object returned when the model fails to generate a response.
366
367* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.
368
369* `id` (string, required) — Unique ID of the response.
370
371* `incomplete_details` (object | object | object)
372
373* `instructions` (string | null) — A system (or developer) message inserted into the model's context.
374
375* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.
376
377* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.
378
379* `metadata` (object, required) — Only included for compatibility.
380
381* `model` (string, required) — Model name used to generate the response.
382
383* `object` (string, required) — The object type of this resource. Always set to \`response\`.
384
385* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.
386
387* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.
388
389* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.
390
391* `previous_response_id` (string | null) — The ID of the previous response from the model.
392
393* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.
394
395* `reasoning` (object)
396
397 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.
398
399 * `generate_summary` (string | null) — Only included for compatibility.
400
401 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.
402
403* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.
404
405* `service_tier` ("default" | "priority", required)
406
407* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.
408
409* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.
410
411* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
412
413* `text` (object, required)
414
415 * `format` (object | object | object)
416
417* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.
418
419 * `name` (string, required) — Name of the function to use.
420
421 * `type` (string, required) — Type is always \`"function"\`.
422
423* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.
424
425* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.
426
427* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.
428
429* `truncation` (string, required) — The truncation strategy to use for the model response.
430
431* `usage` (object)
432
433 * `context_details` (object)
434
435 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from
436 \`SamplingUsage.context\_prompt\_tokens\`).
437
438 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from
439 \`SamplingUsage.context\_output\_tokens\`).
440
441 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.
442
443 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:
444 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000
445 which means there is 10'000'000'000 ticks in one \*dollar\*.
446
447 * `input_tokens` (integer, required) — Number of input tokens used.
448
449 * `input_tokens_details` (object, required)
450
451 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.
452
453 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.
454
455 * `num_sources_used` (integer, required) — Number of sources used (for live search).
456
457 * `output_tokens` (integer, required) — Number of output tokens used.
458
459 * `output_tokens_details` (object, required)
460
461 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.
462
463 * `server_side_tool_usage_details` (object)
464
465 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.
466
467 * `document_search_calls` (integer, required) — Number of document search calls.
468
469 * `file_search_calls` (integer, required) — Number of file search calls.
470
471 * `image_generation_calls` (integer, required) — Number of image generation calls.
472
473 * `mcp_calls` (integer, required) — Number of MCP calls.
474
475 * `web_search_calls` (integer, required) — Number of web search calls.
476
477 * `x_search_calls` (integer, required) — Number of X search calls.
478
479 * `total_tokens` (integer, required) — Total tokens used.
480
481* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.
482
483\*\*Response example:\*\*
484
485```json
486{
487 "created_at": 1754475266,
488 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",
489 "max_output_tokens": null,
490 "model": "latest",
491 "object": "response",
492 "output": [
493 {
494 "content": [
495 {
496 "type": "output_text",
497 "text": "101 multiplied by 3 is 303.",
498 "logprobs": null,
499 "annotations": []
500 }
501 ],
502 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",
503 "role": "assistant",
504 "type": "message",
505 "status": "completed"
506 },
507 {
508 "id": "",
509 "summary": [
510 {
511 "text": "First, the user asked: \"What is 101*3?\"\n\nThis is a simple multiplication: 101 multiplied by 3.\n\nCalculating: 100 * 3 = 300, and 1 * 3 = 3, so 300 + 3 = 303.\n\nI should respond helpfully and directly, as per my system prompt: \"You are a helpful assistant that can answer questions and help with tasks.\"\n\nKeep the response concise and accurate. No need for extra fluff unless it adds value.\n\nFinal answer: 303.",
512 "type": "summary_text"
513 }
514 ],
515 "type": "reasoning",
516 "status": "completed"
517 }
518 ],
519 "parallel_tool_calls": true,
520 "previous_response_id": null,
521 "reasoning": null,
522 "temperature": null,
523 "text": {
524 "format": {
525 "type": "text"
526 }
527 },
528 "tool_choice": "auto",
529 "tools": [],
530 "top_p": null,
531 "usage": {
532 "prompt_tokens": 32,
533 "completion_tokens": 9,
534 "total_tokens": 151,
535 "prompt_tokens_details": {
536 "text_tokens": 32,
537 "audio_tokens": 0,
538 "image_tokens": 0,
539 "cached_tokens": 8
540 },
541 "completion_tokens_details": {
542 "reasoning_tokens": 110,
543 "audio_tokens": 0,
544 "accepted_prediction_tokens": 0,
545 "rejected_prediction_tokens": 0
546 },
547 "num_sources_used": 0
548 },
549 "user": null,
550 "incomplete_details": null,
551 "status": "completed",
552 "store": true
553}
554```
555
556***
557
558## GET /v1/responses/\{response\_id}/input\_items
559
560List input items for a previously generated response.
561
562### Path Parameters
563
564* `response_id` (string, required) — The response id returned by a previous create response request.
565
566### Query Parameters
567
568* `limit` (integer) — Maximum number of items to return (1-100, default 20).
569
570* `order` ("asc" | "desc") — Sort order: asc or desc. Default asc.
571
572* `after` (string) — Cursor for pagination. Returns items after this item ID.
573
574### Response Body
575
576* `data` (array\<object>, required) — The list of input items.
577
578* `first_id` (string | null) — The ID of the first item in the list.
579
580* `has_more` (boolean, required) — Whether there are more items beyond this page.
581
582* `last_id` (string | null) — The ID of the last item in the list.
583
584* `object` (string, required) — The object type, always \`list\`.
585
586\*\*Response example:\*\*
587
588```json
589{}
590```
591
592***
593
594## DELETE /v1/responses/\{response\_id}
595
596Delete a previously generated response.
597
598### Path Parameters
599
600* `response_id` (string, required) — The response id returned by a previous create response request.
601
602### Response Body
603
604* `deleted` (boolean, required) — Whether the response was successfully deleted.
605
606* `id` (string, required) — The response\_id to be deleted.
607
608* `object` (string, required) — The deleted object type, which is always \`response\`.
609
610\*\*Response example:\*\*
611
612```json
613{
614 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",
615 "object": "response",
616 "deleted": true
617}
618```
619
620***
621
622## POST /v1/responses/compact
623
624Compacts a full Responses API input window into a shorter canonical window.
625
626### Request Body
627
628* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.
629
630* `model` (string, required) — Model to use for compaction summarization (required).
631
632### Response Body
633
634* `created_at` (integer, required) — Unix timestamp (in seconds) when the compacted conversation was created.
635
636* `id` (string, required) — Unique ID for this compaction (e.g. \`cmp\_\<uuid>\`).
637
638* `model` (string, required) — Model used for the compaction summary.
639
640* `object` (string, required) — Always \`"response.compaction"\`.
641
642* `output` (array\<object | object | object | object | object>, required) — Compacted output containing a single compaction item.
643 Pass this verbatim as input to the next \`/v1/responses\` call.
644
645* `usage` (object)
646
647 * `dropped_message_count` (integer, required) — Number of messages dropped/summarized during compaction.
648
649 * `input_tokens` (integer, required) — Number of tokens in the original (pre-compaction) input.
650
651 * `input_tokens_details` (object, required) — Breakdown of input tokens for a compaction call.
652
653 * `cached_tokens` (integer, required) — Number of input tokens that were served from the prompt cache.
654
655 * `output_tokens` (integer, required) — Number of tokens in the compacted output.
656
657 * `output_tokens_details` (object, required) — Breakdown of output tokens for a compaction call.
658
659 * `reasoning_tokens` (integer, required) — Number of reasoning tokens generated during compaction.
660
661 * `total_tokens` (integer, required) — Total number of tokens used (input + output, including reasoning).
662
663### Code Examples
664
665```bash
666curl -s https://api.x.ai/v1/responses/compact \
667 -H "Content-Type: application/json" \
668 -H "Authorization: Bearer $XAI_API_KEY" \
669 -d '{
670 "model": "grok-4.6",
671 "input": [
672 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},
673 {"role": "user", "content": "What is the Higgs boson and why is it important?"},
674 {"role": "assistant", "content": "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."},
675 {"role": "user", "content": "How does the Higgs mechanism actually work?"},
676 {"role": "assistant", "content": "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless."}
677 ]
678 }'
679```
680
681```pythonOpenAISDK
682import os
683
684from openai import OpenAI
685
686client = OpenAI(
687 api_key=os.environ["XAI_API_KEY"],
688 base_url="https://api.x.ai/v1",
689)
690
691compacted = client.responses.compact(
692 model="grok-4.6",
693 input=[
694 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},
695 {"role": "user", "content": "What is the Higgs boson and why is it important?"},
696 {
697 "role": "assistant",
698 "content": (
699 "The Higgs boson is an elementary particle in the Standard Model, predicted by "
700 "Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation "
701 "of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."
702 ),
703 },
704 {"role": "user", "content": "How does the Higgs mechanism actually work?"},
705 {
706 "role": "assistant",
707 "content": (
708 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, "
709 "and particles acquire mass in proportion to how strongly they couple to it. Photons "
710 "do not couple, which is why they remain massless."
711 ),
712 },
713 ],
714)
715
716print(compacted.model_dump_json(indent=2))
717```
718
719```javascriptOpenAISDK
720import OpenAI from "openai";
721
722const client = new OpenAI({
723 apiKey: process.env.XAI_API_KEY,
724 baseURL: "https://api.x.ai/v1",
725});
726
727const compacted = await client.responses.compact({
728 model: "grok-4.6",
729 input: [
730 { role: "system", content: "You are a concise and knowledgeable science tutor." },
731 { role: "user", content: "What is the Higgs boson and why is it important?" },
732 {
733 role: "assistant",
734 content:
735 "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism.",
736 },
737 { role: "user", content: "How does the Higgs mechanism actually work?" },
738 {
739 role: "assistant",
740 content:
741 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless.",
742 },
743 ],
744});
745
746console.log(JSON.stringify(compacted, null, 2));
747```
748
749\*\*Response example:\*\*
750
751```json
752{}
753```