SpyBara
Go Premium

Documentation 2026-09-03 16:59 UTC to 2026-09-07 22:57 UTC

24 files changed +1,503 −29. View all changes and history on the product overview
2026
Wed 30 23:57 Tue 29 23:59 Mon 28 23:57 Sun 27 22:59 Sat 26 23:59 Fri 25 23:01 Thu 24 23:59 Wed 23 23:59 Tue 22 23:58 Mon 21 23:00 Sun 20 23:01 Sat 19 23:59 Fri 18 23:59 Thu 17 10:04 Wed 16 19:01 Tue 15 17:00 Mon 14 06:00 Sun 13 05:00 Fri 11 21:00 Tue 8 21:00 Mon 7 22:57 Thu 3 16:59 Wed 2 22:03
Details

985## Related985## Related

986 986 

987* [API Reference: Batch endpoints](/developers/rest-api-reference/inference/batches#create-a-new-batch)987* [API Reference: Batch endpoints](/developers/rest-api-reference/inference/batches#create-a-new-batch)

988* [gRPC Reference: Batch management](/developers/grpc-api-reference#batch-management)988* [gRPC Reference: Batch Management](/developers/grpc-api-reference/batches)

989* [Pricing — Batch API Pricing](/developers/pricing#batch-api-pricing)989* [Pricing — Batch API Pricing](/developers/pricing#batch-api-pricing)

990* [xAI Python SDK](https://github.com/xai-org/xai-sdk-python)990* [xAI Python SDK](https://github.com/xai-org/xai-sdk-python)

Details

267 267 

268* [Generate Text — Responses API](/developers/model-capabilities/text/generate-text) — the primary endpoint that compaction feeds into.268* [Generate Text — Responses API](/developers/model-capabilities/text/generate-text) — the primary endpoint that compaction feeds into.

269* [Prompt Caching](/developers/advanced-api-usage/prompt-caching) — a complementary cost-reduction lever for unchanged prompt prefixes.269* [Prompt Caching](/developers/advanced-api-usage/prompt-caching) — a complementary cost-reduction lever for unchanged prompt prefixes.

270* [Chat API Reference](/developers/rest-api-reference/inference/chat) — full request/response schema for the Compaction API.270* [Responses API Reference](/developers/rest-api-reference/inference/responses#compact-a-conversation) — full request/response schema for the Compaction API.

Details

203}203}

204```204```

205 205 

206For more details, refer to [Chat completions](/developers/rest-api-reference/inference/chat#chat-completions) and [Get deferred chat completions](/developers/rest-api-reference/inference/chat#get-deferred-chat-completions) in our REST API Reference.206For more details, refer to [Chat completions](/developers/rest-api-reference/inference/chat-completions#chat-completions) and [Get deferred chat completions](/developers/rest-api-reference/inference/chat-completions#get-deferred-chat-completions) in our REST API Reference.

Details

24 24 

25After the WebSocket upgrade succeeds, every turn is initiated by the client sending a25After the WebSocket upgrade succeeds, every turn is initiated by the client sending a

26`response.create` message. The body is the same shape as the26`response.create` message. The body is the same shape as the

27[Responses create body](/developers/rest-api-reference/inference/chat#create-new-response), minus27[Responses create body](/developers/rest-api-reference/inference/responses#create-new-response), minus

28transport-only fields like `stream` and `background` (responses are always streamed back as28transport-only fields like `stream` and `background` (responses are always streamed back as

29events on the socket).29events on the socket).

30 30 


236 236 

237* [Streaming](/developers/model-capabilities/text/streaming)237* [Streaming](/developers/model-capabilities/text/streaming)

238* [Function Calling](/developers/tools/function-calling)238* [Function Calling](/developers/tools/function-calling)

239* [Responses API Reference](/developers/rest-api-reference/inference/chat#create-new-response)239* [Responses API Reference](/developers/rest-api-reference/inference/responses#create-new-response)

grok-4-6.md +2 −2

Details

75|----------|-------|75|----------|-------|

76| Model name | `grok-4.6` |76| Model name | `grok-4.6` |

77| Context window | 500,000 tokens |77| Context window | 500,000 tokens |

78| Knowledge cutoff | January 2026 |78| Knowledge cutoff | February 1, 2026 |

79| Modalities | Text and image input; text output |79| Modalities | Text and image input; text output |

80| Output limit | No text output limit |80| Output limit | No text output limit |

81| Input price | $2.00 / 1M tokens |81| Input price | $2.00 / 1M tokens |

82| Output price | $6.00 / 1M tokens |82| Output price | $6.00 / 1M tokens |

83| Reasoning | Low, medium, high (default), or xhigh |83| Reasoning | Low, medium, high (default), or xhigh |

84| APIs | [Responses API](/developers/rest-api-reference/inference/chat#create-new-response), [Chat Completions](/developers/rest-api-reference/inference/chat#chat-completions) |84| APIs | [Responses API](/developers/rest-api-reference/inference/responses#create-new-response), [Chat Completions](/developers/rest-api-reference/inference/chat-completions#chat-completions) |

85| Tools | [Function calling](/developers/tools/function-calling), [web search](/developers/tools/web-search), [X search](/developers/tools/x-search), [code execution](/developers/tools/code-execution) |85| Tools | [Function calling](/developers/tools/function-calling), [web search](/developers/tools/web-search), [X search](/developers/tools/x-search), [code execution](/developers/tools/code-execution) |

86 86 

87Rate limits and live pricing for your team are on the [model detail page](/developers/models/grok-4.6) and [Pricing](/developers/pricing).87Rate limits and live pricing for your team are on the [model detail page](/developers/models/grok-4.6) and [Pricing](/developers/pricing).

Details

1# gRPC API Reference1#### gRPC API

2 2 

3The xAI gRPC API is a robust, high-performance gRPC interface designed for seamless integration into existing systems.3# Overview

4 4 

5The base url for all services is at `api.x.ai`. For all services, you have to authenticate with the header `Authorization: Bearer <your xAI API key>`.5The xAI gRPC API exposes the same models and services as the REST API over gRPC. The base URL for all services is `api.x.ai`, and every call must carry the header `Authorization: Bearer <your xAI API key>`.

6 6 

7Visit [xAI API Protobuf Definitions](https://github.com/xai-org/xai-proto) to view and download our protobuf definitions.7The protobuf definitions are published in [xai-org/xai-proto](https://github.com/xai-org/xai-proto). The [xAI Python SDK](https://github.com/xai-org/xai-sdk-python) (`xai-sdk`) uses gRPC natively; install it with `pip install xai-sdk`.

8 

9The [xAI Python SDK](https://github.com/xai-org/xai-sdk-python) (`xai-sdk`) uses gRPC natively. Install with `pip install xai-sdk`.

10 8 

11## Using buf curl9## Using buf curl

12 10 


17cd xai-proto15cd xai-proto

18```16```

19 17 

20All `buf curl` examples below assume you run from inside the cloned `xai-proto` directory.18All `buf curl` examples in this reference assume you run from inside the cloned `xai-proto` directory.

19 

20## Services

21 21 

22***22* [Chat](/developers/grpc-api-reference/chat) — `xai_api.Chat`

23* [Image](/developers/grpc-api-reference/image) — `xai_api.Image`

24* [Video](/developers/grpc-api-reference/video) — `xai_api.Video`

25* [Batch Management](/developers/grpc-api-reference/batches) — `xai_api.BatchMgmt`

26* [Models](/developers/grpc-api-reference/models) — `xai_api.Models`

27* [Auth](/developers/grpc-api-reference/auth) — `xai_api.Auth`

28* [Tokenize](/developers/grpc-api-reference/tokenize) — `xai_api.Tokenize`

29* [Raw Sampling](/developers/grpc-api-reference/sample) — `xai_api.Sample`

rate-limits.md +2 −2

Details

40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |40| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |41| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |42| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

43| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

44| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |43| grok-imagine-image-2.0 | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

45| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |44| grok-imagine-image | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

46| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |45| grok-imagine-image-quality | T0: 6, T1: 12, T2: 25, T3: 50, T4: 100 | — |

47| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |46| grok-imagine-video-1.5 | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

47| grok-imagine-video | T0: 10, T1: 20, T2: 39, T3: 79, T4: 158 | — |

48 48 

49### What counts toward TPM49### What counts toward TPM

50 50 

Details

9* [Upload](/developers/rest-api-reference/files/upload)9* [Upload](/developers/rest-api-reference/files/upload)

10* [Manage](/developers/rest-api-reference/files/manage)10* [Manage](/developers/rest-api-reference/files/manage)

11* [Download](/developers/rest-api-reference/files/download)11* [Download](/developers/rest-api-reference/files/download)

12* [Public URLs](/developers/rest-api-reference/files/public-urls)

Details

1#### Files API

2 

3# Public URLs

4 

5See the [Public URLs guide](/developers/files/public-urls) for expiry behaviour, idempotency, and end-to-end examples.

6 

7***

8 

9## POST /v1/files/\{file\_id}/public-url

10 

11Create a permanent, unauthenticated public URL for an existing file. The

12underlying file is unaffected and can still be fetched through the

13authenticated content endpoint. Use this when you want to share a stored

14asset (image, video, PDF) outside your API-keyed environment. Public URLs

15can be revoked at any time via \`POST /v1/files/\{file\_id}/public-url/revoke\`.

16 

17### Path Parameters

18 

19* `file_id` (string, required) — The file's \`id\`.

20 

21### Request Body

22 

23* `expires_after` (integer | null) — Seconds from now until the public URL expires. Must be between \`3600\` (1

24 hour) and \`2592000\` (30 days). Omit to inherit the file's expiry (if it

25 has one) or to make the URL valid indefinitely.

26 

27### Response Body

28 

29* `expires_at` (integer | null) — Unix timestamp (seconds) when the public URL expires. Present when

30 the public URL has an expiry, either from an explicit \`expires\_after\`

31 in the request or inherited from the file's TTL. Absent when the

32 public URL is valid indefinitely.

33 

34* `public_url` (string, required) — The full public URL.

35 

36\*\*Response example:\*\*

37 

38```json

39{

40 "public_url": "https://files-cdn.x.ai/ZsqeMtdcSYWPPHTQdxXDKQ/file_a128090d-f0c9-4873-bd84-e499777e7417.png",

41 "expires_at": 1755600000

42}

43```

44 

45***

46 

47## POST /v1/files/\{file\_id}/public-url/revoke

48 

49Revoke the active public URL for a file. The underlying file remains

50available through the authenticated content endpoint. Revoke is idempotent

51— calling it on a file without an active public URL returns

52\`revoked: false\` without an error.

53 

54### Path Parameters

55 

56* `file_id` (string, required) — The file's \`id\`.

57 

58### Response Body

59 

60* `id` (string, required) — The file ID whose public URL was revoked.

61 

62* `public_url` (string | null) — The full public URL that was revoked. Only present when \`revoked\` is \`true\`.

63 

64* `revoked` (boolean, required) — Whether a public URL was actually revoked. \`false\` if the file had no

65 active public URL (no-op).

66 

67\*\*Response example:\*\*

68 

69```json

70{

71 "id": "file_a128090d-f0c9-4873-bd84-e499777e7417",

72 "revoked": true,

73 "public_url": "https://files-cdn.x.ai/ZsqeMtdcSYWPPHTQdxXDKQ/file_a128090d-f0c9-4873-bd84-e499777e7417.png"

74}

75```

Details

1#### Inference API1#### API Reference

2 2 

3# Inference REST API Overview3# Overview

4 4 

5The xAI Inference REST API is a robust, high-performance RESTful interface designed for seamless integration into existing systems.5The xAI REST API is compatible with the OpenAI REST API. This reference is generated from the OpenAPI specification and organised by resource.

6It offers advanced AI capabilities with full compatibility with the OpenAI REST API.

7 6 

8The base for all routes is at `https://api.x.ai`. For all routes, you have to authenticate with the header `Authorization: Bearer <your xAI API key>`.7## Base URLs and authentication

9 8 

10* [Chat](/developers/rest-api-reference/inference/chat)9| API | Base URL | Authenticate with |

10| --- | --- | --- |

11| Inference (responses, chat completions, embeddings, images, videos, voice, files, batches, models) | `https://api.x.ai` | `Authorization: Bearer <xAI API key>` |

12| Collections management | `https://management-api.x.ai` | `Authorization: Bearer <xAI Management API key>` |

13| Collections search | `https://api.x.ai` | `Authorization: Bearer <xAI API key>` |

14| Management (API keys, teams, billing, audit) | `https://management-api.x.ai` | `Authorization: Bearer <xAI Management API key>` |

15 

16API keys are created on the [API Keys page](https://console.x.ai/team/default/api-keys?utm_source=docs\&utm_medium=referral\&utm_campaign=developers-rest-api-reference-inference\&utm_content=api-keys) of the xAI Console. Management keys are created on the [Management Keys page](https://console.x.ai/team/default/management-keys?utm_source=docs\&utm_medium=referral\&utm_campaign=developers-rest-api-reference-inference\&utm_content=management-keys); see [Using Management API](/developers/management-api-guide).

17 

18## Inference API

19 

20* [Responses](/developers/rest-api-reference/inference/responses)

21* [Chat Completions](/developers/rest-api-reference/inference/chat-completions)

22* [Embeddings](/developers/rest-api-reference/inference/embeddings)

11* [Images](/developers/rest-api-reference/inference/images)23* [Images](/developers/rest-api-reference/inference/images)

12* [Videos](/developers/rest-api-reference/inference/videos)24* [Videos](/developers/rest-api-reference/inference/videos)

13* [Voice](/developers/rest-api-reference/inference/voice)25* [Voice](/developers/rest-api-reference/inference/voice)

14* [Models](/developers/rest-api-reference/inference/models)

15* [Files](/developers/rest-api-reference/files)26* [Files](/developers/rest-api-reference/files)

16* [Batches](/developers/rest-api-reference/inference/batches)27* [Batches](/developers/rest-api-reference/inference/batches)

17* [Other](/developers/rest-api-reference/inference/other)28* [Models](/developers/rest-api-reference/inference/models)

29* [Account](/developers/rest-api-reference/inference/other)

18* [Legacy & Deprecated](/developers/rest-api-reference/inference/legacy)30* [Legacy & Deprecated](/developers/rest-api-reference/inference/legacy)

31 

32## Management-key APIs

33 

34* [Collections API](/developers/rest-api-reference/collections)

35* [Management API](/developers/rest-api-reference/management)

36 

37## Other protocols

38 

39* [gRPC API](/developers/grpc-api-reference)

40 

41For status codes and their likely causes, see [Debugging Errors](/developers/debugging).

Details

303 303 

304 * `presence_penalty` (number | null) — (Not supported by \`grok-3\` and reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.304 * `presence_penalty` (number | null) — (Not supported by \`grok-3\` and reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

305 305 

306 * `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Not supported by \`grok-4\` and will result in error if used with \`grok-4\`. Possible values are \`low\` (uses fewer reasoning tokens) and \`high\` (uses more reasoning tokens).306 * `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.

307 307 

308 * `response_format` (object | object | object)308 * `response_format` (object | object | object)

309 309 

Details

1#### Inference API

2 

3# Chat Completions

4 

5The Chat Completions API is the stateless, OpenAI-compatible predecessor of the [Responses API](/developers/rest-api-reference/inference/responses). New integrations should use Responses; see [Migrating from Chat Completions](/developers/model-capabilities/text/comparison).

6 

7***

8 

9## POST /v1/chat/completions

10 

11Create a chat response from text/image chat prompts. This is the endpoint for making requests to chat and image understanding models.

12 

13### Request Body

14 

15* `deferred` (boolean | null) — If set to \`true\`, the request returns a \`request\_id\`. You can then get the deferred response by GET \`/v1/chat/deferred-completion/\{request\_id}\`.

16 

17* `frequency_penalty` (number | null) — (Not supported by reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.

18 

19* `logit_bias` (object | null) — (Unsupported) A JSON object that maps tokens (specified by their token ID in the tokenizer) to an associated bias value from -100 to 100. Mathematically, the bias is added to the logits generated by the model prior to sampling. The exact effect will vary per model, but values between -1 and 1 should decrease or increase likelihood of selection; values like -100 or 100 should result in a ban or exclusive selection of the relevant token.

20 

21* `logprobs` (boolean | null) — Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

22 

23* `max_completion_tokens` (integer | null) — An upper bound for the number of tokens that can be generated for a completion, only applies to visible output tokens (i.e. does not apply to tokens used for reasoning or function calls). Defaults to 128,000 when unset; set a larger value to allow longer generations.

24 

25* `max_tokens` (integer | null) — \\\[DEPRECATED\\] The maximum number of tokens that can be generated in the chat completion. Deprecated in favor of \`max\_completion\_tokens\`.

26 

27* `messages` (array\<object | object | object | object | object>) — A list of messages that make up the chat conversation. Different models support different message types, such as image and text.

28 

29* `model` (string) — Model name for the model to use. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

30 

31* `n` (integer | null) — How many chat completion choices to generate for each input message. Note that you will be charged based on the number of generated tokens across all of the choices. Keep n as 1 to minimize costs.

32 

33* `parallel_tool_calls` (boolean | null) — If set to false, the model can perform maximum one tool call.

34 

35* `presence_penalty` (number | null) — (Not supported by \`grok-3\` and reasoning models) Number between -2.0 and 2.0. Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

36 

37* `prompt_cache_key` (string | null) — A stable cache key for best-effort sticky routing / prompt-cache hits

38 across requests sharing a prompt prefix. Plumbed to \`x-grok-conv-id\`,

39 same as on \`/v1/responses\`.

40 

41* `reasoning_effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.

42 

43* `response_format` (object | object | object)

44 

45* `search_parameters` (object)

46 

47 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See

48 \<https://en.wikipedia.org/wiki/ISO\_8601>.

49 

50 * `max_search_results` (integer | null) — Maximum number of search results to use.

51 

52 * `mode` (string | null) — Choose the mode to query realtime data:

53 \* \`off\`: no search performed and no external will be considered.

54 \* \`on\` (default): the model will search in every sources for relevant data.

55 \* \`auto\`: the model choose whether to search data or not and where to search the data.

56 

57 * `return_citations` (boolean | null) — Whether to return citations in the response or not.

58 

59 * `sources` (array | null) — List of sources to search in. If no sources specified, the model will look over the web and X by default.

60 

61 * `to_date` (string | null) — Date up to which to consider the results in ISO-8601 YYYY-MM-DD. See

62 \<https://en.wikipedia.org/wiki/ISO\_8601>.

63 

64* `seed` (integer | null) — If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same \`seed\` and parameters should return the same result. Determinism is not guaranteed, and you should refer to the \`system\_fingerprint\` response parameter to monitor changes in the backend.

65 

66* `service_tier` ("default" | "priority")

67 

68* `stop` (array | null) — (Not supported by reasoning models) Up to 4 sequences where the API will stop generating further tokens.

69 

70* `stream` (boolean | null) — If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a \`data: \[DONE]\` message.

71 

72* `stream_options` (object)

73 

74 * `include_usage` (boolean, required) — Set an additional chunk to be streamed before the \`data: \[DONE]\` message. The other chunks will return \`null\` in \`usage\` field.

75 

76* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

77 

78* `tool_choice` (string | object)

79 

80* `tools` (array | null) — A list of tools the model may call in JSON-schema. Currently, only functions are supported as a tool. Use this to provide a list of functions the model may generate JSON inputs for. A max of 128 functions are supported.

81 

82* `top_logprobs` (integer | null) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position, each with an associated log probability. logprobs must be set to true if this parameter is used. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

83 

84* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

85 

86* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

87 

88* `web_search_options` (object)

89 

90 * `filters` (object) — Only included for compatibility.

91 

92 * `search_context_size` (string | null) — This field included for compatibility reason with OpenAI's API. It is mapped to \`max\_search\`.

93 

94 * `user_location` (object) — Only included for compatibility.

95 

96### Response Body

97 

98* `choices` (array\<object>, required) — A list of response choices from the model. The length corresponds to the \`n\` in request body (default to 1).

99 

100 * `finish_reason` (string | null) — Finish reason. \`"stop"\` means the inference has reached a model-defined or user-supplied stop sequence in \`stop\`. \`"length"\` means the inference result has reached models' maximum allowed token length or user defined value in \`max\_tokens\`. \`"end\_turn"\` or \`null\` in streaming mode when the chunk is not the last.

101 

102 * `index` (integer, required) — Index of the choice within the response choices, starting from 0.

103 

104 * `logprobs` (object)

105 

106 * `content` (array | null) — An array the log probabilities of each output token returned.

107 

108 * `message` (object, required)

109 

110 * `content` (string | null) — The content of the message.

111 

112 * `reasoning_content` (string | null) — The reasoning trace generated by the model.

113 

114 * `refusal` (string | null) — The reason given by model if the model is unable to generate a response. null if model is able to generate.

115 

116 * `role` (string, required) — The role that the message belongs to, the response from model is always \`"assistant"\`.

117 

118 * `tool_calls` (array | null) — A list of tool calls asked by model for user to perform.

119 

120* `citations` (array | null) — List of all the external pages used by the model to answer.

121 

122* `created` (integer, required) — The chat completion creation time in Unix timestamp.

123 

124* `id` (string, required) — A unique ID for the chat response.

125 

126* `model` (string, required) — Model ID used to create chat completion.

127 

128* `object` (string, required) — The object type, which is always \`"chat.completion"\`.

129 

130* `output_files` (array | null) — Files generated during the response (e.g., by the code execution tool).

131 Only populated when \`code\_execution\_files\_output\` is included.

132 

133* `service_tier` ("default" | "priority", required) — Processing tier for a request. Determines scheduling priority and billing.

134 

135* `system_fingerprint` (string | null) — System fingerprint, used to indicate xAI system configuration changes.

136 

137* `usage` (object)

138 

139 * `completion_tokens` (integer, required) — Total completion token used.

140 

141 * `completion_tokens_details` (object, required) — Details of completion usage.

142 

143 * `accepted_prediction_tokens` (integer, required) — The number of tokens in the prediction that appeared in the completion.

144 

145 * `audio_tokens` (integer, required) — Audio input tokens generated by the model.

146 

147 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

148 

149 * `rejected_prediction_tokens` (integer, required) — The number of tokens in the prediction that did not appear in the completion.

150 

151 * `cost_in_usd_ticks` (integer, required) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

152 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

153 which means there is 10'000'000'000 ticks in one \*dollar\*.

154 

155 * `num_sources_used` (integer, required) — Number of individual live search source used.

156 

157 * `prompt_tokens` (integer, required) — Total prompt token used.

158 

159 * `prompt_tokens_details` (object, required) — Details of prompt usage.

160 

161 * `audio_tokens` (integer, required) — Audio prompt token used.

162 

163 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

164 

165 * `image_tokens` (integer, required) — Image prompt token used.

166 

167 * `text_tokens` (integer, required) — Total text prompt token used (cached + non-cached text tokens).

168 

169 * `total_tokens` (integer, required) — Total token used, the sum of prompt token and completion token amount.

170 

171\*\*Request example:\*\*

172 

173```json

174{

175 "messages": [

176 {

177 "role": "system",

178 "content": "You are a helpful assistant that can answer questions and help with tasks."

179 },

180 {

181 "role": "user",

182 "content": "What is 101*3?"

183 }

184 ],

185 "model": "latest"

186}

187```

188 

189\*\*Response example:\*\*

190 

191```json

192{

193 "id": "a3d1008e-4544-40d4-d075-11527e794e4a",

194 "object": "chat.completion",

195 "created": 1752854522,

196 "model": "latest",

197 "choices": [

198 {

199 "index": 0,

200 "message": {

201 "role": "assistant",

202 "content": "101 multiplied by 3 is 303.",

203 "refusal": null

204 },

205 "finish_reason": "stop"

206 }

207 ],

208 "usage": {

209 "prompt_tokens": 32,

210 "completion_tokens": 9,

211 "total_tokens": 135,

212 "prompt_tokens_details": {

213 "text_tokens": 32,

214 "audio_tokens": 0,

215 "image_tokens": 0,

216 "cached_tokens": 6

217 },

218 "completion_tokens_details": {

219 "reasoning_tokens": 94,

220 "audio_tokens": 0,

221 "accepted_prediction_tokens": 0,

222 "rejected_prediction_tokens": 0

223 },

224 "num_sources_used": 0

225 },

226 "system_fingerprint": "fp_3a7881249c"

227}

228```

229 

230***

231 

232## GET /v1/chat/deferred-completion/\{request\_id}

233 

234Tries to fetch a result for a previously-started deferred completion. Returns \`200 Success\` with the response body, if the request has been completed. Returns \`202 Accepted\` when the request is pending processing.

235 

236### Path Parameters

237 

238* `request_id` (string, required) — The deferred request id returned by a previous deferred chat request.

239 

240### Response Body

241 

242* `choices` (array\<object>, required) — A list of response choices from the model. The length corresponds to the \`n\` in request body (default to 1).

243 

244 * `finish_reason` (string | null) — Finish reason. \`"stop"\` means the inference has reached a model-defined or user-supplied stop sequence in \`stop\`. \`"length"\` means the inference result has reached models' maximum allowed token length or user defined value in \`max\_tokens\`. \`"end\_turn"\` or \`null\` in streaming mode when the chunk is not the last.

245 

246 * `index` (integer, required) — Index of the choice within the response choices, starting from 0.

247 

248 * `logprobs` (object)

249 

250 * `content` (array | null) — An array the log probabilities of each output token returned.

251 

252 * `message` (object, required)

253 

254 * `content` (string | null) — The content of the message.

255 

256 * `reasoning_content` (string | null) — The reasoning trace generated by the model.

257 

258 * `refusal` (string | null) — The reason given by model if the model is unable to generate a response. null if model is able to generate.

259 

260 * `role` (string, required) — The role that the message belongs to, the response from model is always \`"assistant"\`.

261 

262 * `tool_calls` (array | null) — A list of tool calls asked by model for user to perform.

263 

264* `citations` (array | null) — List of all the external pages used by the model to answer.

265 

266* `created` (integer, required) — The chat completion creation time in Unix timestamp.

267 

268* `id` (string, required) — A unique ID for the chat response.

269 

270* `model` (string, required) — Model ID used to create chat completion.

271 

272* `object` (string, required) — The object type, which is always \`"chat.completion"\`.

273 

274* `output_files` (array | null) — Files generated during the response (e.g., by the code execution tool).

275 Only populated when \`code\_execution\_files\_output\` is included.

276 

277* `service_tier` ("default" | "priority", required) — Processing tier for a request. Determines scheduling priority and billing.

278 

279* `system_fingerprint` (string | null) — System fingerprint, used to indicate xAI system configuration changes.

280 

281* `usage` (object)

282 

283 * `completion_tokens` (integer, required) — Total completion token used.

284 

285 * `completion_tokens_details` (object, required) — Details of completion usage.

286 

287 * `accepted_prediction_tokens` (integer, required) — The number of tokens in the prediction that appeared in the completion.

288 

289 * `audio_tokens` (integer, required) — Audio input tokens generated by the model.

290 

291 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

292 

293 * `rejected_prediction_tokens` (integer, required) — The number of tokens in the prediction that did not appear in the completion.

294 

295 * `cost_in_usd_ticks` (integer, required) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

296 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

297 which means there is 10'000'000'000 ticks in one \*dollar\*.

298 

299 * `num_sources_used` (integer, required) — Number of individual live search source used.

300 

301 * `prompt_tokens` (integer, required) — Total prompt token used.

302 

303 * `prompt_tokens_details` (object, required) — Details of prompt usage.

304 

305 * `audio_tokens` (integer, required) — Audio prompt token used.

306 

307 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

308 

309 * `image_tokens` (integer, required) — Image prompt token used.

310 

311 * `text_tokens` (integer, required) — Total text prompt token used (cached + non-cached text tokens).

312 

313 * `total_tokens` (integer, required) — Total token used, the sum of prompt token and completion token amount.

314 

315\*\*Response example:\*\*

316 

317```json

318{

319 "id": "335b92e4-afa5-48e7-b99c-b9a4eabc1c8e",

320 "object": "chat.completion",

321 "created": 1743770624,

322 "model": "latest",

323 "choices": [

324 {

325 "index": 0,

326 "message": {

327 "role": "assistant",

328 "content": "101 multiplied by 3 is 303.",

329 "refusal": null

330 },

331 "finish_reason": "stop"

332 }

333 ],

334 "usage": {

335 "prompt_tokens": 31,

336 "completion_tokens": 11,

337 "total_tokens": 42,

338 "prompt_tokens_details": {

339 "text_tokens": 31,

340 "audio_tokens": 0,

341 "image_tokens": 0,

342 "cached_tokens": 0

343 },

344 "completion_tokens_details": {

345 "reasoning_tokens": 0,

346 "audio_tokens": 0,

347 "accepted_prediction_tokens": 0,

348 "rejected_prediction_tokens": 0

349 }

350 },

351 "system_fingerprint": "fp_156d35dcaa"

352}

353```

Details

1#### Inference API

2 

3# Embeddings

4 

5***

6 

7## POST /v1/embeddings

8 

9Create an embedding vector representation corresponding to the input text. This is the endpoint for making requests to embedding models.

10 

11### Request Body

12 

13* `dimensions` (integer | null) — The number of dimensions the resulting output embeddings should have.

14 

15* `encoding_format` (string | null) — The format to return the embeddings in. Can be either \`float\` or \`base64\`.

16 

17* `input` (object | object | object | object)

18 

19 * `String` (string, required) — A strings to be embedded. For best performance, prepend "query: " in front of query content and prepend "passage: " in front of passage/text

20 

21 * `StringArray` (array\<string>, required) — An array of strings to be embedded

22 

23 * `Ints` (array\<integer>, required) — A token in integer to be embedded

24 

25 * `IntsArray` (array\<array\<integer>>, required) — An array of tokens in integers to be embedded

26 

27* `model` (string) — ID of the model to use.

28 

29* `preview` (boolean | null) — Flag to use the new format of the API.

30 

31* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

32 

33### Response Body

34 

35* `data` (array\<object>, required) — A list of embedding objects.

36 

37 * `embedding` (string | array\<number>, required)

38 

39 * `index` (integer, required) — Index of the embedding object in the data list.

40 

41 * `object` (string, required) — The object type, which is always \`"embedding"\`.

42 

43* `model` (string, required) — Model ID used to create embedding.

44 

45* `object` (string, required) — The object type of \`data\` field, which is always \`"list"\`.

46 

47* `usage` (object)

48 

49 * `prompt_tokens` (integer, required) — Prompt token used.

50 

51 * `total_tokens` (integer, required) — Total token used.

52 

53\*\*Request example:\*\*

54 

55```json

56"{\n \"input\": [\"This is an example content to embed...\"],\n \"model\": \"v1\",\n \"encoding_format\": \"float\"\n }"

57```

58 

59\*\*Response example:\*\*

60 

61```json

62{

63 "object": "list",

64 "model": "v1",

65 "data": [

66 {

67 "index": 0,

68 "embedding": [

69 0.01567895,

70 0.063257694,

71 0.045925662

72 ],

73 "object": "embedding"

74 }

75 ],

76 "usage": {

77 "prompt_tokens": 1,

78 "total_tokens": 1

79 }

80}

81```

82 

83***

84 

85## GET /v1/embedding-models

86 

87List all embedding models available to the authenticating API key with full information. Additional information compared to /v1/models includes modalities, fingerprint and alias(es).

88 

89### Response Body

90 

91* `models` (array\<object>, required) — Array of available embedding models.

92 

93 * `aliases` (array\<string>, required) — Alias ID(s) of the model that user can use in a request's model field.

94 

95 * `created` (integer, required) — Model creation time in Unix timestamp.

96 

97 * `fingerprint` (string, required) — Fingerprint of the xAI system configuration hosting the model.

98 

99 * `id` (string, required) — Model ID. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

100 

101 * `input_modalities` (array\<string>, required) — The input modalities supported by the model.

102 

103 * `object` (string, required) — Object type, should be model.

104 

105 * `output_modalities` (array\<string>, required) — The output modalities supported by the model.

106 

107 * `owned_by` (string, required) — Owner of the model.

108 

109 * `prompt_image_token_price` (integer, required) — Price of the prompt image token in USD cents per million token.

110 

111 * `prompt_text_token_price` (integer, required) — Price of the prompt text token in USD cents per million token.

112 

113 * `version` (string, required) — Version of the model.

114 

115\*\*Response example:\*\*

116 

117```json

118{

119 "models": [

120 {

121 "id": "v1",

122 "fingerprint": "fp_df37966059",

123 "created": 1725148800,

124 "object": "model",

125 "owned_by": "xai",

126 "version": "0.1.0",

127 "input_modalities": [

128 "text"

129 ],

130 "prompt_text_token_price": 100,

131 "prompt_image_token_price": 0,

132 "aliases": []

133 }

134 ]

135}

136```

137 

138***

139 

140## GET /v1/embedding-models/\{model\_id}

141 

142Get full information about an embedding model with its model\_id.

143 

144### Path Parameters

145 

146* `model_id` (string, required) — ID of the model to get.

147 

148### Response Body

149 

150* `aliases` (array\<string>, required) — Alias ID(s) of the model that user can use in a request's model field.

151 

152* `created` (integer, required) — Model creation time in Unix timestamp.

153 

154* `fingerprint` (string, required) — Fingerprint of the xAI system configuration hosting the model.

155 

156* `id` (string, required) — Model ID. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

157 

158* `input_modalities` (array\<string>, required) — The input modalities supported by the model.

159 

160* `object` (string, required) — Object type, should be model.

161 

162* `output_modalities` (array\<string>, required) — The output modalities supported by the model.

163 

164* `owned_by` (string, required) — Owner of the model.

165 

166* `prompt_image_token_price` (integer, required) — Price of the prompt image token in USD cents per million token.

167 

168* `prompt_text_token_price` (integer, required) — Price of the prompt text token in USD cents per million token.

169 

170* `version` (string, required) — Version of the model.

171 

172\*\*Response example:\*\*

173 

174```json

175{

176 "id": "v1",

177 "created": 1725148800,

178 "object": "model",

179 "owned_by": "xai",

180 "version": "0.1.0",

181 "input_modalities": [

182 "text"

183 ],

184 "prompt_text_token_price": 10,

185 "prompt_image_token_price": 0,

186 "aliases": []

187}

188```

Details

150 150 

151> [!WARNING]151> [!WARNING]

152>152>

153> **Deprecated**: The Anthropic SDK compatibility is fully deprecated. Please migrate to the [Responses API](/developers/rest-api-reference/inference/chat#create-new-response) or [gRPC](/developers/grpc-api-reference).153> **Deprecated**: The Anthropic SDK compatibility is fully deprecated. Please migrate to the [Responses API](/developers/rest-api-reference/inference/responses#create-new-response) or [gRPC](/developers/grpc-api-reference).

154 154 

155## POST /v1/messages155## POST /v1/messages

156 156 


258 258 

259> [!WARNING]259> [!WARNING]

260>260>

261> **Deprecated**: The Anthropic SDK compatibility is fully deprecated. Please migrate to the [Responses API](/developers/rest-api-reference/inference/chat#create-new-response) or [gRPC](/developers/grpc-api-reference).261> **Deprecated**: The Anthropic SDK compatibility is fully deprecated. Please migrate to the [Responses API](/developers/rest-api-reference/inference/responses#create-new-response) or [gRPC](/developers/grpc-api-reference).

262 262 

263## POST /v1/complete263## POST /v1/complete

264 264 

Details

1#### Inference API1#### Inference API

2 2 

3# Other3# Account

4 

5***

6 

7## GET /v1/me

8 

9Get information about the currently authenticated caller.

10Works with both API keys and OAuth tokens. Returns identity, team, and ZDR status.

11 

12### Response Body

13 

14* `api_key` (object)

15 

16 * `api_key_id` (string, required) — The API key ID.

17 

18 * `blocked` (boolean, required) — Whether the API key is blocked.

19 

20 * `disabled` (boolean, required) — Whether the API key is disabled.

21 

22 * `redacted_api_key` (string, required) — The redacted API key.

23 

24* `oauth` (object)

25 

26 * `client_id` (string, required) — The OAuth client\_id of the application.

27 

28* `team_blocked` (boolean, required) — Whether the team is blocked from making API requests.

29 

30* `team_id` (string, required) — Team ID associated with the credentials.

31 

32* `user_id` (string, required) — User ID associated with the credentials.

33 

34* `zdr_status` ("no\_zdr" | "zdr" | "pii\_scrubbing", required) — Zero Data Retention status for a team.

35 

36\*\*Response example:\*\*

37 

38```json

39{

40 "user_id": "59fbe5f2-040b-46d5-8325-868bb8f23eb2",

41 "team_id": "5ea6f6bd-7815-4b8a-9135-28b2d7ba6722",

42 "zdr_status": "no_zdr",

43 "team_blocked": false,

44 "api_key": {

45 "redacted_api_key": "xai-...b14o",

46 "api_key_id": "ae1e1841-4326-4b36-a8a9-8a1a7237db11",

47 "blocked": false,

48 "disabled": false

49 }

50}

51```

52 

53***

4 54 

5## GET /v1/api-key55## GET /v1/api-key

6 56 

Details

1#### Inference API

2 

3# Responses

4 

5The Responses API is the primary interface for text generation, reasoning, and tool use. See the [Text Generation guide](/developers/model-capabilities/text/generate-text) for usage.

6 

7***

8 

9## POST /v1/responses

10 

11Generates a response based on text or image prompts. The response ID can be used to retrieve the response later or to continue the conversation without repeating prior context. New responses will be stored for 30 days and then permanently deleted.

12 

13### Request Body

14 

15* `background` (boolean | null) — (Unsupported) Whether to process the response asynchronously in the background.

16 

17* `context_management` (array | null) — Optional context-management directives (e.g. compaction). Parsed but not yet executed.

18 

19* `include` (array | null) — What additional output data to include in the response. Supported values include

20 \`reasoning.encrypted\_content\` (encrypted reasoning tokens) and tool-output options.

21 OpenAI's \`message.output\_text.logprobs\` is accepted for compatibility but silently ignored.

22 

23* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.

24 

25* `instructions` (string | null) — An alternate way to specify the system prompt. Note that this cannot be used alongside \`previous\_response\_id\`, where the system prompt of the previous message will be used.

26 

27* `logprobs` (boolean | null) — Whether to return log probabilities of the output tokens or not. If true, returns the log probabilities of each output token returned in the content of message. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

28 

29* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens. Defaults to 128,000 when unset; set a larger value to allow longer generations.

30 

31* `max_turns` (integer | null) — Maximum number of agentic tool calling turns allowed for this request.

32 If not set, defaults to the server's global cap.

33 This parameter will be ignored for any non-agentic requests.

34 

35* `metadata` (object) — Not supported. Only maintained for compatibility reasons.

36 

37* `min_p` (number | null) — Min-p sampling: tokens whose probability is below \`min\_p\` times the probability of the most likely token are excluded from sampling. Disabled when unset.

38 

39* `model` (string) — Model name for the model to use. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

40 

41* `parallel_tool_calls` (boolean | null) — Whether to allow the model to run parallel tool calls.

42 

43* `previous_response_id` (string | null) — The ID of the previous response from the model.

44 

45* `prompt_cache_key` (string | null) — Plumbed to x-grok-conv-id for Open Responses compatibility, used for routing.

46 

47* `reasoning` (object)

48 

49 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.

50 

51 * `generate_summary` (string | null) — Only included for compatibility.

52 

53 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.

54 

55* `reasoning_effort` (string | null) — reasoning\_effort alternative to reasoning configuration. This is a non-standard field meant to ease user experience. We only look at this if the reasoning field is unset.

56 

57* `search_parameters` (object)

58 

59 * `from_date` (string | null) — Date from which to consider the results in ISO-8601 YYYY-MM-DD. See

60 \<https://en.wikipedia.org/wiki/ISO\_8601>.

61 

62 * `max_search_results` (integer | null) — Maximum number of search results to use.

63 

64 * `mode` (string | null) — Choose the mode to query realtime data:

65 \* \`off\`: no search performed and no external will be considered.

66 \* \`on\` (default): the model will search in every sources for relevant data.

67 \* \`auto\`: the model choose whether to search data or not and where to search the data.

68 

69 * `return_citations` (boolean | null) — Whether to return citations in the response or not.

70 

71 * `sources` (array | null) — List of sources to search in. If no sources specified, the model will look over the web and X by default.

72 

73 * `to_date` (string | null) — Date up to which to consider the results in ISO-8601 YYYY-MM-DD. See

74 \<https://en.wikipedia.org/wiki/ISO\_8601>.

75 

76* `service_tier` ("default" | "priority")

77 

78* `store` (boolean | null) — Whether to store the input message(s) and model response for later retrieval.

79 

80* `stream` (boolean | null) — If set, partial message deltas will be sent. Tokens will be sent as data-only server-sent events as they become available, with the stream terminated by a \`data: \[DONE]\` message.

81 

82* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

83 

84* `text` (object)

85 

86 * `format` (object | object | object)

87 

88* `tool_choice` (string | object)

89 

90* `tools` (array | null) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.\`web\_search\_preview\` tool, if specified, will be overridden by \`search\_parameters\`.

91 

92* `top_k` (integer | null) — Top-k sampling: only the \`top\_k\` most probable tokens are considered at each sampling step. Disabled when unset.

93 

94* `top_logprobs` (integer | null) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position, each with an associated log probability. logprobs must be set to true if this parameter is used. Not supported by models \`grok-4.20\` and newer; the field will be silently ignored if set.

95 

96* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

97 

98* `truncation` (string | null) — Not supported. Only maintained for compatibility reasons.

99 

100* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

101 

102### Response Body

103 

104* `background` (boolean, required) — OpenResponses compatibility fields.

105 Not used at the moment. Just for OpenResponses compatibility.

106 Whether to process the response asynchronously in the background.

107 

108* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.

109 

110* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.

111 

112* `error` (object) — An error object returned when the model fails to generate a response.

113 

114* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.

115 

116* `id` (string, required) — Unique ID of the response.

117 

118* `incomplete_details` (object | object | object)

119 

120* `instructions` (string | null) — A system (or developer) message inserted into the model's context.

121 

122* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.

123 

124* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.

125 

126* `metadata` (object, required) — Only included for compatibility.

127 

128* `model` (string, required) — Model name used to generate the response.

129 

130* `object` (string, required) — The object type of this resource. Always set to \`response\`.

131 

132* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.

133 

134* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.

135 

136* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

137 

138* `previous_response_id` (string | null) — The ID of the previous response from the model.

139 

140* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.

141 

142* `reasoning` (object)

143 

144 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.

145 

146 * `generate_summary` (string | null) — Only included for compatibility.

147 

148 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.

149 

150* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.

151 

152* `service_tier` ("default" | "priority", required)

153 

154* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.

155 

156* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.

157 

158* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

159 

160* `text` (object, required)

161 

162 * `format` (object | object | object)

163 

164* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.

165 

166 * `name` (string, required) — Name of the function to use.

167 

168 * `type` (string, required) — Type is always \`"function"\`.

169 

170* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.

171 

172* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.

173 

174* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

175 

176* `truncation` (string, required) — The truncation strategy to use for the model response.

177 

178* `usage` (object)

179 

180 * `context_details` (object)

181 

182 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from

183 \`SamplingUsage.context\_prompt\_tokens\`).

184 

185 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from

186 \`SamplingUsage.context\_output\_tokens\`).

187 

188 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.

189 

190 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

191 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

192 which means there is 10'000'000'000 ticks in one \*dollar\*.

193 

194 * `input_tokens` (integer, required) — Number of input tokens used.

195 

196 * `input_tokens_details` (object, required)

197 

198 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

199 

200 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.

201 

202 * `num_sources_used` (integer, required) — Number of sources used (for live search).

203 

204 * `output_tokens` (integer, required) — Number of output tokens used.

205 

206 * `output_tokens_details` (object, required)

207 

208 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

209 

210 * `server_side_tool_usage_details` (object)

211 

212 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.

213 

214 * `document_search_calls` (integer, required) — Number of document search calls.

215 

216 * `file_search_calls` (integer, required) — Number of file search calls.

217 

218 * `image_generation_calls` (integer, required) — Number of image generation calls.

219 

220 * `mcp_calls` (integer, required) — Number of MCP calls.

221 

222 * `web_search_calls` (integer, required) — Number of web search calls.

223 

224 * `x_search_calls` (integer, required) — Number of X search calls.

225 

226 * `total_tokens` (integer, required) — Total tokens used.

227 

228* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

229 

230### Code Examples

231 

232```bash

233curl -s https://api.x.ai/v1/responses \

234 -H "Content-Type: application/json" \

235 -H "Authorization: Bearer $XAI_API_KEY" \

236 -d '{

237 "model": "grok-4.6",

238 "input": "What is the meaning of life?"

239 }'

240```

241 

242```javascriptAISDK

243import { xai } from "@ai-sdk/xai";

244import { generateText } from "ai";

245 

246const result = await generateText({

247 model: xai.responses("grok-4.6"),

248 prompt: "What is the meaning of life?",

249});

250 

251console.log(JSON.stringify(result, null, 2));

252```

253 

254```pythonOpenAISDK

255import os

256 

257from openai import OpenAI

258 

259client = OpenAI(

260 api_key=os.environ["XAI_API_KEY"],

261 base_url="https://api.x.ai/v1",

262)

263 

264response = client.responses.create(

265 model="grok-4.6",

266 input="What is the meaning of life?",

267)

268 

269print(response.model_dump_json(indent=2))

270```

271 

272```javascriptOpenAISDK

273import OpenAI from "openai";

274 

275const client = new OpenAI({

276 apiKey: process.env.XAI_API_KEY,

277 baseURL: "https://api.x.ai/v1",

278});

279 

280const response = await client.responses.create({

281 model: "grok-4.6",

282 input: "What is the meaning of life?",

283});

284 

285console.log(JSON.stringify(response, null, 2));

286```

287 

288\*\*Response example:\*\*

289 

290```json

291{

292 "created_at": 1754475266,

293 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",

294 "max_output_tokens": null,

295 "model": "latest",

296 "object": "response",

297 "output": [

298 {

299 "content": [

300 {

301 "type": "output_text",

302 "text": "101 multiplied by 3 is 303.",

303 "logprobs": null,

304 "annotations": []

305 }

306 ],

307 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",

308 "role": "assistant",

309 "type": "message",

310 "status": "completed"

311 }

312 ],

313 "parallel_tool_calls": true,

314 "previous_response_id": null,

315 "reasoning": null,

316 "temperature": null,

317 "text": {

318 "format": {

319 "type": "text"

320 }

321 },

322 "tool_choice": "auto",

323 "tools": [],

324 "top_p": null,

325 "usage": {

326 "input_tokens": 32,

327 "input_tokens_details": {

328 "cached_tokens": 8

329 },

330 "output_tokens": 9,

331 "output_tokens_details": {

332 "reasoning_tokens": 110

333 },

334 "total_tokens": 151,

335 "num_sources_used": 0,

336 "num_server_side_tools_used": 0

337 },

338 "user": null,

339 "incomplete_details": null,

340 "status": "completed",

341 "store": true

342}

343```

344 

345***

346 

347## GET /v1/responses/\{response\_id}

348 

349Retrieve a previously generated response.

350 

351### Path Parameters

352 

353* `response_id` (string, required) — The response id returned by a previous create response request.

354 

355### Response Body

356 

357* `background` (boolean, required) — OpenResponses compatibility fields.

358 Not used at the moment. Just for OpenResponses compatibility.

359 Whether to process the response asynchronously in the background.

360 

361* `completed_at` (integer | null) — The Unix timestamp (in seconds) for the response completion time. Only set when the response is completed.

362 

363* `created_at` (integer, required) — The Unix timestamp (in seconds) for the response creation time.

364 

365* `error` (object) — An error object returned when the model fails to generate a response.

366 

367* `frequency_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on their existing frequency in the text so far, decreasing the model's likelihood to repeat the same line verbatim.

368 

369* `id` (string, required) — Unique ID of the response.

370 

371* `incomplete_details` (object | object | object)

372 

373* `instructions` (string | null) — A system (or developer) message inserted into the model's context.

374 

375* `max_output_tokens` (integer | null) — Max number of tokens that can be generated in a response. This includes both output and reasoning tokens.

376 

377* `max_tool_calls` (integer | null) — The maximum number of tool calls allowed for this response.

378 

379* `metadata` (object, required) — Only included for compatibility.

380 

381* `model` (string, required) — Model name used to generate the response.

382 

383* `object` (string, required) — The object type of this resource. Always set to \`response\`.

384 

385* `output` (array\<object | object | object | object | object | object | object | object | object | object | object | object>, required) — The response generated by the model.

386 

387* `parallel_tool_calls` (boolean, required) — Whether to allow the model to run parallel tool calls.

388 

389* `presence_penalty` (number, required) — (NOT SUPPORTED in Responses API) Positive values penalize new tokens based on whether they appear in the text so far, increasing the model's likelihood to talk about new topics.

390 

391* `previous_response_id` (string | null) — The ID of the previous response from the model.

392 

393* `prompt_cache_key` (string | null) — The cache key used for the prompt for routing to the correct engine.

394 

395* `reasoning` (object)

396 

397 * `effort` (string | null) — Constrains how hard a reasoning model thinks before responding. Supported by some models; models that do not support it reject the request with an error. Possible values are \`none\` (disables reasoning completely), \`low\`, \`medium\`, \`high\` (uses the most reasoning tokens) and \`xhigh\`. The accepted values and the default used when unspecified vary per model. See the model's documentation page for details.

398 

399 * `generate_summary` (string | null) — Only included for compatibility.

400 

401 * `summary` (string | null) — A summary of the model's reasoning process. Possible values are \`auto\`, \`concise\` and \`detailed\`. Only included for compatibility. The model shall always return \`detailed\`.

402 

403* `safety_identifier` (string | null) — A stable identifier used to help detect users of your application that may be violating xAI's usage policies.

404 

405* `service_tier` ("default" | "priority", required)

406 

407* `status` (string, required) — Status of the response. One of \`completed\`, \`in\_progress\` or \`incomplete\`.

408 

409* `store` (boolean, required) — Whether to store the input message(s) and model response for later retrieval.

410 

411* `temperature` (number | null) — What sampling temperature to use, between 0 and 2. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.

412 

413* `text` (object, required)

414 

415 * `format` (object | object | object)

416 

417* `tool_choice` (string | object, required) — Parameter to control how model chooses the tools.

418 

419 * `name` (string, required) — Name of the function to use.

420 

421 * `type` (string, required) — Type is always \`"function"\`.

422 

423* `tools` (array\<object | object | object | object | object | object | object | object | object>, required) — A list of tools the model may call in JSON-schema. Currently, only functions and web search are supported as tools. A max of 128 tools are supported.

424 

425* `top_logprobs` (integer, required) — An integer between 0 and 8 specifying the number of most likely tokens to return at each token position.

426 

427* `top_p` (number | null) — An alternative to sampling with \`temperature\`, called nucleus sampling, where the model considers the results of the tokens with \`top\_p\` probability mass. So 0.1 means only the tokens comprising the top 10% probability mass are considered. It is generally recommended to alter this or \`temperature\` but not both.

428 

429* `truncation` (string, required) — The truncation strategy to use for the model response.

430 

431* `usage` (object)

432 

433 * `context_details` (object)

434 

435 * `input_tokens` (integer, required) — Prompt tokens in the latest context (sourced from

436 \`SamplingUsage.context\_prompt\_tokens\`).

437 

438 * `output_tokens` (integer, required) — Completion + reasoning tokens in the latest context (sourced from

439 \`SamplingUsage.context\_output\_tokens\`).

440 

441 * `cost_in_nano_usd` (integer | null) — Cost in nano US dollars for this request.

442 

443 * `cost_in_usd_ticks` (integer | null) — Accurate cost of this request in USD ticks, where "tick" is defined as follows:

444 TICKS\_IN\_USD\_CENT: i64 = 100\_000\_000

445 which means there is 10'000'000'000 ticks in one \*dollar\*.

446 

447 * `input_tokens` (integer, required) — Number of input tokens used.

448 

449 * `input_tokens_details` (object, required)

450 

451 * `cached_tokens` (integer, required) — Token cached by xAI from previous requests and reused for this request.

452 

453 * `num_server_side_tools_used` (integer, required) — Number of server side tools used.

454 

455 * `num_sources_used` (integer, required) — Number of sources used (for live search).

456 

457 * `output_tokens` (integer, required) — Number of output tokens used.

458 

459 * `output_tokens_details` (object, required)

460 

461 * `reasoning_tokens` (integer, required) — Tokens generated by the model for reasoning.

462 

463 * `server_side_tool_usage_details` (object)

464 

465 * `code_interpreter_calls` (integer, required) — Number of code interpreter calls.

466 

467 * `document_search_calls` (integer, required) — Number of document search calls.

468 

469 * `file_search_calls` (integer, required) — Number of file search calls.

470 

471 * `image_generation_calls` (integer, required) — Number of image generation calls.

472 

473 * `mcp_calls` (integer, required) — Number of MCP calls.

474 

475 * `web_search_calls` (integer, required) — Number of web search calls.

476 

477 * `x_search_calls` (integer, required) — Number of X search calls.

478 

479 * `total_tokens` (integer, required) — Total tokens used.

480 

481* `user` (string | null) — A unique identifier representing your end-user, which can help xAI to monitor and detect abuse.

482 

483\*\*Response example:\*\*

484 

485```json

486{

487 "created_at": 1754475266,

488 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",

489 "max_output_tokens": null,

490 "model": "latest",

491 "object": "response",

492 "output": [

493 {

494 "content": [

495 {

496 "type": "output_text",

497 "text": "101 multiplied by 3 is 303.",

498 "logprobs": null,

499 "annotations": []

500 }

501 ],

502 "id": "msg_ad5663da-63e6-86c6-e0be-ff15effa8357",

503 "role": "assistant",

504 "type": "message",

505 "status": "completed"

506 },

507 {

508 "id": "",

509 "summary": [

510 {

511 "text": "First, the user asked: \"What is 101*3?\"\n\nThis is a simple multiplication: 101 multiplied by 3.\n\nCalculating: 100 * 3 = 300, and 1 * 3 = 3, so 300 + 3 = 303.\n\nI should respond helpfully and directly, as per my system prompt: \"You are a helpful assistant that can answer questions and help with tasks.\"\n\nKeep the response concise and accurate. No need for extra fluff unless it adds value.\n\nFinal answer: 303.",

512 "type": "summary_text"

513 }

514 ],

515 "type": "reasoning",

516 "status": "completed"

517 }

518 ],

519 "parallel_tool_calls": true,

520 "previous_response_id": null,

521 "reasoning": null,

522 "temperature": null,

523 "text": {

524 "format": {

525 "type": "text"

526 }

527 },

528 "tool_choice": "auto",

529 "tools": [],

530 "top_p": null,

531 "usage": {

532 "prompt_tokens": 32,

533 "completion_tokens": 9,

534 "total_tokens": 151,

535 "prompt_tokens_details": {

536 "text_tokens": 32,

537 "audio_tokens": 0,

538 "image_tokens": 0,

539 "cached_tokens": 8

540 },

541 "completion_tokens_details": {

542 "reasoning_tokens": 110,

543 "audio_tokens": 0,

544 "accepted_prediction_tokens": 0,

545 "rejected_prediction_tokens": 0

546 },

547 "num_sources_used": 0

548 },

549 "user": null,

550 "incomplete_details": null,

551 "status": "completed",

552 "store": true

553}

554```

555 

556***

557 

558## GET /v1/responses/\{response\_id}/input\_items

559 

560List input items for a previously generated response.

561 

562### Path Parameters

563 

564* `response_id` (string, required) — The response id returned by a previous create response request.

565 

566### Query Parameters

567 

568* `limit` (integer) — Maximum number of items to return (1-100, default 20).

569 

570* `order` ("asc" | "desc") — Sort order: asc or desc. Default asc.

571 

572* `after` (string) — Cursor for pagination. Returns items after this item ID.

573 

574### Response Body

575 

576* `data` (array\<object>, required) — The list of input items.

577 

578* `first_id` (string | null) — The ID of the first item in the list.

579 

580* `has_more` (boolean, required) — Whether there are more items beyond this page.

581 

582* `last_id` (string | null) — The ID of the last item in the list.

583 

584* `object` (string, required) — The object type, always \`list\`.

585 

586\*\*Response example:\*\*

587 

588```json

589{}

590```

591 

592***

593 

594## DELETE /v1/responses/\{response\_id}

595 

596Delete a previously generated response.

597 

598### Path Parameters

599 

600* `response_id` (string, required) — The response id returned by a previous create response request.

601 

602### Response Body

603 

604* `deleted` (boolean, required) — Whether the response was successfully deleted.

605 

606* `id` (string, required) — The response\_id to be deleted.

607 

608* `object` (string, required) — The deleted object type, which is always \`response\`.

609 

610\*\*Response example:\*\*

611 

612```json

613{

614 "id": "ad5663da-63e6-86c6-e0be-ff15effa8357",

615 "object": "response",

616 "deleted": true

617}

618```

619 

620***

621 

622## POST /v1/responses/compact

623 

624Compacts a full Responses API input window into a shorter canonical window.

625 

626### Request Body

627 

628* `input` (string | array\<object | object | object | object | object>, required) — Content of the input passed to a \`/v1/response\` request.

629 

630* `model` (string, required) — Model to use for compaction summarization (required).

631 

632### Response Body

633 

634* `created_at` (integer, required) — Unix timestamp (in seconds) when the compacted conversation was created.

635 

636* `id` (string, required) — Unique ID for this compaction (e.g. \`cmp\_\<uuid>\`).

637 

638* `model` (string, required) — Model used for the compaction summary.

639 

640* `object` (string, required) — Always \`"response.compaction"\`.

641 

642* `output` (array\<object | object | object | object | object>, required) — Compacted output containing a single compaction item.

643 Pass this verbatim as input to the next \`/v1/responses\` call.

644 

645* `usage` (object)

646 

647 * `dropped_message_count` (integer, required) — Number of messages dropped/summarized during compaction.

648 

649 * `input_tokens` (integer, required) — Number of tokens in the original (pre-compaction) input.

650 

651 * `input_tokens_details` (object, required) — Breakdown of input tokens for a compaction call.

652 

653 * `cached_tokens` (integer, required) — Number of input tokens that were served from the prompt cache.

654 

655 * `output_tokens` (integer, required) — Number of tokens in the compacted output.

656 

657 * `output_tokens_details` (object, required) — Breakdown of output tokens for a compaction call.

658 

659 * `reasoning_tokens` (integer, required) — Number of reasoning tokens generated during compaction.

660 

661 * `total_tokens` (integer, required) — Total number of tokens used (input + output, including reasoning).

662 

663### Code Examples

664 

665```bash

666curl -s https://api.x.ai/v1/responses/compact \

667 -H "Content-Type: application/json" \

668 -H "Authorization: Bearer $XAI_API_KEY" \

669 -d '{

670 "model": "grok-4.6",

671 "input": [

672 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},

673 {"role": "user", "content": "What is the Higgs boson and why is it important?"},

674 {"role": "assistant", "content": "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."},

675 {"role": "user", "content": "How does the Higgs mechanism actually work?"},

676 {"role": "assistant", "content": "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless."}

677 ]

678 }'

679```

680 

681```pythonOpenAISDK

682import os

683 

684from openai import OpenAI

685 

686client = OpenAI(

687 api_key=os.environ["XAI_API_KEY"],

688 base_url="https://api.x.ai/v1",

689)

690 

691compacted = client.responses.compact(

692 model="grok-4.6",

693 input=[

694 {"role": "system", "content": "You are a concise and knowledgeable science tutor."},

695 {"role": "user", "content": "What is the Higgs boson and why is it important?"},

696 {

697 "role": "assistant",

698 "content": (

699 "The Higgs boson is an elementary particle in the Standard Model, predicted by "

700 "Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation "

701 "of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism."

702 ),

703 },

704 {"role": "user", "content": "How does the Higgs mechanism actually work?"},

705 {

706 "role": "assistant",

707 "content": (

708 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, "

709 "and particles acquire mass in proportion to how strongly they couple to it. Photons "

710 "do not couple, which is why they remain massless."

711 ),

712 },

713 ],

714)

715 

716print(compacted.model_dump_json(indent=2))

717```

718 

719```javascriptOpenAISDK

720import OpenAI from "openai";

721 

722const client = new OpenAI({

723 apiKey: process.env.XAI_API_KEY,

724 baseURL: "https://api.x.ai/v1",

725});

726 

727const compacted = await client.responses.compact({

728 model: "grok-4.6",

729 input: [

730 { role: "system", content: "You are a concise and knowledgeable science tutor." },

731 { role: "user", content: "What is the Higgs boson and why is it important?" },

732 {

733 role: "assistant",

734 content:

735 "The Higgs boson is an elementary particle in the Standard Model, predicted by Peter Higgs in 1964 and confirmed at CERN in 2012. It is the quantum excitation of the Higgs field, which gives mass to fundamental particles via the Higgs mechanism.",

736 },

737 { role: "user", content: "How does the Higgs mechanism actually work?" },

738 {

739 role: "assistant",

740 content:

741 "Through spontaneous symmetry breaking. The Higgs field has a nonzero vacuum value, and particles acquire mass in proportion to how strongly they couple to it. Photons do not couple, which is why they remain massless.",

742 },

743 ],

744});

745 

746console.log(JSON.stringify(compacted, null, 2));

747```

748 

749\*\*Response example:\*\*

750 

751```json

752{}

753```