SpyBara
Go Premium

Documentation 2026-08-08 23:00 UTC to 2026-08-11 20:57 UTC

5 files changed +71 −4. View all changes and history on the product overview
2026
Mon 31 20:58 Tue 25 22:58 Fri 21 18:57 Thu 20 15:58 Wed 19 18:02 Tue 18 04:01 Thu 13 22:00 Wed 12 23:59 Tue 11 20:57 Sat 8 23:00 Fri 7 17:57 Thu 6 20:01 Mon 3 23:00 Sat 1 01:59
Details

2 2 

3# Image Generation3# Image Generation

4 4 

5Generate images from text prompts with Grok Imagine models. The API supports batch generation of multiple images, and control over aspect ratio and resolution.5Generate images from text prompts with Grok Imagine models. The API supports batch generation of multiple images, and control over aspect ratio, resolution, and quality.

6 6 

7## Quick Start7## Quick Start

8 8 


333}'333}'

334```334```

335 335 

336### Quality

337 

338Control generation quality with the optional `quality` parameter. Allowed values are `low` and `medium`. When omitted, the default is `medium`. The parameter is only supported for `grok-imagine-image-2.0`.

339 

336### Base64 Output340### Base64 Output

337 341 

338For embedding images directly without downloading, request base64:342For embedding images directly without downloading, request base64:

quickstart.md +3 −1

Details

113console.log(response.output_text);113console.log(response.output_text);

114```114```

115 115 

116For multi-turn chat, reasoning, and [structured outputs](/developers/model-capabilities/text/structured-outputs), see the [Text Generation Guide](/developers/model-capabilities/text/generate-text). For agentic coding workflows, see the [Grok Build overview](/build/overview).116For multi-turn chat, reasoning, and [structured outputs](/developers/model-capabilities/text/structured-outputs), see the [Text Generation Guide](/developers/model-capabilities/text/generate-text). For agentic coding workflows, see the [Grok Build overview](/build/overview). For AI teammates on a cloud computer, see [Grok Bot](/grok-bot/overview).

117 117 

118## Step 5: Generate an image118## Step 5: Generate an image

119 119 


201* [Tools](/developers/tools/overview) - Web search, X search, code execution, and function calling201* [Tools](/developers/tools/overview) - Web search, X search, code execution, and function calling

202* [Models](/developers/models) - Compare available models and their capabilities202* [Models](/developers/models) - Compare available models and their capabilities

203* [Pricing](/developers/pricing) - Tools, batch API, and other platform pricing203* [Pricing](/developers/pricing) - Tools, batch API, and other platform pricing

204* [Grok Bot](/grok-bot/overview) - AI teammates on a persistent cloud computer

205* [Grok Build](/build/overview) - Agentic coding CLI and API

rate-limits.md +2 −1

Details

39| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |39| grok-4.20-0309-non-reasoning | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

40| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |40| grok-build-0.1 | T0: 37, T1: 50, T2: 75, T3: 125, T4: 208 | T0: 10M, T1: 15M, T2: 25M, T3: 45M, T4: 85M |

41| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |41| grok-4.20-multi-agent-0309 | T0: 9, T1: 12, T2: 18, T3: 31, T4: 56 | T0: 2.5M, T1: 3.7M, T2: 6.2M, T3: 11M, T4: 21M |

42| grok-imagine-image-quality | 5 | — |

43| grok-imagine-image | 5 | — |42| grok-imagine-image | 5 | — |

43| grok-imagine-image-quality | 5 | — |

44| grok-imagine-image-2.0 | 5 | — |

44| grok-imagine-video-1.5 | 10 | — |45| grok-imagine-video-1.5 | 10 | — |

45| grok-imagine-video | 10 | — |46| grok-imagine-video | 10 | — |

46 47 

Details

27 * `id` (string, required) — Model ID. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.27 * `id` (string, required) — Model ID. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

28 28 

29 * `image_price` (integer | null) — Price per image in USD cents (image generation models).29 * `image_price` (integer | null) — Price per image in USD cents (image generation models).

30 The default tier: medium quality at the default (1k) resolution. See

31 \`pricing\` for the full quality/resolution matrix.

30 32 

31 * `long_context_threshold` (integer | null) — Token count at or above which the long context prices apply.33 * `long_context_threshold` (integer | null) — Token count at or above which the long context prices apply.

32 34 


34 36 

35 * `owned_by` (string, required) — Owner of the model.37 * `owned_by` (string, required) — Owner of the model.

36 38 

39 * `pricing` (array\<object>) — Per-image prices by (quality, resolution) tier (image generation

40 models). Omitted when the model prices all qualities identically

41 (see \`image\_price\`).

42 

43 * `price_per_image` (integer, required) — Price per generated image, in 1/100,000,000ths of a USD cent.

44 

45 * `quality` (string, required) — Quality tier this price applies to: \`"low"\`, \`"medium"\`, or \`"high"\`.

46 Medium is the default quality a request serves at when it leaves

47 \`quality\` unset.

48 

49 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the

50 default when a request leaves \`resolution\` unset.

51 

37 * `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.52 * `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.

38 53 

39 * `prompt_text_token_price` (integer | null) — Price of the prompt text token in USD cents per 100 million tokens.54 * `prompt_text_token_price` (integer | null) — Price of the prompt text token in USD cents per 100 million tokens.


117* `id` (string, required) — Model ID. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.132* `id` (string, required) — Model ID. Obtainable from \<https://console.x.ai/team/default/models> or \<https://docs.x.ai/docs/models>.

118 133 

119* `image_price` (integer | null) — Price per image in USD cents (image generation models).134* `image_price` (integer | null) — Price per image in USD cents (image generation models).

135 The default tier: medium quality at the default (1k) resolution. See

136 \`pricing\` for the full quality/resolution matrix.

120 137 

121* `long_context_threshold` (integer | null) — Token count at or above which the long context prices apply.138* `long_context_threshold` (integer | null) — Token count at or above which the long context prices apply.

122 139 


124 141 

125* `owned_by` (string, required) — Owner of the model.142* `owned_by` (string, required) — Owner of the model.

126 143 

144* `pricing` (array\<object>) — Per-image prices by (quality, resolution) tier (image generation

145 models). Omitted when the model prices all qualities identically

146 (see \`image\_price\`).

147 

148 * `price_per_image` (integer, required) — Price per generated image, in 1/100,000,000ths of a USD cent.

149 

150 * `quality` (string, required) — Quality tier this price applies to: \`"low"\`, \`"medium"\`, or \`"high"\`.

151 Medium is the default quality a request serves at when it leaves

152 \`quality\` unset.

153 

154 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the

155 default when a request leaves \`resolution\` unset.

156 

127* `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.157* `prompt_image_token_price` (integer | null) — Price of the prompt image token in USD cents per 100 million tokens.

128 158 

129* `prompt_text_token_price` (integer | null) — Price of the prompt text token in USD cents per 100 million tokens.159* `prompt_text_token_price` (integer | null) — Price of the prompt text token in USD cents per 100 million tokens.


354 * `id` (string, required) — Model ID.384 * `id` (string, required) — Model ID.

355 385 

356 * `image_price` (integer, required) — Price of a single image in USD cents.386 * `image_price` (integer, required) — Price of a single image in USD cents.

387 The default tier: medium quality at the default (1k) resolution. See

388 \`pricing\` for the full quality/resolution matrix.

357 389 

358 * `input_modalities` (array\<string>, required) — The input modalities supported by the model.390 * `input_modalities` (array\<string>, required) — The input modalities supported by the model.

359 391 


365 397 

366 * `owned_by` (string, required) — Owner of the model.398 * `owned_by` (string, required) — Owner of the model.

367 399 

400 * `pricing` (array\<object>) — Per-image prices by (quality, resolution) tier. One entry per

401 combination the model serves; omitted when the model prices all

402 qualities identically (see \`image\_price\`).

403 

404 * `price_per_image` (integer, required) — Price per generated image, in 1/100,000,000ths of a USD cent.

405 

406 * `quality` (string, required) — Quality tier this price applies to: \`"low"\`, \`"medium"\`, or \`"high"\`.

407 Medium is the default quality a request serves at when it leaves

408 \`quality\` unset.

409 

410 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the

411 default when a request leaves \`resolution\` unset.

412 

368 * `version` (string, required) — Version of the model.413 * `version` (string, required) — Version of the model.

369 414 

370\*\*Response example:\*\*415\*\*Response example:\*\*


410* `id` (string, required) — Model ID.455* `id` (string, required) — Model ID.

411 456 

412* `image_price` (integer, required) — Price of a single image in USD cents.457* `image_price` (integer, required) — Price of a single image in USD cents.

458 The default tier: medium quality at the default (1k) resolution. See

459 \`pricing\` for the full quality/resolution matrix.

413 460 

414* `input_modalities` (array\<string>, required) — The input modalities supported by the model.461* `input_modalities` (array\<string>, required) — The input modalities supported by the model.

415 462 


421 468 

422* `owned_by` (string, required) — Owner of the model.469* `owned_by` (string, required) — Owner of the model.

423 470 

471* `pricing` (array\<object>) — Per-image prices by (quality, resolution) tier. One entry per

472 combination the model serves; omitted when the model prices all

473 qualities identically (see \`image\_price\`).

474 

475 * `price_per_image` (integer, required) — Price per generated image, in 1/100,000,000ths of a USD cent.

476 

477 * `quality` (string, required) — Quality tier this price applies to: \`"low"\`, \`"medium"\`, or \`"high"\`.

478 Medium is the default quality a request serves at when it leaves

479 \`quality\` unset.

480 

481 * `resolution` (string, required) — Output resolution this price applies to: \`"1k"\` or \`"2k"\`. 1k is the

482 default when a request leaves \`resolution\` unset.

483 

424* `version` (string, required) — Version of the model.484* `version` (string, required) — Version of the model.

425 485 

426\*\*Response example:\*\*486\*\*Response example:\*\*

Details

393 393 

394### Request Body394### Request Body

395 395 

396* `text` (string, required) — The text to convert to speech. Maximum 15,000 characters. Supports inline speech tags for expressive output: \`\[pause]\`, \`\[long-pause]\`, \`\[hum-tune]\`, \`\[laugh]\`, \`\[chuckle]\`, \`\[giggle]\`, \`\[cry]\`, \`\[tsk]\`, \`\[tongue-click]\`, \`\[lip-smack]\`, \`\[breath]\`, \`\[inhale]\`, \`\[exhale]\`, \`\[sigh]\`. Also supports wrapping tags for style control: \`\<soft>\`, \`\<whisper>\`, \`\<loud>\`, \`\<build-intensity>\`, \`\<decrease-intensity>\`, \`\<higher-pitch>\`, \`\<lower-pitch>\`, \`\<slow>\`, \`\<fast>\`, \`\<sing-song>\`, \`\<singing>\`, \`\<laugh-speak>\`, \`\<emphasis>\`.396* `text` (string, required) — The text to convert to speech. Maximum 15,000 characters. Supports inline speech tags for expressive output: \`\[pause]\`, \`\[long-pause]\`, \`\[hum-tune]\`, \`\[laugh]\`, \`\[chuckle]\`, \`\[giggle]\`, \`\[cry]\`, \`\[tsk]\`, \`\[tongue-click]\`, \`\[lip-smack]\`, \`\[breath]\`, \`\[inhale]\`, \`\[exhale]\`, \`\[sigh]\`. Also supports wrapping tags for style control: \`\<soft>\`, \`\<whisper>\`, \`\<loud>\`, \`\<build-intensity>\`, \`\<decrease-intensity>\`, \`\<higher-pitch>\`, \`\<lower-pitch>\`, \`\<slow>\`, \`\<fast>\`, \`\<sing-song>\`, \`\<singing>\`, \`\<emphasis>\`.

397 397 

398* `voice_id` (string) — Voice identifier. Use a built-in voice from \`GET /v1/tts/voices\` (e.g. \`eve\`, \`ara\`) or a custom voice ID. Defaults to \`eve\` when omitted.398* `voice_id` (string) — Voice identifier. Use a built-in voice from \`GET /v1/tts/voices\` (e.g. \`eve\`, \`ara\`) or a custom voice ID. Defaults to \`eve\` when omitted.

399 399