11 11
12 12
13 13
14In this guide, you will learn about building applications involving images with the OpenAI API.14<a id="a-tour-of-image-related-use-cases"></a>
15If you know what you want to build, find your use case below to get started. If you're not sure where to start, continue reading to get an overview.
16
17### A tour of image-related use cases
18 15
19Recent language models can process image inputs and analyze them—a capability known as **vision**. GPT Image models can use text and image inputs to create new images or edit existing ones.16Recent language models can process image inputs and analyze them—a capability known as **vision**. GPT Image models can use text and image inputs to create new images or edit existing ones.
20 17
21The OpenAI API offers several endpoints to process images as input or generate them as output, enabling you to build powerful multimodal applications.18Choose an endpoint based on whether you want to analyze images or generate them:
22 19
23| API | Supported use cases |20| API | Supported use cases |
24| ---------------------------------------------------- | --------------------------------------------------------------------- |21| ---------------------------------------------------- | -------------------------------------------------------------------------- |
25| [Responses API](https://developers.openai.com/api/reference/resources/responses) | Analyze images and use them as input and/or generate images as output |22| [Responses API](https://developers.openai.com/api/reference/resources/responses) | Analyze images, or generate and edit images with the image generation tool |
26| [Images API](https://developers.openai.com/api/reference/resources/images) | Generate images as output, optionally using images as input |23| [Images API](https://developers.openai.com/api/reference/resources/images) | Generate images as output, optionally using images as input |
27| [Chat Completions API](https://developers.openai.com/api/reference/resources/chat) | Analyze images and use them as input to generate text or audio |24| [Chat Completions API](https://developers.openai.com/api/reference/resources/chat) | Analyze images and generate text responses |
28 25
29To learn more about the input and output modalities supported by our models, refer to our [models page](https://developers.openai.com/api/docs/models).26To learn more about the input and output modalities supported by our models, refer to our [models page](https://developers.openai.com/api/docs/models).
30 27
31## Generate or edit images28## Generate or edit images
32 29
33You can generate or edit images using the Image API or the Responses API.30With the Images API, choose `gpt-image-2` to generate images from text or edit existing images. With the Responses API, choose a mainline model that supports the image generation tool; the tool handles GPT Image model selection.
34
35The state-of-the-art image generation model, `gpt-image-2`, can understand text and images and use broad world knowledge to generate images with strong instruction following and contextual awareness.
36 31
37 32
38 33
158Files.write(Path.of("cat_and_otter.png"), Base64.getDecoder().decode(imageResult));153Files.write(Path.of("cat_and_otter.png"), Base64.getDecoder().decode(imageResult));
159```154```
160 155
156```csharp
157using OpenAI.Responses;
158#pragma warning disable OPENAI001
159
160string key = Environment.GetEnvironmentVariable("OPENAI_API_KEY")!;
161ResponsesClient client = new(key);
162
163CreateResponseOptions options = new()
164{
165 Model = "gpt-5.6",
166};
167options.InputItems.Add(
168 ResponseItem.CreateUserMessageItem(
169 "Generate an image of a gray tabby cat hugging an otter with an orange scarf."
170 )
171);
172options.Tools.Add(
173 ResponseTool.CreateImageGenerationTool(model: "gpt-image-2")
174);
175
176ResponseResult response = await client.CreateResponseAsync(options);
177ImageGenerationCallResponseItem image = response
178 .OutputItems.OfType<ImageGenerationCallResponseItem>()
179 .FirstOrDefault()
180 ?? throw new InvalidOperationException("No generated image was returned.");
181await File.WriteAllBytesAsync(
182 "cat_and_otter.png",
183 image.ImageResultBytes.ToArray()
184);
185```
186
161```ruby187```ruby
162require "base64"188require "base64"
163require "openai"189require "openai"
200 226
201### Using world knowledge for image generation227### Using world knowledge for image generation
202 228
203GPT Image models can use visual understanding of the world to generate lifelike images including real-life details without a reference.229GPT Image models can draw on world knowledge without a reference image. For example, a prompt for a cabinet of semi-precious stones can produce a scene containing recognizable gemstones such as amethyst, rose quartz, and jade.
204
205For example, if you prompt GPT Image to generate an image of a glass cabinet with the most popular semi-precious stones, the model knows enough to select gemstones like amethyst, rose quartz, jade, etc, and depict them in a realistic way.
206 230
207## Analyze images231## Analyze images
208 232
209**Vision** is the ability for a model to "see" and understand images. If there is text in an image, the model can also understand the text.233Use a vision-capable model to describe images, read visible text, and answer questions about objects, shapes, colors, or textures. Account for the model's [limitations](#limitations) when using its answers.
210It can understand most visual elements, including objects, shapes, colors, and textures, even if there are some [limitations](#limitations).
211 234
212### Giving a model images as input235### Giving a model images as input
213 236
215 238
216 239
217 240
218You can provide images as input to generation requests in multiple ways:241Provide an image for analysis in any of these ways:
219 242
220- By providing a fully qualified URL to an image file243- By providing a fully qualified URL to an image file
221- By providing an image as a Base64-encoded data URL244- By providing an image as a Base64-encoded data URL
938 961
939### Image input requirements962### Image input requirements
940 963
941Input images must meet the following requirements to be used in the API.964Use supported image files that are clear enough for the model to analyze.
942 965
943<table>966| Requirement | Supported inputs |
944 <tr>967| ------------ | ------------------------------------------------------------------------------------- |
945 <td>Supported file types</td>968| File types | PNG (`.png`), JPEG (`.jpeg` or `.jpg`), WEBP (`.webp`), and non-animated GIF (`.gif`) |
946 <td>969| Request size | Up to 512 MB total payload per request |
947 - PNG (`.png`) - JPEG (`.jpeg` and `.jpg`) - WEBP (`.webp`) - Non-animated970| Image count | Up to 1,500 images per request |
948 GIF (`.gif`)971
949 </td>972Image tokens and the rest of your prompt must also fit the model's input and context limits. A token estimate does not guarantee that a request meets every input limit. Image use must comply with our [usage policies](https://openai.com/policies/usage-policies/).
950 </tr>
951 <tr>
952 <td>Size limits</td>
953 <td>
954 - Up to 512 MB total payload size per request - Up to 1500 individual
955 image inputs per request
956 </td>
957 </tr>
958 <tr>
959 <td>Other requirements</td>
960 <td>
961 - No watermarks or logos - No NSFW content - Clear enough for a human to
962 understand
963 </td>
964 </tr>
965</table>
966 973
967### Choose an image detail level974### Choose an image detail level
968 975
969The `detail` parameter tells the model what level of detail to use when processing and understanding the image (`low`, `high`, `original`, or `auto`). If you skip the parameter, the model will use `auto`. This behavior is the same in both the Responses API and the Chat Completions API. On `gpt-5.5` and GPT-5.6 models, `auto` and the default omitted behavior are equivalent to `original`.976The `detail` parameter controls image preprocessing. Supported values depend on the model: `low`, `high`, `original`, or `auto`. If you omit the parameter, it defaults to `auto` in both the Responses API and the Chat Completions API. The [model sizing table](#model-sizing-behavior) shows the corresponding behavior.
970 977
971 978
972 979
984Use the following guidance to choose a detail level:991Use the following guidance to choose a detail level:
985 992
986| Detail level | Best for |993| Detail level | Best for |
987| ------------ | ---------------------------------------------------------------------------------------------------------------------------------------------- |994| ------------ | --------------------------------------------------------------------------------------------------------------------------- |
988| `low` | Fast, low-cost understanding when fine visual detail is not important. The model receives a low-resolution 512px x 512px version of the image. |995| `low` | Coarse image understanding. Resizing and token use depend on the model; `low` does not always use fewer tokens than `high`. |
989| `high` | Standard high-fidelity image understanding when precise original-image coordinates are not required. |996| `high` | Standard high-fidelity image understanding when precise original-image coordinates are not required. |
990| `original` | Large, dense, spatially sensitive, or computer-use images. Available on `gpt-5.4` and future models. |997| `original` | Large, dense, spatially sensitive, or computer-use images, when supported by the model. |
991| `auto` | Automatic detail selection. On `gpt-5.5` and GPT-5.6 models, `auto` and the omitted/default behavior are equivalent to `original`. |998| `auto` | Use the model's default sizing behavior, shown in the model sizing table. |
992 999
993For high-accuracy tasks that require fine visual detail or precise coordinates in the original image, such as optical character recognition (OCR), small-object detection, bounding boxes, localization, or computer use, set `"detail": "original"` when supported. The `low` and `high` detail levels may resize the image before analysis, which can obscure small details and cause model-generated coordinates to no longer match the original image. On `gpt-5.4` and `gpt-5.5`, `original` can also resize images that exceed the model's patch or dimension limits; for coordinate-sensitive tasks, resize those images before sending them and remap returned coordinates to the original image. Use `low` or `high` when lower cost or latency is more important than fine-detail recognition or spatial accuracy. See the [Computer use guide](https://developers.openai.com/api/docs/guides/tools-computer-use) for more detail.1000For tasks that require fine visual detail or precise coordinates, such as optical character recognition (OCR), small-object detection, or computer use, use `"detail": "original"` when supported. Original detail can still resize images that exceed the model's limits. For coordinate-sensitive tasks, resize images to fit those limits before sending them and map returned coordinates back to the original image. See the [Computer use guide](https://developers.openai.com/api/docs/guides/tools-computer-use) for coordinate handling.
994
995Read more about how models resize images in the [Model sizing
996 behavior](#model-sizing-behavior) section, and about token costs in the
997 [Calculating costs](#calculating-costs) section below.
998 1001
999### Model sizing behavior1002### Model sizing behavior
1000 1003
1001Different models use different resizing rules before image tokenization:1004The following table covers the general-purpose vision models available in the [image input cost calculator](#image-input-cost-calculator). Other models and specialized variants can use different limits. All resizing preserves aspect ratio without enlarging smaller images.
1002 1005
1003<table>1006<table>
1004 <tr>1007 <tr>
1007 <th>Patch and resizing behavior</th>1010 <th>Patch and resizing behavior</th>
1008 </tr>1011 </tr>
1009 <tr>1012 <tr>
1010 <td>GPT-5.6 family</td>1013 <td>
1014 `gpt-5.6-sol`, `gpt-5.6-terra`,
1015 `gpt-5.6-luna`
1016 </td>
1011 <td>1017 <td>
1012 `low`, `high`, `original`,1018 `low`, `high`, `original`,
1013 `auto`1019 `auto`
1014 </td>1020 </td>
1015 <td>1021 <td>
1016 `low` and `high` can resize images under their1022 `low` fits within 512 × 512 pixels. `high` fits
1017 finite limits. `original` preserves the input dimensions and1023 within 2048 × 2048 pixels and 2,500 patches. `original` fits
1018 does not resize the image to a pixel-dimension or patch-budget limit.1024 within 65,535 × 65,535 pixels, with no patch-budget limit.
1019 `auto` and omitted `detail` use the same sizing1025 `auto` uses the same sizing behavior as `original`.
1020 behavior as `original`. Request payload and other image-input
1021 limits still apply.
1022 </td>1026 </td>
1023 </tr>1027 </tr>
1024 <tr>1028 <tr>
1030 `auto`1034 `auto`
1031 </td>1035 </td>
1032 <td>1036 <td>
1033 `high` allows up to 2,500 patches or a 2048-pixel maximum1037 `low` fits within 512 × 512 pixels. `high` allows up
1034 dimension. `original` allows up to 10,000 patches or a1038 to 2,500 patches and a 2048-pixel maximum dimension. `original`
1035 6000-pixel maximum dimension. If either limit is exceeded, we resize the1039 allows up to 10,000 patches and a 6000-pixel maximum dimension. Both
1036 image while preserving aspect ratio to fit within the lesser of those two1040 limits apply. `auto` uses the same sizing behavior as
1037 constraints for the selected detail level. `auto` and omitted1041 `original`.
1038 `detail` use the same sizing behavior as
1039 `original`. [Full resizing details
1040 below.](#patch-based-image-tokenization)
1041 </td>1042 </td>
1042 </tr>1043 </tr>
1043 <tr>1044 <tr>
1044 <td>1045 <td>
1045 `gpt-5.4`1046 `gpt-5.4`, `gpt-5.4-mini`, `gpt-5.4-nano`
1046 </td>1047 </td>
1047 <td>1048 <td>
1048 `low`, `high`, `original`,1049 `low`, `high`, `original`,
1049 `auto`1050 `auto`
1050 </td>1051 </td>
1051 <td>1052 <td>
1052 `high` allows up to 2,500 patches or a 2048-pixel maximum1053 `low` uses a 2048-pixel maximum dimension and a 6,144-patch
1053 dimension. `original` allows up to 10,000 patches or a1054 budget, so it can use more tokens than `high`.
1054 6000-pixel maximum dimension. If either limit is exceeded, we resize the1055 `high` allows up to 2,500 patches and a 2048-pixel maximum
1055 image while preserving aspect ratio to fit within the lesser of those two1056 dimension. `original` allows up to 10,000 patches and a
1056 constraints for the selected detail level. `auto` and omitted1057 6000-pixel maximum dimension. Both limits apply. `auto` uses
1057 `detail` use the same sizing behavior as1058 the same sizing behavior as `high`.
1058 `high`. [Full resizing details
1059 below.](#patch-based-image-tokenization)
1060 </td>1059 </td>
1061 </tr>1060 </tr>
1062 <tr>1061 <tr>
1063 <td>1062 <td>
1064 `gpt-5.4-mini`, `gpt-5.4-nano`,1063 `gpt-5.2`, `gpt-4.1-mini`
1065 `gpt-5-mini`, `gpt-5-nano`, `gpt-5.2`,
1066 `gpt-5.3-codex`, `gpt-5-codex-mini`,
1067 `gpt-5.1-codex-mini`, `gpt-5.2-codex`,
1068 `gpt-5.2-chat-latest`, `o4-mini`, and the
1069 `gpt-4.1-mini` and `gpt-4.1-nano` 2025-04-14
1070 snapshot variants
1071 </td>1064 </td>
1072 <td>1065 <td>
1073 `low`, `high`, `auto`1066 `low`, `high`, `auto`
1074 </td>1067 </td>
1075 <td>1068 <td>
1076 `high` allows up to 1,536 patches or a 2048-pixel maximum1069 These detail levels use the same sizing limits: a 2048-pixel maximum
1077 dimension. If either limit is exceeded, we resize the image while1070 dimension and a 6,144-patch budget. `original` is not
1078 preserving aspect ratio to fit within the lesser of those two constraints.1071 supported.
1079 [Full resizing details below.](#patch-based-image-tokenization)
1080 </td>1072 </td>
1081 </tr>1073 </tr>
1082 <tr>1074 <tr>
1083 <td>1075 <td>
1084 `GPT-4o`, `GPT-4.1`, `GPT-4o-mini`,1076 `gpt-5.1`, `gpt-4.1`, `gpt-4o`,
1085 `computer-use-preview`, and o-series models except1077 `gpt-4o-mini`
1086 `o4-mini`
1087 </td>1078 </td>
1088 <td>1079 <td>
1089 `low`, `high`, `auto`1080 `low`, `high`, `auto`
1090 </td>1081 </td>
1091 <td>1082 <td>
1092 Use tile-based resizing behavior. See 1083 `low` uses a fixed token count. `high` and
1093 [the detailed behavior below](#gpt-4o-gpt-41-gpt-4o-mini-cua-and-o-series-except-o4-mini)1084 `auto` use the
1085 [tile-based sizing rules](#tile-based-image-tokenization).
1094 </td>1086 </td>
1095 </tr>1087 </tr>
1096</table>1088</table>
1097 1089
1098## Calculating costs1090## Calculating costs
1099 1091
1100Image inputs are metered and charged in token units similar to text inputs. How images are converted to text token inputs varies based on the model. You can find a vision pricing calculator in the FAQ section of the [pricing page](https://openai.com/api/pricing/).1092Vision models convert image inputs into billable input tokens. The calculator and patch/tile rules in this section cover vision-model inputs, not GPT Image generation or editing. See [GPT Image model inputs](#gpt-image-model-inputs) for that separate pricing.
1093
1094Image tokens also count toward your [tokens per minute (TPM) limits](https://developers.openai.com/api/docs/guides/rate-limits). The calculator estimates one image at standard input rates; it does not include the rest of your prompt or model output.
1095
1096### Image input cost calculator
1097
1098Estimate input tokens and cost for one image.
1101 1099
1102### Patch-based image tokenization1100### Patch-based image tokenization
1103 1101
1104Some models tokenize images by covering them with 32px x 32px patches. Many model and detail-level combinations define a maximum patch budget. The token cost of an image is determined as follows:1102Some models tokenize images by covering them with 32px x 32px patches. Many model and detail-level combinations define a maximum patch budget. First, the API fits the image within the selected detail level's pixel-dimension limit, preserving aspect ratio and rounding to integer pixels without enlarging smaller images. The token cost is then determined as follows:
1105 1103
1106A. Compute how many 32px x 32px patches are needed to cover the original image. A patch may extend beyond the image boundary.1104A. Compute how many 32px x 32px patches are needed to cover the image after applying the pixel-dimension limit. A patch may extend beyond the image boundary.
1107 1105
1108```1106```
1109original_patch_count = ceil(width/32)×ceil(height/32)1107patch_count = ceil(width/32)×ceil(height/32)
1110```1108```
1111 1109
1112For GPT-5.6 models with `detail` set to `original` or `auto`, the service uses the original patch count without resizing the image to a patch budget or pixel-dimension limit. This means large images can use more input tokens than they did with earlier models. To control token use and latency, resize the image before sending it or select `low` or `high` detail.1110GPT-5.6 Sol, Terra, and Luna have no patch-budget limit for `original` or `auto`. After applying their pixel-dimension limit, skip the patch-budget resizing step. Large images can therefore use more tokens than with earlier models; resize them before sending or select `low` or `high` to control token use.
1113 1111
1114B. If the original image would exceed the model's patch budget, scale it down proportionally until it fits within that budget. Then adjust the scale so the final resized image stays within budget after converting to integer pixel dimensions and computing patch coverage.1112B. When a patch budget applies and the image exceeds it, scale the image down proportionally. Adjust the scale to stay within budget after converting to integer pixel dimensions and computing patch coverage. Keep full precision until calculating the final dimensions.
1115 1113
1116```1114```
1117shrink_factor = sqrt((32^2 * patch_budget) / (width * height))1115shrink_factor = sqrt((32^2 * patch_budget) / (width * height))
1121)1119)
1122```1120```
1123 1121
1124C. Convert the adjusted scale into integer pixel dimensions, then compute the number of patches needed to cover the resized image. This resized patch count is the image-token count before applying the model multiplier, and it is capped by the model's patch budget.1122C. If step B resized the image, round down the final scaled width and height to integer pixels. Compute the patches needed to cover the resulting image. This is the image-token count before applying the model multiplier. When a patch budget applies, this count stays within that budget.
1125 1123
1126```1124```
1127resized_patch_count = ceil(resized_width/32)×ceil(resized_height/32)1125resized_patch_count = ceil(resized_width/32)×ceil(resized_height/32)
1128```1126```
1129 1127
1130D. Apply a multiplier based on the model to get the total tokens:1128D. Multiply the patch count by the model's multiplier and round up to get the billable image input tokens. Apply the model's input price to those tokens once; the multiplier does not apply to other prompt tokens or to the price again.
1131 1129
1132| Model | Multiplier |1130| Model | Multiplier |
1133| --------------- | ---------- |1131| -------------------------------------- | ---------- |
1134| `gpt-5.6-sol` | 1.2 |1132| `gpt-5.6-sol` | 1.2 |
1135| `gpt-5.6-terra` | 1.2 |1133| `gpt-5.6-terra` | 1.2 |
1136| `gpt-5.6-luna` | 1.2 |1134| `gpt-5.6-luna` | 1.2 |
1137| `gpt-5.5` | 1.2 |1135| `gpt-5.5` | 1.2 |
1136| `gpt-5.4` | 1.2 |
1138| `gpt-5.4-mini` | 1.2 |1137| `gpt-5.4-mini` | 1.2 |
1139| `gpt-5.4-nano` | 1.2 |1138| `gpt-5.4-nano` | 1.2 |
1140| `gpt-5-mini` | 1.2 |1139| `gpt-5.2` | 1.2 |
1141| `gpt-5-nano` | 1.5 |1140| `gpt-5-mini`\* | 1.2 |
1142| `gpt-4.1-mini*` | 1.62 |1141| `gpt-5-nano`\* | 1.5 |
1143| `gpt-4.1-nano*` | 2.46 |1142| `gpt-4.1-mini` | 1.62 |
1144| `o4-mini` | 1.72 |1143| `gpt-4.1-nano`\* (2025-04-14 snapshot) | 2.46 |
1145 1144| `o4-mini`\* | 1.72 |
1146_For `gpt-4.1-mini` and `gpt-4.1-nano`, this applies to the 2025-04-14 snapshot variants._
1147
1148**Cost calculation examples for a model with a 1,536-patch budget**
1149
1150- A 1024 × 1024 image has a post-resize patch count of **1024**
1151 - A. `original_patch_count = ceil(1024 / 32) * ceil(1024 / 32) = 32 * 32 = 1024`
1152 - B. `1024` is below the `1,536` patch budget, so no resize is needed.
1153 - C. `resized_patch_count = 1024`
1154 - Resized patch count before the model multiplier: `1024`
1155 - Multiply by the model's token multiplier to get the billed token units.
1156- A 1800 × 2400 image has a post-resize patch count of **1452**
1157 - A. `original_patch_count = ceil(1800 / 32) * ceil(2400 / 32) = 57 * 75 = 4275`
1158 - B. `4275` exceeds the `1,536` patch budget, so we first compute `shrink_factor = sqrt((32^2 * 1536) / (1800 * 2400)) = 0.603`.
1159 - We then adjust that scale so the final integer pixel dimensions stay within budget after patch counting: `adjusted_shrink_factor = 0.603 * min(floor(1800 * 0.603 / 32) / (1800 * 0.603 / 32), floor(2400 * 0.603 / 32) / (2400 * 0.603 / 32)) = 0.586`.
1160 - Resized image dimensions: `1056 × 1408`
1161 - C. `resized_patch_count = ceil(1056 / 32) * ceil(1408 / 32) = 33 * 44 = 1452`
1162 - Resized patch count before the model multiplier: `1452`
1163 - Multiply by the model's token multiplier to get the billed token units.
1164 1145
1165### Tile-based image tokenization1146_For `gpt-4.1-mini`, this applies to the 2025-04-14 snapshot._
1166 1147
1167#### GPT-4o, GPT-4.1, GPT-4o-mini, CUA, and o-series (except o4-mini)1148\* Deprecated and scheduled for shutdown. See the [deprecation schedule](https://developers.openai.com/api/docs/deprecations) for dates and replacements. These models aren't included in the calculator or the model sizing table above.
1168 1149
1169The token cost of an image is determined by two factors: size and detail.1150**Cost calculation examples for `gpt-5.4` with `detail: high`**
1170 1151
1171Any image with `"detail": "low"` costs a set, base number of tokens. This amount varies by model. To calculate the cost of an image with `"detail": "high"`, we do the following:1152This combination uses a 2048-pixel maximum dimension, a 2,500-patch budget, and a 1.2× multiplier.
1172 1153
1173- Scale to fit in a 2048px x 2048px square, maintaining original aspect ratio1154- A 1024 × 1024 image needs `32 × 32 = 1024` patches. No resizing is needed. The billable image input is `ceil(1024 × 1.2) = 1229` tokens.
1174- Scale so that the image's shortest side is 768px long1155- A 2048 × 2048 image initially needs `64 × 64 = 4096` patches. The patch budget reduces it to 1600 × 1600 pixels, or `50 × 50 = 2500` patches. The estimate is `ceil(2500 × 1.2) = 3000` tokens.
1175- Count the number of 512px squares in the image. Each square costs a set amount of tokens, shown below.1156
1176- Add the base tokens to the total1157Floating-point rounding in billing can make the final count differ from the estimate by one token.
1158
1159### Tile-based image tokenization
1160
1161<a id="gpt-4o-gpt-41-gpt-4o-mini-cua-and-o-series-except-o4-mini"></a>
1162
1163The models in this table use a base token count plus tokens for image tiles:
1177 1164
1178| Model | Base tokens | Tile tokens |1165| Model | Base tokens | Tile tokens |
1179| ------------------------------ | ----------- | ----------- |1166| -------------------------- | ----------- | ----------- |
1180| `gpt-5`, `gpt-5-chat-latest` | 70 | 140 |1167| `gpt-5.1` | 70 | 140 |
1181| `gpt-4o`, `gpt-4.1`, `gpt-4.5` | 85 | 170 |1168| `gpt-5`\* | 70 | 140 |
1169| `gpt-4o`, `gpt-4.1` | 85 | 170 |
1182| `gpt-4o-mini` | 2833 | 5667 |1170| `gpt-4o-mini` | 2833 | 5667 |
1183| `o1`, `o1-pro`, `o3` | 75 | 150 |1171| `o1`\*, `o1-pro`\*, `o3`\* | 75 | 150 |
1184| `computer-use-preview` | 65 | 129 |1172
1173\* Deprecated and scheduled for shutdown. See the [deprecation schedule](https://developers.openai.com/api/docs/deprecations) for dates and replacements. These models aren't included in the calculator or the model sizing table above.
1174
1175With `"detail": "low"`, an image costs only the model's base tokens, regardless of dimensions. With `"detail": "high"` or `"detail": "auto"`:
1185 1176
1186### GPT Image 11177- Scale down to fit in a 2048px x 2048px square, maintaining aspect ratio. Smaller images are not enlarged.
1178- If the shortest side exceeds 768px, scale it down to 768px and round down the other dimension.
1179- Count the 512px squares needed to cover the image. Each square uses the model's tile tokens.
1180- Add the model's base tokens to the tile tokens.
1187 1181
1188For GPT Image 1, we calculate the cost of an image input the same way as described above, except that we scale down the image so that the shortest side is 512px instead of 768px.1182### GPT Image model inputs
1189The price depends on the dimensions of the image and the [input fidelity](https://developers.openai.com/api/docs/guides/image-generation?image-generation-model=gpt-image-1#image-input-fidelity).1183
1184GPT Image models use separate image-token pricing for generation and editing. The vision calculator does not estimate their input or output costs. For current rates, see [image generation pricing](https://developers.openai.com/api/docs/pricing#image-generation); for generation and editing workflows, see the [Image generation guide](https://developers.openai.com/api/docs/guides/image-generation).
1185
1186#### GPT Image 1
1187
1188The following input-token rules apply to `gpt-image-1`. Use tile-based image sizing, but scale the shortest side down to 512px instead of 768px. Token use depends on the image dimensions and the `input_fidelity` parameter in the [Images API](https://developers.openai.com/api/reference/resources/images/methods/edit).
1190 1189
1191When input fidelity is set to low, the base cost is 65 image tokens, and each tile costs 129 image tokens.1190When input fidelity is set to low, the base cost is 65 image tokens, and each tile costs 129 image tokens.
1192When using high input fidelity, we add a set number of tokens based on the image's aspect ratio in addition to the image tokens described above.1191When using high input fidelity, we add a set number of tokens based on the image's aspect ratio in addition to the image tokens described above.
1198 1197
1199## Limitations1198## Limitations
1200 1199
1201While models with vision capabilities are powerful and can be used in many situations, it's important to understand the limitations of these models. Here are some known limitations:1200Vision models can make mistakes. Account for these limitations when designing your application:
1202 1201
1203- **Medical images**: The model is not suitable for interpreting specialized medical images like CT scans and shouldn't be used for medical advice.1202- **Medical images**: The model is not suitable for interpreting specialized medical images like CT scans and shouldn't be used for medical advice.
1204- **Non-English**: The model may not perform optimally when handling images with text of non-Latin alphabets, such as Japanese or Korean.1203- **Non-English**: The model may not perform optimally when handling images with text of non-Latin alphabets, such as Japanese or Korean.
1208- **Spatial reasoning**: The model struggles with tasks requiring precise spatial localization, such as identifying chess positions.1207- **Spatial reasoning**: The model struggles with tasks requiring precise spatial localization, such as identifying chess positions.
1209- **Accuracy**: The model may generate incorrect descriptions or captions in certain scenarios.1208- **Accuracy**: The model may generate incorrect descriptions or captions in certain scenarios.
1210- **Image shape**: The model struggles with panoramic and fisheye images.1209- **Image shape**: The model struggles with panoramic and fisheye images.
1211- **Metadata and resizing**: The model doesn't process original file names or metadata. `low` and `high` detail, and models with finite image budgets, may resize images before analysis. GPT-5.6 models preserve the input dimensions with `original` and `auto` detail.1210- **Metadata and resizing**: The model doesn't process original file names or metadata. Images may be resized before analysis, including with `original` detail. See [Model sizing behavior](#model-sizing-behavior) for the limits that apply to each model.
1212- **Counting**: The model may give approximate counts for objects in images.1211- **Counting**: The model may give approximate counts for objects in images.
1213- **CAPTCHAs**: For safety reasons, our system blocks the submission of CAPTCHAs.1212- **CAPTCHAs**: For safety reasons, our system blocks the submission of CAPTCHAs.
1214
1215
1216We process images at the token level, so each image we process counts towards your tokens per minute (TPM) limit.
1217
1218For the most precise and up-to-date estimates for image processing, please use our image pricing calculator available [here](https://openai.com/api/pricing/).