2 2
3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.3> For the complete documentation index, see [llms.txt](/llms.txt). Markdown versions of documentation pages are available by appending `.md` to the page URL.
4 4
5Choosing the right model, whether [`gpt-6-astra`](https://developers.openai.com/api/docs/models/gpt-6-astra) or a smaller option like [`gpt-5.6-terra`](https://developers.openai.com/api/docs/models/gpt-5.6-terra), requires balancing **accuracy**, **latency**, and **cost**. This guide explains key principles to help you make informed decisions, along with a practical example.5## Meet the models
6 6
7## Core principles7Availability, tools, reasoning settings, and usage limits differ by product and
8model version. Check the [models available in ChatGPT](https://developers.openai.com/codex/models) or the
9[API model catalog](https://developers.openai.com/api/docs/models).
8 10
9The principles for model selection are simple:11## Find the right model for your workflow
10 12
11- **Optimize for accuracy first:** Optimize for accuracy until you hit your accuracy target.13Choose your work and task and get a recommendation.
12- **Optimize for cost and latency second:** Then aim to maintain accuracy with the cheapest, fastest model possible.
13 14
14### 1. Focus on accuracy first15## How to think about models and reasoning effort
15 16
16Begin by setting a clear accuracy goal for your use case, where you're clear on the accuracy that would be "good enough" for this use case to go to production. You can accomplish this through:
17 17
18- **Setting a clear accuracy target:** Identify what your target accuracy statistic is going to be.
19 - For example, 90% of customer service calls need to be triaged correctly at the first interaction.
20- **Developing an evaluation dataset:** Create a dataset that allows you to measure the model's performance against these goals.
21 - To extend the example above, capture 100 interaction examples where we have what the user asked for, what the LLM triaged them to, what the correct triage should be, and whether this was correct or not.
22- **Using the most powerful model to optimize:** Start with the most capable model available to achieve your accuracy targets. Log all responses so we can use them for distillation of a smaller model.
23 - Use retrieval-augmented generation to optimize for accuracy
24 - Use fine-tuning to optimize for consistency and behavior
25 18
26During this process, collect prompt and completion pairs for use in evaluations, few-shot learning, or fine-tuning. This practice, known as **prompt baking**, helps you produce high-quality examples for future use.19 Luna is the most cost-efficient model, while Astra is our state-of-the-art,
20 most powerful model. If cost and latency aren't a concern, you can default to
21 Astra. To reduce costs or latency, use the guidance below to choose a model
22 and reasoning effort for your needs.
27 23
28For more methods and tools here, see our [Accuracy Optimization Guide](https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy).
29 24
30#### Setting a realistic accuracy target
31 25
32Calculate a realistic accuracy target by evaluating the financial impact of model decisions. For example, in a fake news classification scenario:
33 26
34- **Correctly classified news:** If the model classifies it correctly, it saves you the cost of a human reviewing it - let's assume **$50**.
35- **Incorrectly classified news:** If it falsely classifies a safe article or misses a fake news article, it may trigger a review process and possible complaint, which might cost us **$300**.
36 27
37Our news classification example would need **85.8%** accuracy to cover costs, so targeting 90% or more ensures an overall return on investment. Use these calculations to set an effective accuracy target based on your specific cost structures.281. **Luna · Low**
38 29
39### 2. Optimize cost and latency30 Fine-grained edits, well-scoped problem-solving, and simple data extraction.
40 31
41Cost and latency are considered secondary because if the model can’t hit your accuracy target then these concerns are moot. However, once you’ve got a model that works for your use case, you can take one of two approaches:322. **Luna · Medium**
42 33
43- **Compare with a smaller model zero- or few-shot:** Swap out the model for a smaller, cheaper one and test whether it maintains accuracy at the lower cost and latency point.34 Creating from clear briefs and making coordinated updates to existing work.
44- **Model distillation:** Fine-tune a smaller model using the data gathered during accuracy optimization.
45 35
46Cost and latency are typically interconnected; reducing tokens and requests generally leads to faster processing.363. **Luna · Extra high**
47 37
48The main strategies to consider here are:38 Finding current context across multiple apps, prioritizing work, and solving problems with clear constraints.
49 39
50- **Reduce requests:** Limit the number of necessary requests to complete tasks.404. **Sol · Low**
51- **Minimize tokens:** Lower the number of input tokens and optimize for shorter model outputs.
52- **Select a smaller model:** Use models that balance reduced costs and latency with maintained accuracy.
53 41
54To dive deeper into these, please refer to our guide on [latency optimization](https://developers.openai.com/api/docs/guides/latency-optimization).42 Focused writing and editing, fact-checking, and straightforward work in apps.
55 43
56#### Exceptions to the rule445. **Sol · Medium**
57 45
58Clear exceptions exist for these principles. If your use case is extremely cost or latency sensitive, establish thresholds for these metrics before beginning your testing, then remove the models that exceed those from consideration. Once benchmarks are set, these guidelines will help you refine model accuracy within your constraints.46 Everyday coding, research, and workflows that need judgment and completeness.
59 47
60## Practical example486. **Sol · Extra high**
61 49
62To demonstrate these principles, we'll develop a fake news classifier with the following target metrics. The experiment below uses historical GPT-4o-family results to show the workflow; for current evaluations, start with [`gpt-6-astra`](https://developers.openai.com/api/docs/models/gpt-6-astra) and compare against smaller or fine-tuned models.50 Deeper analysis, thorough verification, and careful review of documents, data, and code.
63 51
64- **Accuracy:** Achieve 90% correct classification527. **Astra · Low**
65- **Cost:** Spend less than $5 per 1,000 articles
66- **Latency:** Maintain processing time under 2 seconds per article
67 53
68### Experiments54 Concise writing and content adaptation that preserve facts and nuance.
69 55
70We ran three experiments to reach our goal:568. **Astra · Medium**
71 57
721. **Zero-shot:** Used `GPT-4o` with a basic prompt for 1,000 records, but missed the accuracy target.58 Ambitious projects that need broad context, reliable interactions, and complete results.
732. **Few-shot learning:** Included 5 few-shot examples, meeting the accuracy target but exceeding cost due to more prompt tokens.
743. **Fine-tuned model:** Fine-tuned `GPT-4o-mini` with 1,000 labeled examples, meeting all targets with similar latency and accuracy but significantly lower costs.
75 59
76| ID | Method | Accuracy | Accuracy target | Cost | Cost target | Avg. latency | Latency target |609. **Astra · Extra high**
77| --- | --------------------------------------- | -------- | --------------- | ------ | ----------- | ------------ | -------------- |
78| 1 | gpt-4o zero-shot | 84.5% | | $1.72 | | < 1s | |
79| 2 | gpt-4o few-shot (n=5) | 91.5% | ✓ | $11.92 | | < 1s | ✓ |
80| 3 | gpt-4o-mini fine-tuned w/ 1000 examples | 91.5% | ✓ | $0.21 | ✓ | < 1s | ✓ |
81 61
82## Conclusion62 Demanding analysis and complex deliverables with exacting requirements.
83 63
84By switching from `gpt-4o` to `gpt-4o-mini` with fine-tuning, we achieved **equivalent performance for less than 2%** of the cost, using only 1,000 labeled examples.64### Experiment
85 65
86This process is important - you often can’t jump right to fine-tuning because you don’t know whether fine-tuning is the right tool for the optimization you need, or you don’t have enough labeled examples. Start with [`gpt-6-astra`](https://developers.openai.com/api/docs/models/gpt-6-astra) to establish your accuracy target, then test smaller or fine-tuned models when cost and latency matter.66Treat the guidance on this page as a starting point. The best way to find the right
67model for your workflow is to experiment with different models and reasoning
68settings to see what works.
69
70Start by considering:
71
72- **How often does your workflow run?** A frequent automation makes usage and cost
73 add up faster than an occasional project.
74- **How quickly do you need the result?** A task you're waiting on may need a
75 faster setting than one that runs overnight.
76- **How will you use the output?** A draft for your review may need less polish
77 than something you'll share externally.
78- **How important is the quality of the result?** Depending on your use case or
79 industry, you might want to use a stronger model to put an emphasis on quality.
80
81If you can, experiment using the same inputs to compare results and keep the
82lightest setting that meets your quality bar.