Alpha
Graders
Run grader
post /fine_tuning/alpha/graders/run
Run a grader.
Body Parameters
-
grader: StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreThe grader used for the fine-tuning job.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
input: stringThe input text. This may include template strings.
-
name: stringThe name of the grader.
-
operation: "eq" or "ne" or "like" or "ilike"The string check operation to perform. One of
eq,ne,like, orilike.-
"eq" -
"ne" -
"like" -
"ilike"
-
-
reference: stringThe reference text. This may include template strings.
-
type: "string_check"The object type, which is always
string_check."string_check"
-
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
evaluation_metric: "cosine" or "fuzzy_match" or "bleu" or 8 moreThe evaluation metric to use. One of
cosine,fuzzy_match,bleu,gleu,meteor,rouge_1,rouge_2,rouge_3,rouge_4,rouge_5, orrouge_l.-
"cosine" -
"fuzzy_match" -
"bleu" -
"gleu" -
"meteor" -
"rouge_1" -
"rouge_2" -
"rouge_3" -
"rouge_4" -
"rouge_5" -
"rouge_l"
-
-
input: stringThe text being graded.
-
name: stringThe name of the grader.
-
reference: stringThe text being graded against.
-
type: "text_similarity"The type of grader.
"text_similarity"
-
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
name: stringThe name of the grader.
-
source: stringThe source code of the python script.
-
type: "python"The object type, which is always
python."python"
-
image_tag: optional stringThe image tag to use for the python script.
-
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
input: array of object { content, role, type }The input messages evaluated by the grader. Supports text, output text, input image, and input audio content blocks, and may include template strings.
-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
model: stringThe model to use for the evaluation.
-
name: stringThe name of the grader.
-
type: "score_model"The object type, which is always
score_model."score_model"
-
range: optional array of numberThe range of the score. Defaults to
[0, 1]. -
sampling_params: optional object { max_completions_tokens, reasoning_effort, seed, 2 more }The sampling parameters for the model.
-
max_completions_tokens: optional number or nullThe maximum number of tokens the grader model may generate in its response.
-
reasoning_effort: optional ReasoningEffort or nullConstrains effort on reasoning for reasoning models. Currently supported values are
none,minimal,low,medium,high,xhigh, andmax. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Not all reasoning models support every value. See the reasoning guide for model-specific support.-
"none" -
"minimal" -
"low" -
"medium" -
"high" -
"xhigh" -
"max"
-
-
seed: optional number or nullA seed value to initialize the randomness, during sampling.
-
temperature: optional number or nullA higher temperature increases randomness in the outputs.
-
top_p: optional number or nullAn alternative to temperature for nucleus sampling; 1.0 includes all tokens.
-
-
-
MultiGrader object { calculate_output, graders, name, type }A MultiGrader object combines the output of multiple graders to produce a single score.
-
calculate_output: stringA formula to calculate the output based on grader results.
-
graders: StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreA StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
LabelModelGrader object { input, labels, model, 3 more }A LabelModelGrader object which uses a model to assign labels to each item in the evaluation.
-
input: array of object { content, role, type }-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
labels: array of stringThe labels to assign to each item in the evaluation.
-
model: stringThe model to use for the evaluation. Must support structured outputs.
-
name: stringThe name of the grader.
-
passing_labels: array of stringThe labels that indicate a passing result. Must be a subset of labels.
-
type: "label_model"The object type, which is always
label_model."label_model"
-
-
-
name: stringThe name of the grader.
-
type: "multi"The object type, which is always
multi."multi"
-
-
-
model_sample: stringThe model sample to be evaluated. This value will be used to populate the
samplenamespace. See the guide for more details. Theoutput_jsonvariable will be populated if the model sample is a valid JSON string. -
item: optional unknownThe dataset item provided to the grader. This will be used to populate the
itemnamespace. See the guide for more details.
Returns
-
metadata: object { errors, execution_time, name, 4 more }-
errors: object { formula_parse_error, invalid_variable_error, model_grader_parse_error, 11 more }-
formula_parse_error: boolean -
invalid_variable_error: boolean -
model_grader_parse_error: boolean -
model_grader_refusal_error: boolean -
model_grader_server_error: boolean -
model_grader_server_error_details: string or null -
other_error: boolean -
python_grader_runtime_error: boolean -
python_grader_runtime_error_details: string or null -
python_grader_server_error: boolean -
python_grader_server_error_type: string or null -
sample_parse_error: boolean -
truncated_observation_error: boolean -
unresponsive_reward_error: boolean
-
-
execution_time: number -
name: string -
sampled_model_name: string or null -
scores: map[unknown] -
token_usage: number or null -
type: string
-
-
model_grader_token_usage_per_model: map[unknown] -
reward: number -
sub_rewards: map[unknown]
Example
curl https://api.openai.com/v1/fine_tuning/alpha/graders/run \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"grader": {
"input": "input",
"name": "name",
"operation": "eq",
"reference": "reference",
"type": "string_check"
},
"model_sample": "model_sample"
}'
Response
{
"metadata": {
"errors": {
"formula_parse_error": true,
"invalid_variable_error": true,
"model_grader_parse_error": true,
"model_grader_refusal_error": true,
"model_grader_server_error": true,
"model_grader_server_error_details": "model_grader_server_error_details",
"other_error": true,
"python_grader_runtime_error": true,
"python_grader_runtime_error_details": "python_grader_runtime_error_details",
"python_grader_server_error": true,
"python_grader_server_error_type": "python_grader_server_error_type",
"sample_parse_error": true,
"truncated_observation_error": true,
"unresponsive_reward_error": true
},
"execution_time": 0,
"name": "name",
"sampled_model_name": "sampled_model_name",
"scores": {
"foo": "bar"
},
"token_usage": 0,
"type": "type"
},
"model_grader_token_usage_per_model": {
"foo": "bar"
},
"reward": 0,
"sub_rewards": {
"foo": "bar"
}
}
Score an audio response
curl -X POST https://api.openai.com/v1/fine_tuning/alpha/graders/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"grader": {
"type": "score_model",
"name": "Audio clarity grader",
"input": [
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Listen to the clip and return a confidence score from 0 to 1 that the speaker said: {{item.target_phrase}}"
},
{
"type": "input_audio",
"input_audio": {
"data": "{{item.audio_clip_b64}}",
"format": "mp3"
}
}
]
}
],
"model": "gpt-audio",
"sampling_params": {
"temperature": 0.2,
"top_p": 1,
"seed": 123
}
},
"item": {
"target_phrase": "Please deliver the package on Tuesday",
"audio_clip_b64": "<base64-encoded mp3>"
},
"model_sample": "Please deliver the package on Tuesday"
}'
Score an image caption
curl -X POST https://api.openai.com/v1/fine_tuning/alpha/graders/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"grader": {
"type": "score_model",
"name": "Image caption grader",
"input": [
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Score how well the provided caption matches the image on a 0-1 scale. Only return the score.\n\nCaption: {{sample.output_text}}"
},
{
"type": "input_image",
"image_url": "https://example.com/dog-catching-ball.png",
"file_id": null,
"detail": "high"
}
]
}
],
"model": "gpt-5-mini",
"sampling_params": {
"temperature": 0.2
}
},
"item": {
"expected_caption": "A golden retriever jumps to catch a tennis ball"
},
"model_sample": "A dog leaps to grab a tennis ball mid-air"
}'
Score text alignment
curl -X POST https://api.openai.com/v1/fine_tuning/alpha/graders/run \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"grader": {
"type": "score_model",
"name": "Example score model grader",
"input": [
{
"role": "user",
"content": [
{
"type": "input_text",
"text": "Score how close the reference answer is to the model answer on a 0-1 scale. Return only the score.\n\nReference answer: {{item.reference_answer}}\n\nModel answer: {{sample.output_text}}"
}
]
}
],
"model": "gpt-5-mini",
"sampling_params": {
"temperature": 1,
"top_p": 1,
"seed": 42
}
},
"item": {
"reference_answer": "fuzzy wuzzy was a bear"
},
"model_sample": "fuzzy wuzzy was a bear"
}'
Response
{
"reward": 1.0,
"metadata": {
"name": "Example score model grader",
"type": "score_model",
"errors": {
"formula_parse_error": false,
"sample_parse_error": false,
"truncated_observation_error": false,
"unresponsive_reward_error": false,
"invalid_variable_error": false,
"other_error": false,
"python_grader_server_error": false,
"python_grader_server_error_type": null,
"python_grader_runtime_error": false,
"python_grader_runtime_error_details": null,
"model_grader_server_error": false,
"model_grader_refusal_error": false,
"model_grader_parse_error": false,
"model_grader_server_error_details": null
},
"execution_time": 4.365238428115845,
"scores": {},
"token_usage": {
"prompt_tokens": 190,
"total_tokens": 324,
"completion_tokens": 134,
"cached_tokens": 0
},
"sampled_model_name": "gpt-4o-2024-08-06"
},
"sub_rewards": {},
"model_grader_token_usage_per_model": {
"gpt-4o-2024-08-06": {
"prompt_tokens": 190,
"total_tokens": 324,
"completion_tokens": 134,
"cached_tokens": 0
}
}
}
Validate grader
post /fine_tuning/alpha/graders/validate
Validate a grader.
Body Parameters
-
grader: StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreThe grader used for the fine-tuning job.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
input: stringThe input text. This may include template strings.
-
name: stringThe name of the grader.
-
operation: "eq" or "ne" or "like" or "ilike"The string check operation to perform. One of
eq,ne,like, orilike.-
"eq" -
"ne" -
"like" -
"ilike"
-
-
reference: stringThe reference text. This may include template strings.
-
type: "string_check"The object type, which is always
string_check."string_check"
-
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
evaluation_metric: "cosine" or "fuzzy_match" or "bleu" or 8 moreThe evaluation metric to use. One of
cosine,fuzzy_match,bleu,gleu,meteor,rouge_1,rouge_2,rouge_3,rouge_4,rouge_5, orrouge_l.-
"cosine" -
"fuzzy_match" -
"bleu" -
"gleu" -
"meteor" -
"rouge_1" -
"rouge_2" -
"rouge_3" -
"rouge_4" -
"rouge_5" -
"rouge_l"
-
-
input: stringThe text being graded.
-
name: stringThe name of the grader.
-
reference: stringThe text being graded against.
-
type: "text_similarity"The type of grader.
"text_similarity"
-
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
name: stringThe name of the grader.
-
source: stringThe source code of the python script.
-
type: "python"The object type, which is always
python."python"
-
image_tag: optional stringThe image tag to use for the python script.
-
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
input: array of object { content, role, type }The input messages evaluated by the grader. Supports text, output text, input image, and input audio content blocks, and may include template strings.
-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
model: stringThe model to use for the evaluation.
-
name: stringThe name of the grader.
-
type: "score_model"The object type, which is always
score_model."score_model"
-
range: optional array of numberThe range of the score. Defaults to
[0, 1]. -
sampling_params: optional object { max_completions_tokens, reasoning_effort, seed, 2 more }The sampling parameters for the model.
-
max_completions_tokens: optional number or nullThe maximum number of tokens the grader model may generate in its response.
-
reasoning_effort: optional ReasoningEffort or nullConstrains effort on reasoning for reasoning models. Currently supported values are
none,minimal,low,medium,high,xhigh, andmax. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Not all reasoning models support every value. See the reasoning guide for model-specific support.-
"none" -
"minimal" -
"low" -
"medium" -
"high" -
"xhigh" -
"max"
-
-
seed: optional number or nullA seed value to initialize the randomness, during sampling.
-
temperature: optional number or nullA higher temperature increases randomness in the outputs.
-
top_p: optional number or nullAn alternative to temperature for nucleus sampling; 1.0 includes all tokens.
-
-
-
MultiGrader object { calculate_output, graders, name, type }A MultiGrader object combines the output of multiple graders to produce a single score.
-
calculate_output: stringA formula to calculate the output based on grader results.
-
graders: StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreA StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
LabelModelGrader object { input, labels, model, 3 more }A LabelModelGrader object which uses a model to assign labels to each item in the evaluation.
-
input: array of object { content, role, type }-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
labels: array of stringThe labels to assign to each item in the evaluation.
-
model: stringThe model to use for the evaluation. Must support structured outputs.
-
name: stringThe name of the grader.
-
passing_labels: array of stringThe labels that indicate a passing result. Must be a subset of labels.
-
type: "label_model"The object type, which is always
label_model."label_model"
-
-
-
name: stringThe name of the grader.
-
type: "multi"The object type, which is always
multi."multi"
-
-
Returns
-
grader: optional StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreThe grader used for the fine-tuning job.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
input: stringThe input text. This may include template strings.
-
name: stringThe name of the grader.
-
operation: "eq" or "ne" or "like" or "ilike"The string check operation to perform. One of
eq,ne,like, orilike.-
"eq" -
"ne" -
"like" -
"ilike"
-
-
reference: stringThe reference text. This may include template strings.
-
type: "string_check"The object type, which is always
string_check."string_check"
-
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
evaluation_metric: "cosine" or "fuzzy_match" or "bleu" or 8 moreThe evaluation metric to use. One of
cosine,fuzzy_match,bleu,gleu,meteor,rouge_1,rouge_2,rouge_3,rouge_4,rouge_5, orrouge_l.-
"cosine" -
"fuzzy_match" -
"bleu" -
"gleu" -
"meteor" -
"rouge_1" -
"rouge_2" -
"rouge_3" -
"rouge_4" -
"rouge_5" -
"rouge_l"
-
-
input: stringThe text being graded.
-
name: stringThe name of the grader.
-
reference: stringThe text being graded against.
-
type: "text_similarity"The type of grader.
"text_similarity"
-
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
name: stringThe name of the grader.
-
source: stringThe source code of the python script.
-
type: "python"The object type, which is always
python."python"
-
image_tag: optional stringThe image tag to use for the python script.
-
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
input: array of object { content, role, type }The input messages evaluated by the grader. Supports text, output text, input image, and input audio content blocks, and may include template strings.
-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
model: stringThe model to use for the evaluation.
-
name: stringThe name of the grader.
-
type: "score_model"The object type, which is always
score_model."score_model"
-
range: optional array of numberThe range of the score. Defaults to
[0, 1]. -
sampling_params: optional object { max_completions_tokens, reasoning_effort, seed, 2 more }The sampling parameters for the model.
-
max_completions_tokens: optional number or nullThe maximum number of tokens the grader model may generate in its response.
-
reasoning_effort: optional ReasoningEffort or nullConstrains effort on reasoning for reasoning models. Currently supported values are
none,minimal,low,medium,high,xhigh, andmax. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Not all reasoning models support every value. See the reasoning guide for model-specific support.-
"none" -
"minimal" -
"low" -
"medium" -
"high" -
"xhigh" -
"max"
-
-
seed: optional number or nullA seed value to initialize the randomness, during sampling.
-
temperature: optional number or nullA higher temperature increases randomness in the outputs.
-
top_p: optional number or nullAn alternative to temperature for nucleus sampling; 1.0 includes all tokens.
-
-
-
MultiGrader object { calculate_output, graders, name, type }A MultiGrader object combines the output of multiple graders to produce a single score.
-
calculate_output: stringA formula to calculate the output based on grader results.
-
graders: StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreA StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
LabelModelGrader object { input, labels, model, 3 more }A LabelModelGrader object which uses a model to assign labels to each item in the evaluation.
-
input: array of object { content, role, type }-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
labels: array of stringThe labels to assign to each item in the evaluation.
-
model: stringThe model to use for the evaluation. Must support structured outputs.
-
name: stringThe name of the grader.
-
passing_labels: array of stringThe labels that indicate a passing result. Must be a subset of labels.
-
type: "label_model"The object type, which is always
label_model."label_model"
-
-
-
name: stringThe name of the grader.
-
type: "multi"The object type, which is always
multi."multi"
-
-
Example
curl https://api.openai.com/v1/fine_tuning/alpha/graders/validate \
-H 'Content-Type: application/json' \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{
"grader": {
"input": "input",
"name": "name",
"operation": "eq",
"reference": "reference",
"type": "string_check"
}
}'
Response
{
"grader": {
"input": "input",
"name": "name",
"operation": "eq",
"reference": "reference",
"type": "string_check"
}
}
Example
curl https://api.openai.com/v1/fine_tuning/alpha/graders/validate \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"grader": {
"type": "string_check",
"name": "Example string check grader",
"input": "{{sample.output_text}}",
"reference": "{{item.label}}",
"operation": "eq"
}
}'
Response
{
"grader": {
"type": "string_check",
"name": "Example string check grader",
"input": "{{sample.output_text}}",
"reference": "{{item.label}}",
"operation": "eq"
}
}
Domain Types
Grader Run Response
-
GraderRunResponse object { metadata, model_grader_token_usage_per_model, reward, sub_rewards }-
metadata: object { errors, execution_time, name, 4 more }-
errors: object { formula_parse_error, invalid_variable_error, model_grader_parse_error, 11 more }-
formula_parse_error: boolean -
invalid_variable_error: boolean -
model_grader_parse_error: boolean -
model_grader_refusal_error: boolean -
model_grader_server_error: boolean -
model_grader_server_error_details: string or null -
other_error: boolean -
python_grader_runtime_error: boolean -
python_grader_runtime_error_details: string or null -
python_grader_server_error: boolean -
python_grader_server_error_type: string or null -
sample_parse_error: boolean -
truncated_observation_error: boolean -
unresponsive_reward_error: boolean
-
-
execution_time: number -
name: string -
sampled_model_name: string or null -
scores: map[unknown] -
token_usage: number or null -
type: string
-
-
model_grader_token_usage_per_model: map[unknown] -
reward: number -
sub_rewards: map[unknown]
-
Grader Validate Response
-
GraderValidateResponse object { grader }-
grader: optional StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreThe grader used for the fine-tuning job.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
input: stringThe input text. This may include template strings.
-
name: stringThe name of the grader.
-
operation: "eq" or "ne" or "like" or "ilike"The string check operation to perform. One of
eq,ne,like, orilike.-
"eq" -
"ne" -
"like" -
"ilike"
-
-
reference: stringThe reference text. This may include template strings.
-
type: "string_check"The object type, which is always
string_check."string_check"
-
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
evaluation_metric: "cosine" or "fuzzy_match" or "bleu" or 8 moreThe evaluation metric to use. One of
cosine,fuzzy_match,bleu,gleu,meteor,rouge_1,rouge_2,rouge_3,rouge_4,rouge_5, orrouge_l.-
"cosine" -
"fuzzy_match" -
"bleu" -
"gleu" -
"meteor" -
"rouge_1" -
"rouge_2" -
"rouge_3" -
"rouge_4" -
"rouge_5" -
"rouge_l"
-
-
input: stringThe text being graded.
-
name: stringThe name of the grader.
-
reference: stringThe text being graded against.
-
type: "text_similarity"The type of grader.
"text_similarity"
-
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
name: stringThe name of the grader.
-
source: stringThe source code of the python script.
-
type: "python"The object type, which is always
python."python"
-
image_tag: optional stringThe image tag to use for the python script.
-
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
input: array of object { content, role, type }The input messages evaluated by the grader. Supports text, output text, input image, and input audio content blocks, and may include template strings.
-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
model: stringThe model to use for the evaluation.
-
name: stringThe name of the grader.
-
type: "score_model"The object type, which is always
score_model."score_model"
-
range: optional array of numberThe range of the score. Defaults to
[0, 1]. -
sampling_params: optional object { max_completions_tokens, reasoning_effort, seed, 2 more }The sampling parameters for the model.
-
max_completions_tokens: optional number or nullThe maximum number of tokens the grader model may generate in its response.
-
reasoning_effort: optional ReasoningEffort or nullConstrains effort on reasoning for reasoning models. Currently supported values are
none,minimal,low,medium,high,xhigh, andmax. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Not all reasoning models support every value. See the reasoning guide for model-specific support.-
"none" -
"minimal" -
"low" -
"medium" -
"high" -
"xhigh" -
"max"
-
-
seed: optional number or nullA seed value to initialize the randomness, during sampling.
-
temperature: optional number or nullA higher temperature increases randomness in the outputs.
-
top_p: optional number or nullAn alternative to temperature for nucleus sampling; 1.0 includes all tokens.
-
-
-
MultiGrader object { calculate_output, graders, name, type }A MultiGrader object combines the output of multiple graders to produce a single score.
-
calculate_output: stringA formula to calculate the output based on grader results.
-
graders: StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreA StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
LabelModelGrader object { input, labels, model, 3 more }A LabelModelGrader object which uses a model to assign labels to each item in the evaluation.
-
input: array of object { content, role, type }-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
labels: array of stringThe labels to assign to each item in the evaluation.
-
model: stringThe model to use for the evaluation. Must support structured outputs.
-
name: stringThe name of the grader.
-
passing_labels: array of stringThe labels that indicate a passing result. Must be a subset of labels.
-
type: "label_model"The object type, which is always
label_model."label_model"
-
-
-
name: stringThe name of the grader.
-
type: "multi"The object type, which is always
multi."multi"
-
-
-