Grader Models
Domain Types
Grader Inputs
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
Label Model Grader
-
LabelModelGrader object { input, labels, model, 3 more }A LabelModelGrader object which uses a model to assign labels to each item in the evaluation.
-
input: array of object { content, role, type }-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
labels: array of stringThe labels to assign to each item in the evaluation.
-
model: stringThe model to use for the evaluation. Must support structured outputs.
-
name: stringThe name of the grader.
-
passing_labels: array of stringThe labels that indicate a passing result. Must be a subset of labels.
-
type: "label_model"The object type, which is always
label_model."label_model"
-
Multi Grader
-
MultiGrader object { calculate_output, graders, name, type }A MultiGrader object combines the output of multiple graders to produce a single score.
-
calculate_output: stringA formula to calculate the output based on grader results.
-
graders: StringCheckGrader or TextSimilarityGrader or PythonGrader or 2 moreA StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
input: stringThe input text. This may include template strings.
-
name: stringThe name of the grader.
-
operation: "eq" or "ne" or "like" or "ilike"The string check operation to perform. One of
eq,ne,like, orilike.-
"eq" -
"ne" -
"like" -
"ilike"
-
-
reference: stringThe reference text. This may include template strings.
-
type: "string_check"The object type, which is always
string_check."string_check"
-
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
evaluation_metric: "cosine" or "fuzzy_match" or "bleu" or 8 moreThe evaluation metric to use. One of
cosine,fuzzy_match,bleu,gleu,meteor,rouge_1,rouge_2,rouge_3,rouge_4,rouge_5, orrouge_l.-
"cosine" -
"fuzzy_match" -
"bleu" -
"gleu" -
"meteor" -
"rouge_1" -
"rouge_2" -
"rouge_3" -
"rouge_4" -
"rouge_5" -
"rouge_l"
-
-
input: stringThe text being graded.
-
name: stringThe name of the grader.
-
reference: stringThe text being graded against.
-
type: "text_similarity"The type of grader.
"text_similarity"
-
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
name: stringThe name of the grader.
-
source: stringThe source code of the python script.
-
type: "python"The object type, which is always
python."python"
-
image_tag: optional stringThe image tag to use for the python script.
-
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
input: array of object { content, role, type }The input messages evaluated by the grader. Supports text, output text, input image, and input audio content blocks, and may include template strings.
-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
model: stringThe model to use for the evaluation.
-
name: stringThe name of the grader.
-
type: "score_model"The object type, which is always
score_model."score_model"
-
range: optional array of numberThe range of the score. Defaults to
[0, 1]. -
sampling_params: optional object { max_completions_tokens, reasoning_effort, seed, 2 more }The sampling parameters for the model.
-
max_completions_tokens: optional number or nullThe maximum number of tokens the grader model may generate in its response.
-
reasoning_effort: optional ReasoningEffort or nullConstrains effort on reasoning for reasoning models. Currently supported values are
none,minimal,low,medium,high,xhigh, andmax. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Not all reasoning models support every value. See the reasoning guide for model-specific support.-
"none" -
"minimal" -
"low" -
"medium" -
"high" -
"xhigh" -
"max"
-
-
seed: optional number or nullA seed value to initialize the randomness, during sampling.
-
temperature: optional number or nullA higher temperature increases randomness in the outputs.
-
top_p: optional number or nullAn alternative to temperature for nucleus sampling; 1.0 includes all tokens.
-
-
-
LabelModelGrader object { input, labels, model, 3 more }A LabelModelGrader object which uses a model to assign labels to each item in the evaluation.
-
input: array of object { content, role, type }-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
labels: array of stringThe labels to assign to each item in the evaluation.
-
model: stringThe model to use for the evaluation. Must support structured outputs.
-
name: stringThe name of the grader.
-
passing_labels: array of stringThe labels that indicate a passing result. Must be a subset of labels.
-
type: "label_model"The object type, which is always
label_model."label_model"
-
-
-
name: stringThe name of the grader.
-
type: "multi"The object type, which is always
multi."multi"
-
Python Grader
-
PythonGrader object { name, source, type, image_tag }A PythonGrader object that runs a python script on the input.
-
name: stringThe name of the grader.
-
source: stringThe source code of the python script.
-
type: "python"The object type, which is always
python."python"
-
image_tag: optional stringThe image tag to use for the python script.
-
Score Model Grader
-
ScoreModelGrader object { input, model, name, 3 more }A ScoreModelGrader object that uses a model to assign a score to the input.
-
input: array of object { content, role, type }The input messages evaluated by the grader. Supports text, output text, input image, and input audio content blocks, and may include template strings.
-
content: string or ResponseInputText or object { text, type } or 3 moreInputs to the model - can contain template strings. Supports text, output text, input images, and input audio, either as a single item or an array of items.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
text: stringThe text input to the model.
-
type: "input_text"The type of the input item. Always
input_text."input_text"
-
prompt_cache_breakpoint: optional object { mode }Marks the exact end of a reusable prompt prefix. The breakpoint inherits its TTL from the request's
prompt_cache_options.ttl; the boundary is not rounded to a token block.-
mode: "explicit"The breakpoint mode. Always
explicit."explicit"
-
-
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
input_audio: object { data, format }-
data: stringBase64-encoded audio data.
-
format: "mp3" or "wav"The format of the audio data. Currently supported formats are
mp3andwav.-
"mp3" -
"wav"
-
-
-
type: "input_audio"The type of the input item. Always
input_audio."input_audio"
-
-
GraderInputs = array of string or ResponseInputText or object { text, type } or 2 moreA list of inputs, each of which may be either an input text, output text, input image, or input audio object.
-
TextInput = stringA text input to the model.
-
ResponseInputText object { text, type, prompt_cache_breakpoint }A text input to the model.
-
OutputText object { text, type }A text output from the model.
-
text: stringThe text output from the model.
-
type: "output_text"The type of the output text. Always
output_text."output_text"
-
-
InputImage object { image_url, type, detail }An image input block used within EvalItem content arrays.
-
image_url: stringThe URL of the image input.
-
type: "input_image"The type of the image input. Always
input_image."input_image"
-
detail: optional stringThe detail level of the image to be sent to the model. One of
high,low, orauto. Defaults toauto.
-
-
ResponseInputAudio object { input_audio, type }An audio input to the model.
-
-
-
role: "user" or "assistant" or "system" or "developer"The role of the message input. One of
user,assistant,system, ordeveloper.-
"user" -
"assistant" -
"system" -
"developer"
-
-
type: optional "message"The type of the message input. Always
message."message"
-
-
model: stringThe model to use for the evaluation.
-
name: stringThe name of the grader.
-
type: "score_model"The object type, which is always
score_model."score_model"
-
range: optional array of numberThe range of the score. Defaults to
[0, 1]. -
sampling_params: optional object { max_completions_tokens, reasoning_effort, seed, 2 more }The sampling parameters for the model.
-
max_completions_tokens: optional number or nullThe maximum number of tokens the grader model may generate in its response.
-
reasoning_effort: optional ReasoningEffort or nullConstrains effort on reasoning for reasoning models. Currently supported values are
none,minimal,low,medium,high,xhigh, andmax. Reducing reasoning effort can result in faster responses and fewer tokens used on reasoning in a response. Not all reasoning models support every value. See the reasoning guide for model-specific support.-
"none" -
"minimal" -
"low" -
"medium" -
"high" -
"xhigh" -
"max"
-
-
seed: optional number or nullA seed value to initialize the randomness, during sampling.
-
temperature: optional number or nullA higher temperature increases randomness in the outputs.
-
top_p: optional number or nullAn alternative to temperature for nucleus sampling; 1.0 includes all tokens.
-
-
String Check Grader
-
StringCheckGrader object { input, name, operation, 2 more }A StringCheckGrader object that performs a string comparison between input and reference using a specified operation.
-
input: stringThe input text. This may include template strings.
-
name: stringThe name of the grader.
-
operation: "eq" or "ne" or "like" or "ilike"The string check operation to perform. One of
eq,ne,like, orilike.-
"eq" -
"ne" -
"like" -
"ilike"
-
-
reference: stringThe reference text. This may include template strings.
-
type: "string_check"The object type, which is always
string_check."string_check"
-
Text Similarity Grader
-
TextSimilarityGrader object { evaluation_metric, input, name, 2 more }A TextSimilarityGrader object which grades text based on similarity metrics.
-
evaluation_metric: "cosine" or "fuzzy_match" or "bleu" or 8 moreThe evaluation metric to use. One of
cosine,fuzzy_match,bleu,gleu,meteor,rouge_1,rouge_2,rouge_3,rouge_4,rouge_5, orrouge_l.-
"cosine" -
"fuzzy_match" -
"bleu" -
"gleu" -
"meteor" -
"rouge_1" -
"rouge_2" -
"rouge_3" -
"rouge_4" -
"rouge_5" -
"rouge_l"
-
-
input: stringThe text being graded.
-
name: stringThe name of the grader.
-
reference: stringThe text being graded against.
-
type: "text_similarity"The type of grader.
"text_similarity"
-