Eval refusal classifier

Judges: model response. Questions: choice:type - answered/refused/hedged/off_topic. Action: Count refusal rates per prompt set.

AA clean typed decision (yes/no, a pick from a list, or a level on a scale) that is concrete, repeats routinely, and has meaningful value.3.82

Key facts

Vertical
Software & tech
Function
Agents & dev
Status
Idea
Volume
routine
Value
meaningful
Risk
low
Evidence
—
Flags
None

Build this with a classifier

Define a typed decision with a bounded answer, then evaluate it on examples.

{
  "decision_type": "choice",
  "question": "Does this input match the decision in “Eval refusal classifier”?",
  "input": "<input to classify>",
  "output": "one label from a fixed list"
}

Related use cases

Cite this

Copy a link in your preferred format.