Multi-criteria LLM eval framework
LLM judges score outputs on multiple criteria (relevance, accuracy, tone); use rubrics for reproducible evaluation at scale
AA clean typed decision (yes/no, a pick from a list, or a level on a scale) that is concrete, repeats routinely, and has meaningful value.4.61
Key facts
- Vertical
- Software & tech
- Function
- Agents & dev
- Status
- Seen in the wild
- Volume
- routine
- Value
- meaningful
- Risk
- low
- Evidence
- described plan
- Flags
- None
Source: https://www.galtea.ai/blog/llm-as-a-judge-the-complete-guide
Build this with a classifier
Define a typed decision with a bounded answer, then evaluate it on examples.
{
"decision_type": "choice",
"question": "Does this input match the decision in “Multi-criteria LLM eval framework”?",
"input": "<input to classify>",
"output": "one label from a fixed list"
}Related use cases
Cite this
Copy a link in your preferred format.