Multi-criteria LLM eval framework

LLM judges score outputs on multiple criteria (relevance, accuracy, tone); use rubrics for reproducible evaluation at scale

AA clean typed decision (yes/no, a pick from a list, or a level on a scale) that is concrete, repeats routinely, and has meaningful value.4.61

Key facts

Vertical
Software & tech
Function
Agents & dev
Status
Seen in the wild
Volume
routine
Value
meaningful
Risk
low
Evidence
described plan
Flags
None

Source: https://www.galtea.ai/blog/llm-as-a-judge-the-complete-guide

Build this with a classifier

Define a typed decision with a bounded answer, then evaluate it on examples.

{
  "decision_type": "choice",
  "question": "Does this input match the decision in “Multi-criteria LLM eval framework”?",
  "input": "<input to classify>",
  "output": "one label from a fixed list"
}

Related use cases

Cite this

Copy a link in your preferred format.