ATLASResearch
methods
ACROSS FAMILIES/ Evaluation strategy

LLM evaluation

Evaluate language systems across tasks and risks

LLM evaluation combines task-specific tests, reference-based metrics, human judgements and sometimes model-based judging to assess a specified system. The system includes prompts, retrieval, tools and decoding settings as well as a model version. Evaluation should cover relevant capabilities, robustness, calibration and harmful failures. Benchmark contamination, prompt sensitivity and judge bias require explicit checks and transparent reporting.

WHEN IT FITS

Choose it when selecting or changing a language-based system and realistic tasks, acceptable outputs and consequential failures can be defined, with representative test cases and reliable adjudication resources.

Strengths

  • Combines broad standardised tests with application-specific criteria
  • Can expose trade-offs hidden by a single benchmark score

Limitations

  • Benchmark contamination can inflate apparent capability
  • Automated judges can favour style, position or their own model family

Know the boundary

A model judge is a fallible measurement instrument; benchmark success does not certify safety or reliability for every deployment.

USED ACROSS
Computer scienceEducationBusiness & MBA