LLM evaluation
Evaluate language systems across tasks and risks
LLM evaluation combines task-specific tests, reference-based metrics, human judgements and sometimes model-based judging to assess a specified system. The system includes prompts, retrieval, tools and decoding settings as well as a model version. Evaluation should cover relevant capabilities, robustness, calibration and harmful failures. Benchmark contamination, prompt sensitivity and judge bias require explicit checks and transparent reporting.
Choose it when selecting or changing a language-based system and realistic tasks, acceptable outputs and consequential failures can be defined, with representative test cases and reliable adjudication resources.
Strengths
- Combines broad standardised tests with application-specific criteria
- Can expose trade-offs hidden by a single benchmark score
Limitations
- Benchmark contamination can inflate apparent capability
- Automated judges can favour style, position or their own model family
Know the boundary
A model judge is a fallible measurement instrument; benchmark success does not certify safety or reliability for every deployment.