Define what good means
Combine built-in metrics with numeric, boolean, categorical, model-judged, and code-based scoring criteria.
Open-source AI infrastructure
Evaluate AI outputs with built-in metrics, LLM-as-judge scoring, sandboxed code scorers, datasets, baselines, scheduling, CI gates, and human review.
Combine built-in metrics with numeric, boolean, categorical, model-judged, and code-based scoring criteria.
Run the same dataset across prompt, model, or application changes and compare the result with an accepted baseline.