#15861·phoenix

[online evals] [evaluators] evaluator playground

Author: mikeldkingCreated Sep 2, 2026Updated Sep 18, 2026

As a user, I want the ability to load an evaluator into the playground - this could be an LLM OR code evaluator.

I then also want to load a dataset that is representative of production data and run it though the playground.

After the first run I want to be able to say the evaluator produced the right answer or the wrong answer.

The evaluator then would be "re-aligned" and run again until it is fully in alignment with the ground truth.

"Evaluator" mode is an implicit state of the playground, entered by selecting an evaluator task from the new task menu.