How do you evaluate LLM output quality without a golden set?

LLM-as-judge is popular but biased; human eval doesn't scale. For a team that needs quick iteration, what evaluation setup is worth building first?

3
0 comments

Comments (0)