How do you evaluate LLM output quality without a golden set?
LLM-as-judge is popular but biased; human eval doesn't scale. For a team that needs quick iteration, what evaluation setup is worth building first?
3
0 commentsLLM-as-judge is popular but biased; human eval doesn't scale. For a team that needs quick iteration, what evaluation setup is worth building first?