Evaluation · bring your own key
Generation and question quality
Eight fixed topics, the same retrieved passages for everyone, and your chosen model. Every generated question is checked for citations, grounded numbers, answerability, distractors and duplicates, and you rate it on a short rubric. Nothing here is published; runs and ratings stay in your browser. Retrieval results are on Evaluation.
Generation evaluation
Same topics, same passages, your model.
Each selected topic retrieves the same five passages for everyone, and your model writes four multiple-choice and three short-answer questions from them. Every item is checked automatically, and you rate it on a short rubric. The summary reports both with intervals, and how often the checks agree with you.