Loading evaluation records
/ Target book context
Target book context
Persona, writing examples, and full writer prompt
Generated rubric
Rubric reasoning
Raw rubric response
Selected judge output
Scoring reasoning
Raw scoring response
Scored samples
Proxy rubric-judge score Gold cached external-judge score against the human reference
Full evaluation
Spearman scaling
M = 16 rubrics, N = 4 judge draws, D = 64 candidates
Candidate selection