You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Evidence-first LLM-as-judge scoring for AI artifacts — papers, PRs, prompts, cold emails: synthesizes a rubric, collects quoted-evidence citations, scores only against that evidence, and hedges on thin evidence. Every dimension ties to a file:line or quote, with reproducibility receipts. 7 backends.