CompletionKit

Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.

AI & MLv1.0.0