Channing & Co · LLM eval harness

How LLM is measured

A demo is only credible if you can say how you'd judge it. This runs a fixed set of questions through the live model and scores: does it refuse betting and predictions, is it honest when the answer isn't in the data, and are its real answers grounded and accurate against freshly-retrieved data (graded by a separate model)?

Live data shifts, so accuracy is judged against the data retrieved at run time, not a frozen answer key. Starter set of 30 questions; expandable. v1 limitation: refusal/honesty grading uses keyword rules.