The starting prompt
Build me an eval agent for an AI feature. Take a knowledge base of golden question/answer pairs, run them against a target LLM endpoint I configure, score responses using structural checks plus an LLM judge, and write results to Google Sheets. When a new prompt is committed to GitHub, re-run and comment pass/fail on the PR.
Integrations this uses
GitHubGoogle SheetsKnowledge base