A few prompts that I am storing in a repo for the purpose of running controlled experiments comparing and benchmarking different LLMs for defined use-cases
-
Updated
Dec 4, 2024 - Python
A few prompts that I am storing in a repo for the purpose of running controlled experiments comparing and benchmarking different LLMs for defined use-cases
Building a framework to run prompt evaluation tasks.
Python harness for evaluating and A/B testing LLM prompt variants against a labeled dataset, with deterministic scoring and regression checks.
AI-powered tool that analyzes YouTube titles for hook effectiveness and exports structured CSV reports.
Claude Code / AI Coding Agent 评估方法论:覆盖 Agent、Prompt、RAG 三层 eval 的可复用框架
Add a description, image, and links to the prompt-eval topic page so that developers can more easily learn about it.
To associate your repository with the prompt-eval topic, visit your repo's landing page and select "manage topics."