AI
Prompt A/B Eval
Two system prompts on one user message — side-by-side outputs plus structured judge scores.
Run this experiment yourself
Demos are not embedded on this site. Deploy a standalone copy on Vercel or run the experiment app locally.
Local development
cd apps/experiments/prompt-ab-eval pnpm install pnpm dev
Then open http://localhost:3010.
This is an experimental demo. Use it as a starting point for your own projects.
Run the same user prompt under two system prompts, compare outputs side-by-side, and score with a judge model.
Features
- Standalone Next.js app under
apps/experiments/prompt-ab-eval/. - Honest degradation when optional services are missing.
- Unit-tested helpers in
logic.ts.
Server Reference
POST /api/eval
{ "systemA", "systemB", "user" }Implementation Details
Run the demo
Use the embedded playground or pnpm dev in the experiment app.
Inspect the handlers
Copy Route Handlers from apps/experiments/prompt-ab-eval/ into your project.
Limitations
- Educational demo — not a production-hardened product.
- Inputs are length-capped.
Use in your project
Copy the experiment folder or individual Route Handlers.
Deployment
Local Development
cd apps/experiments/prompt-ab-eval
pnpm install
pnpm devConfiguration
| Variable | Required | Purpose |
|---|---|---|
AI_GATEWAY_API_KEY | Yes | Lab configuration |
Vercel / Next.js Features Used
- AI SDK, AI Gateway, Zod