Trick

Model-as-judge: compare two outputs like a teacher

evaluation · works with Any model · via dair-ai/Prompt-Engineering-Guide

The prompt

Can you compare the two outputs below as if you were a teacher?

Output from ChatGPT: {output 1}

Output from GPT-4: {output 2}

Expected result

A structured, teacher-style critique comparing both outputs — argument coherence, style fidelity, rigor — rather than just a one-line preference.

Why it works

A model-as-judge template for comparing two LLM outputs on the same task: framing it as "compare as if you were a teacher" pulls out a structured critique instead of a bare preference, useful for any A/B comparison of generated text.

Open in PrompVite to copy, rate, and discuss this prompt →