Trick
Model-as-judge: compare two outputs like a teacher
The prompt
Can you compare the two outputs below as if you were a teacher?
Output from ChatGPT: {output 1}
Output from GPT-4: {output 2}Expected result
A structured, teacher-style critique comparing both outputs — argument coherence, style fidelity, rigor — rather than just a one-line preference.
Why it works
A model-as-judge template for comparing two LLM outputs on the same task: framing it as "compare as if you were a teacher" pulls out a structured critique instead of a bare preference, useful for any A/B comparison of generated text.