Trick
Few-shot labels teach format more than meaning
The prompt
This is awesome! // Negative This is bad! // Positive Wow that movie was rad! // Positive What a horrible show! //
Expected result
"Negative" for the final line — notably, the model tends to get this right even though the few-shot labels shown are deliberately randomized/wrong, because the output FORMAT is what's really being taught, not the labels.
Why it works
A deliberately mischievous few-shot example: the demonstration labels are randomized (awesome=Negative, bad=Positive), yet the model still classifies the final line correctly — evidence that few-shot examples teach format/structure at least as much as they teach the actual label mapping.