Special token
Explicit stop sequence
The prompt
Set a stop token such as <|endoftext|> or a custom sequence like ###END### explicitly in your API call — it prevents the model from continuing to generate past the point you actually need.
Expected result
Generation halts exactly at the stop sequence instead of continuing to produce extra, unrequested content after the answer.
Why it works
Without an explicit stop sequence, the model keeps generating until it hits its max-token limit or decides on its own to stop — often later than needed.