Skip to content

< CurriculumModel APIs · 9 of 46 ·63 · Level 01, LLM APIs

Handle the Reply That Got Cut Off

medium · implement · LLM Fundamentals

A summary feature has been shipping reports that end mid-sentence, because nobody looked at the field that said the text was not finished. Every reply carries a stop_reason, and text is only an answer when that reason says so.

Implement complete(llm, messages, max_tokens=200, max_continues=3), returning {"text": ..., "status": ...}. Call llm.messages.create(messages=..., max_tokens=max_tokens) and act on response.stop_reason:

  • "end_turn": the model finished. Status "complete".
  • "max_tokens": the reply is cut off. Extend the conversation with the partial reply as an "assistant" message, then a "user" message asking it to carry on from where it stopped, and call again. Join the pieces in order, adding nothing between them.
  • "tool_use": the model wants a tool run. This function has no tools: stop, status "tool_use".
  • "refusal": the model declined: stop, status "refused".
  • Anything else: raise ValueError. A stop reason you have never seen is not success.

Continue at most max_continues times. If the reply is still cut off after that, return what you have with status "truncated".

text is always everything received so far, whatever the status. Do not modify the list the caller passed in.