Handle the Reply That Got Cut Off
A summary feature has been shipping reports that end mid-sentence. Nothing raised. The call returned text, the code saved the text, and nobody looked at the one field that said the text was not finished.
Every reply carries a stop_reason: why the model stopped writing. Text is only an answer when that reason says so.
Implement complete(llm, messages, max_tokens=200, max_continues=3), returning {"text": ..., "status": ...}.
Call llm.messages.create(messages=..., max_tokens=max_tokens) and act on response.stop_reason:
"end_turn": the model finished. Status"complete"."max_tokens": the reply hit your cap and is cut off. Keep what you have and ask for the rest: extend the conversation with the partial reply as an"assistant"message, then a"user"message asking it to carry on from where it stopped, and call again. Join the pieces in order, adding nothing between them."tool_use": the model wants a tool run. This function has no tools, so that is the caller's business: stop, status"tool_use"."refusal": the model declined. Asking it to continue will not change its mind: stop, status"refused".- Anything else: raise
ValueError. Providers add stop reasons over time, and code that treats one it has never seen as success is how the mid-sentence reports shipped.
Continue at most max_continues times. If the reply is still cut off after that, return what you have with status "truncated", so the caller knows not to trust the ending.
text is always everything received so far, whatever the status. Do not modify the list the caller passed in.