Make Your First Model Call
Everything on this site sits on one request: you send a model some messages, it sends one back. Before parsing, retrying or looping, make that request yourself and look at everything that comes back, because the reply is more than its text.
Implement first_call(llm, question, system, max_tokens=100).
Make one call with llm.messages.create(...):
systemgoes in thesystemargument. It is the instruction the model reads before the conversation, and it is not one of the messages.questiongoes inmessages, as a single message whose role is"user".max_tokenscaps how long the reply may be. Pass it through.
The response object has .text, .stop_reason and .usage (with .input_tokens and .output_tokens). Return a dict with four keys:
"text": the reply, with surrounding whitespace stripped."stop_reason": why the model stopped."end_turn"means it finished;"max_tokens"means it hit your cap and the text is cut off."input_tokens"and"output_tokens": what the call cost you. You pay for both, at different prices.
This one runs against a real model. The tests check properties of what you return (the answer is in the text, the call was made once, a tiny cap really does cut the reply off), not exact words.