Make Your First Model Call
Make one request to a model and return everything that comes back, not just the text.
Implement first_call(llm, question, system, max_tokens=100). Make one call with llm.messages.create(...):
systemgoes in thesystemargument, not in a message.questiongoes inmessages, as a single message whose role is"user".max_tokensis passed through.
The response object has .text, .stop_reason and .usage (with .input_tokens and .output_tokens). Return a dict with four keys:
"text": the reply, with surrounding whitespace stripped."stop_reason":"end_turn"means the model finished;"max_tokens"means it hit your cap and the text is cut off."input_tokens"and"output_tokens": from.usage.
This runs against a real model, so the tests check properties (the answer is in the text, the call was made once, a tiny cap really does cut the reply off), not exact words.