Skip to content

< all problems60 · Level 01, LLM APIs

Make Your First Model Call

easy · implement · LLM Fundamentals

Everything on this site sits on one request: you send a model some messages, it sends one back. Before parsing, retrying or looping, make that request yourself and look at everything that comes back, because the reply is more than its text.

Implement first_call(llm, question, system, max_tokens=100).

Make one call with llm.messages.create(...):

  • system goes in the system argument. It is the instruction the model reads before the conversation, and it is not one of the messages.
  • question goes in messages, as a single message whose role is "user".
  • max_tokens caps how long the reply may be. Pass it through.

The response object has .text, .stop_reason and .usage (with .input_tokens and .output_tokens). Return a dict with four keys:

  • "text": the reply, with surrounding whitespace stripped.
  • "stop_reason": why the model stopped. "end_turn" means it finished; "max_tokens" means it hit your cap and the text is cut off.
  • "input_tokens" and "output_tokens": what the call cost you. You pay for both, at different prices.

This one runs against a real model. The tests check properties of what you return (the answer is in the text, the call was made once, a tiny cap really does cut the reply off), not exact words.