Skip to content

< all problems36 · Level 01, LLM APIs

Parse JSON From a Stream

medium · implement · LLM Fundamentals

A streamed model reply arrives as fragments that can split anywhere, even mid-number. When the reply is JSON, json.loads fails on every fragment but the last, and you do not know which one is last.

Implement parse_stream(chunks) returning (parsed_object, chunks_consumed):

  1. Accumulate fragments. Once the text is a complete JSON object, parse it and stop reading: return how many fragments you needed. Trailing text after the JSON ("Let me know if...") must not be read.
  2. The model may wrap the JSON in a markdown fence. Handle it.
  3. If the stream ends before the JSON completes, return (None, len(chunks)). A truncated stream is a real outcome, not an exception.

The catch: knowing when the object is complete is the actual work. json.loads after every fragment is correct and quadratic; brace depth is linear but gets confused by braces inside strings. {"text": "a } b"} has a } that does not close anything.