Our customer support chatbot runs as a web service. Many customers chat with it at the same time, and every request arrives as chat(session_id, user_message), where session_id identifies one customer's conversation.
The bot answers the first message fine, but on every follow-up it has forgotten what the customer already said, so people keep repeating themselves.
llm.chat(messages=...) takes the usual list of {"role": ..., "content": ...} dicts (roles: system, user, assistant) and returns an object with a .content string. It keeps no memory of its own between calls.
Your task: find the bug and fix it so each customer's conversation keeps its own history. Keep the chat(session_id, user_message) signature, since the web layer calls it. Keeping history in memory is fine for this exercise; mention what you would change for production.