# Streaming the agent

By default, `POST /v2/agent/chat` responds with a Server-Sent Events stream. A single turn can call several tools one after another, and each result is sent to you as soon as it exists. Your interface can show the opening sentence, then product cards, then the closing line, instead of waiting for the whole turn.

This guide covers the wire format, every event the server sends, their order, how to carry conversation state, and how to handle errors and retries.

## Before the stream: check the response

Authentication, quota, rate-limit and validation failures are all decided before any stream opens. Those responses are ordinary JSON with a real HTTP status:

```http
HTTP/1.1 422 Unprocessable Entity
Content-Type: application/json

{"error":"messages cannot be empty."}
```

Always check both the status and the `Content-Type` before you start reading the body as a stream:

```javascript
const type = res.headers.get('content-type') || '';
if (!res.ok || !type.includes('text/event-stream')) {
  const body = await res.json().catch(() => ({}));
  throw new Error(body.error || `HTTP ${res.status}`);
}
```

A successful stream responds with `200`, `Content-Type: text/event-stream`, `Cache-Control: no-cache` and the usual `X-RateLimit-*` headers.

## Frame format

Each event is one frame: an `event:` line, one `data:` line holding a JSON object, and a blank line.

```text
event: text_delta
data: {"delta":"Let me pull together a few options."}

event: done
data: {"output":"...","history":[...]}

```

- The server never sends comments, `id:` lines or `retry:` lines.
- `data` is always a single line of JSON.
- Split the byte stream on a blank line (`\n\n`). A network chunk can end partway through a frame, so keep whatever follows the last blank line in a buffer.

The endpoint is a `POST`, so the browser's built-in `EventSource` cannot call it. Use `fetch` with a stream reader, or an SSE client library that supports POST.

## Event catalogue

| Event | Sent when | Payload |
| --- | --- | --- |
| `thinking` | The model starts or finishes reasoning (every turn, unless `thinking` is `none`) | `{ status }` |
| `thinking_delta` | A chunk of the reasoning summary (same condition) | `{ delta }` |
| `thinking_break` | A new summary paragraph (same condition) | `{}` |
| `text_delta` | The model writes text | `{ delta }` |
| `tool_called` | The model calls a tool | `{ name, arguments }` |
| `search_results` | The agent puts a row of products on screen | `{ label, kind, query, count, results }` |
| `suggestions` | The agent offers follow-up chips | `{ suggestions }` |
| `question` | The agent asks one or more multiple-choice questions on one card | `{ questionId, question, options, allowFreeText, steps? }` |
| `looks` | The agent composes outfits to render | `{ looks: [{ label, items, customInstructions }] }` |
| `cart_action` | The agent reads or wants to change the cart | `{ action, ... }` |
| `tool_output` | A tool call finishes | `{ name, output }` |
| `done` | The turn finished | `{ output, history }` |
| `error` | The run failed after the stream opened | `{ error }` |

### thinking, thinking_delta, thinking_break

Only on a turn that reasons, which by default is every turn (`low`). A request that sets `thinking` to `none` gets no reasoning and none of these events. A client that ignores them loses nothing.

`thinking` marks the start and end of a stretch of reasoning: `{ "status": "started" }` then `{ "status": "done" }`. Show a thinking state between the two. A summary may stream in between as `thinking_delta` chunks, with `thinking_break` between paragraphs, or the provider may send no summary at all for a short stretch — the start/done pair is the reliable signal, the summary is a bonus. The stretch is over when text or a tool call begins.

```json
{ "status": "started" }
{ "delta": "They want a dinner look; I'll check dresses and heels first." }
{ "status": "done" }
```

Set `thinking` in the request to choose for yourself: `none` to switch it off on a turn, `true` for the lightest effort, or a level (`low`, `medium`, `high`, `xhigh`, `max`). A level the model does not accept is a `422` before the stream opens. Reasoning adds about 700 ms to the wait before the first word at `low`; `none` gives that back, at the cost of an occasional turn that ends on a promise with nothing on screen.

### text_delta

A piece of the assistant's reply. Append the pieces in order. A turn often writes text in two separate runs: a short opening line before the tools run, and a closing line afterwards. Treat `tool_called` as the end of the current paragraph.

```json
{ "delta": "Let me pull together a few options." }
```

### tool_called

The model called a tool. `arguments` is the model's raw JSON string, or `null`. Use it to show progress, for example "Searching..." while `lookup_products` runs.

```json
{
  "name": "lookup_products",
  "arguments": "{\"query\":\"linen wedding shirt\",\"slot\":null,\"family\":[\"shirt\"],\"age_group\":null,\"gender\":\"mens\",\"colors\":null,\"pattern\":null,\"fit\":null,\"materials\":[\"linen\"],\"features\":null,\"minPrice\":null,\"maxPrice\":300,\"onSale\":null,\"exclude_terms\":null,\"limit\":8}"
}
```

These tools are always available: `get_articles`, `get_colors`, `get_details`, `lookup_products`, `find_product`, `show_products`, `create_looks`, `ask_question` and `suggest_replies`. When the organization's catalogue is synced from a Shopify store, `list_store_pages` and `read_store_page` let the agent answer policy, FAQ and sizing questions from the store's own pages. When the request includes a `cart`, five more are added: `get_cart`, `add_to_cart`, `set_cart_quantity`, `change_cart_size` and `remove_from_cart`.

### Links in replies

The agent's text can contain markdown links. It is instructed to link only to the store's own site: a page it read (as a path, such as `[Our story](https://stylor.ai/pages/our-story)`) or a product. Resolve a path against the storefront page your UI is on.

Do not rely on the instruction alone. Before you make a link clickable, check that it stays on the store's site: the same host, or a subdomain of it. Draw anything else as plain text. The hosted widget does this. `mailto:` and `tel:` links are safe to keep.

### tool_output

A tool finished. `output` is what the model received. It is usually an object with a `note` written for the model, and sometimes an `error` key. Product searches return short summaries here; the full product records arrive on `search_results`.

One exception: when a store page is a picture (a size chart saved as an image), `read_store_page` shows the model that picture. Images are never sent to you. In that case `output` is an array of text parts, with a line of text standing where each picture was. The same applies to `history`: pictures a tool showed the model are replaced by that line before the history is returned.

```json
{
  "name": "lookup_products",
  "output": {
    "count": 8,
    "items": [
      { "id": "8f2c1a9e-4b7d", "description": "Coastal Linen Shirt, navy, linen, relaxed fit, $89" }
    ],
    "note": "The shopper saw NOTHING. This was for your eyes only."
  }
}
```

An `error` inside `output` is not a failed request. The model reads it and adapts. An invalid dataset scope, for example, shows up only here. When the model's arguments fail validation, `output` is a plain string that starts with `An error occurred while running the tool.` You can show these payloads in a debug view, but don't show them to shoppers.

### search_results

A row of products the agent chose to show. `kind` is `outfit` when the items are worn together (at most one per slot) and `list` when they are alternatives, so lay the two out differently. `query` repeats `label`. `results` holds full product records in the agent's chosen order, and that order is part of the recommendation. A turn can include several rows.

```json
{
  "label": "Wedding shirts",
  "kind": "list",
  "query": "Wedding shirts",
  "count": 3,
  "results": [
    {
      "itemId": "8f2c1a9e-4b7d-4c1e-9a3f-2d6b7e5c1a90",
      "datasetId": "9c21f0aa-3e4b-4d2a-b6d1-0f7e8a9b1c2d",
      "status": "active",
      "metadata": { "gender": "mens", "color_family": "navy", "slot": "top", "family": "shirt" },
      "product": {
        "name": "Coastal Linen Shirt",
        "url": "https://shop.example.com/products/coastal-linen-shirt",
        "price": 89,
        "currency": "USD",
        "sizes": [{ "label": "M / Navy", "sku": "CLS-NVY-M", "price": 89, "inStock": true }],
        "assets": { "images": [{ "uploaded": { "key": "org_5f3a/ds_9c21/8f2c1a9e-4b7d/0.jpg" } }] }
      }
    }
  ]
}
```

Image `key` values are storage keys, not URLs. Build the URL as `https://stylor.ai/api/v1/image/<key>`, or use `originalImageUrl` instead.

### suggestions

Two to four follow-up replies, written as the shopper would say them. They are sent alongside `search_results` or `looks` when the agent attaches them, and by `suggest_replies` after a text-only answer. If a turn sends more than one set, keep the latest. When the shopper taps a chip, send its text as the next user message. The agent never sends chips in the same turn as a `question`.

```json
{ "suggestions": ["Show me trousers to match", "Something more formal", "Cheaper options"] }
```

### question

A multiple-choice card. The agent has stopped to wait for the answer, so the turn ends shortly after this event. When the shopper picks an option, send that option's `label` as the next user message. If `allowFreeText` is `true`, also let them type their own answer.

```json
{
  "questionId": "5f0e8c3a-2b1d-4e7f-9a6c-3d2e1f0a9b8c",
  "question": "Who am I styling for?",
  "options": [
    { "label": "Menswear", "value": "mens" },
    { "label": "Womenswear", "value": "womens" }
  ],
  "allowFreeText": false
}
```

A card can ask several questions at once, most often a size for each piece of an outfit going into the cart. It then carries a `steps` array with every question, the first included. `question`, `options` and `allowFreeText` repeat the first step, so a client that ignores `steps` still shows something answerable.

```json
{
  "questionId": "0b9c1f7e-4d2a-4c1b-8e55-7a1d2c3b4e5f",
  "question": "Which size in the shirt?",
  "options": [{ "label": "S", "value": "s" }, { "label": "M", "value": "m" }],
  "allowFreeText": false,
  "steps": [
    { "topic": "Shirt size", "question": "Which size in the shirt?", "options": [{ "label": "S", "value": "s" }, { "label": "M", "value": "m" }], "allowFreeText": false },
    { "topic": "Shoe size", "question": "Which size in the loafers?", "options": [{ "label": "9", "value": "9" }, { "label": "10", "value": "10" }], "allowFreeText": false }
  ]
}
```

Let the shopper answer every step, then reply once. Send one user message with a line per step, in order, as `<topic>: <label>`:

```
Shirt size: M
Shoe size: 10
```

Do not send each answer as its own message. The agent acts on the whole set, for example adding every piece to the cart in one turn.

### looks

One to four complete outfits, each with two to six full product records and optional rendering direction. Nothing has been rendered yet. For each look, call `POST /v2/agent/looks` with the items' `itemId` values, the `label` and `customInstructions`, plus the same `datasetIds` (and `sessionId`/`session`) you sent to the chat request, and send those calls in parallel. A look can include items that never appeared as cards.

```json
{
  "looks": [
    {
      "label": "Garden wedding",
      "items": [
        { "itemId": "8f2c1a9e-4b7d-4c1e-9a3f-2d6b7e5c1a90", "product": { "name": "Coastal Linen Shirt", "price": 89, "currency": "USD" } },
        { "itemId": "1b7e33c0-9a2d-4f8e-b5c6-7d8e9f0a1b2c", "product": { "name": "Everyday Chino", "price": 79, "currency": "USD" } },
        { "itemId": "c4d5e6f7-0a1b-4c2d-8e9f-0a1b2c3d4e5f", "product": { "name": "Suede Loafer", "price": 129, "currency": "USD" } }
      ],
      "customInstructions": "Shirt untucked with the top button open; outdoor daylight setting."
    }
  ]
}
```

(The items above are shortened. Real payloads carry full product records.)

### cart_action

Sent only when the request included a `cart`. A `read` has already been answered from your snapshot, so you can show it as a receipt. `add`, `quantity`, `swap` and `remove` are requests your client must carry out on the storefront. The [cart guide](https://stylor.ai/guides/rest/v2/cart.md) covers every action.

```json
{
  "action": "add",
  "itemId": "8f2c1a9e-4b7d-4c1e-9a3f-2d6b7e5c1a90",
  "name": "Coastal Linen Shirt",
  "url": "https://shop.example.com/products/coastal-linen-shirt",
  "sku": "CLS-NVY-M",
  "size": "M / Navy",
  "quantity": 1,
  "price": 89,
  "currency": "USD",
  "image": "org_5f3a/ds_9c21/8f2c1a9e-4b7d/0.jpg",
  "intent": "eyJ2IjoxLCJqdGkiOiI3ZjFlIn0.q8vT2rN9..."
}
```

### done

The turn finished and the stream closes. `output` is the final assistant text. It can be an empty string when a card, row or look was the whole answer. `history` is the complete conversation, ready to send back.

```json
{
  "output": "All three are linen, so they will breathe on a warm afternoon.",
  "history": [
    { "role": "user", "content": "I need something for a summer wedding, menswear, under $300." },
    { "type": "message", "role": "assistant", "status": "completed", "content": [{ "type": "output_text", "text": "Let me pull together a few options." }] },
    { "type": "function_call", "callId": "call_8kP2wN7r", "name": "show_products", "arguments": "{\"label\":\"Wedding shirts\",\"kind\":\"list\",\"itemIds\":[\"8f2c1a9e-4b7d\"],\"suggestions\":[\"Show me trousers to match\",\"Something more formal\"]}", "status": "completed" },
    { "type": "function_call_result", "callId": "call_8kP2wN7r", "name": "show_products", "status": "completed", "output": { "type": "text", "text": "{\"shown\":1,\"note\":\"1 card is now on the shopper's screen.\"}" } },
    { "type": "message", "role": "assistant", "status": "completed", "content": [{ "type": "output_text", "text": "All three are linen, so they will breathe on a warm afternoon." }] }
  ]
}
```

### error

The run failed after the stream had already opened, for example because the model hit the ten-call limit or the model provider returned an error. The HTTP status is already `200`, so this event is the only sign of failure. No `done` follows, and the stream closes.

```json
{ "error": "Max turns (10) exceeded" }
```

## Ordering

A typical turn arrives in this order:

```text
text_delta ...                 opening line
tool_called                    one per tool the model called in this step
  search_results | looks | question | suggestions | cart_action
tool_output                    one per tool, after the step's tools finish
... more tool steps ...
text_delta ...                 closing line (optional)
done                           or error
```

Rules you can rely on:

- In each model step, the text streams first, then that step's tool calls.
- A tool's display events always come after its own `tool_called` and before its own `tool_output`.
- When the model calls several tools in one step, all of that step's `tool_called` events arrive first. The display events follow as the tools produce them, then all the `tool_output` events. Tell display events apart by their event name, not their position.
- `done` or `error` is always the last frame. Exactly one of the two is sent whenever the stream opens.
- A turn can contain several `search_results` and several `suggestions`. At most one `question` is expected, and the turn ends soon after it.
- The first model call can take a moment to start. Don't treat a few seconds without events at the start as a failure.

## Keeping history across turns

The server keeps no conversation state. After each turn:

1. Store `done.history` as it is.
2. On the next turn, send `messages` as that history followed by the shopper's new message.

```javascript
const messages = [...history, { role: 'user', content: 'Show me trousers to match' }];
```

Replay the history exactly as you received it:

- Keep every item, including types you don't recognise, and fields such as `id` and `providerData`.
- Never remove a `function_call` without also removing its `function_call_result` (they share the same `callId`).
- Don't reorder or edit items. If you have to shorten a long conversation, cut whole turns from the start, beginning at a user message.
- Empty assistant messages are fine to send back; the server filters them out.

To attach a photo, send the user message's `content` as a list of parts:

```json
{
  "role": "user",
  "content": [
    { "type": "input_text", "text": "What would go with this jacket?" },
    { "type": "input_image", "image": "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQ..." }
  ]
}
```

## Error handling

| Where | How it arrives | What to do |
| --- | --- | --- |
| Before the stream (auth, quota, rate limit, validation) | JSON body `{ "error": "..." }` with a `4xx` or `5xx` status | Check the status. For `429`, wait `Retry-After` seconds. Fix `422` errors in the request body. |
| Unparseable request body | JSON `400`: `Request body must be valid JSON.` | Send valid JSON. |
| During the stream | An `error` event, then the stream closes | Keep the previous history; offer the shopper a retry. |
| Inside a tool | An `error` key or a string in `tool_output.output` | Nothing. The agent handles it within the turn. |
| Stream cut off (network) | The body ends without `done` or `error` | Treat it the same as an `error` event. |
| Rendering a look | JSON error from `/v2/agent/looks` | Show a fallback for that look; the others are unaffected. |

When a turn fails, no new `history` comes back. Your stored history is still the last good state, so a retry sends that history plus the same user message again.

## Reconnection and retries

The stream cannot be resumed. There are no event ids, and a new request always starts a new turn. When a stream fails or is cut off:

1. Throw away the partial turn in your interface, or mark it as failed.
2. Resend the last good `history` plus the same user message as a new request.
3. Each attempt uses one unit of `api.v2.agent.chat` quota, so retry once automatically at most. After that, let the shopper decide.
4. On `429`, honour `Retry-After` before trying again.

Take care with cart actions from a failed turn. If your client already carried out an `add` or `remove`, the cart has changed, so send a fresh cart snapshot with the retry to avoid applying the change twice.

To cancel a turn (for example, when the shopper closes the panel), abort the request. Partial output is not saved anywhere.

## JavaScript reader

```javascript
/**
 * Streams one agent turn and calls onEvent(name, data) for each frame.
 * Resolves with the new history; rejects on any failure.
 */
export async function streamAgentTurn({ history, message, cart, signal, onEvent }) {
  const res = await fetch('https://stylor.ai/api/v2/agent/chat', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.STYLOR_API_KEY}`,
      'Content-Type': 'application/json',
    },
    body: JSON.stringify({
      messages: [...history, { role: 'user', content: message }],
      ...(cart ? { cart } : {}),
    }),
    signal,
  });

  const type = res.headers.get('content-type') || '';
  if (!res.ok || !type.includes('text/event-stream')) {
    const body = await res.json().catch(() => ({}));
    const err = new Error(body.error || `HTTP ${res.status}`);
    err.status = res.status;
    err.retryAfter = Number(res.headers.get('retry-after')) || null;
    throw err;
  }

  const reader = res.body.getReader();
  const decoder = new TextDecoder();
  let buffer = '';
  let finished = null;

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;
    buffer += decoder.decode(value, { stream: true });

    let boundary;
    while ((boundary = buffer.indexOf('\n\n')) !== -1) {
      const frame = buffer.slice(0, boundary);
      buffer = buffer.slice(boundary + 2);

      let name = '';
      const dataLines = [];
      for (const line of frame.split('\n')) {
        if (line.startsWith('event:')) name = line.slice(6).trim();
        else if (line.startsWith('data:')) dataLines.push(line.slice(5).trimStart());
      }
      if (!name || dataLines.length === 0) continue;

      const data = JSON.parse(dataLines.join('\n'));
      onEvent(name, data);

      if (name === 'done') finished = data;
      if (name === 'error') throw new Error(data.error);
    }
  }

  if (!finished) throw new Error('Stream ended before the turn finished.');
  return finished.history;
}
```

## Python reader

```python
import codecs
import json
import os

import requests

API = "https://stylor.ai/api"


def stream_agent_turn(history, message, on_event, cart=None):
    """Streams one agent turn. Returns the new history or raises."""
    body = {"messages": [*history, {"role": "user", "content": message}]}
    if cart is not None:
        body["cart"] = cart

    with requests.post(
        f"{API}/v2/agent/chat",
        headers={
            "Authorization": f"Bearer {os.environ['STYLOR_API_KEY']}",
            "Content-Type": "application/json",
        },
        json=body,
        stream=True,
        timeout=(10, 120),
    ) as res:
        content_type = res.headers.get("Content-Type", "")
        if res.status_code != 200 or "text/event-stream" not in content_type:
            try:
                error = res.json().get("error")
            except ValueError:
                error = None
            raise RuntimeError(error or f"HTTP {res.status_code}")

        # Decode UTF-8 incrementally; a multi-byte character can span chunks.
        decoder = codecs.getincrementaldecoder("utf-8")()
        buffer = ""
        finished = None

        for chunk in res.iter_content(chunk_size=None):
            buffer += decoder.decode(chunk)

            while "\n\n" in buffer:
                frame, buffer = buffer.split("\n\n", 1)
                name, data_lines = "", []
                for line in frame.split("\n"):
                    if line.startswith("event:"):
                        name = line[6:].strip()
                    elif line.startswith("data:"):
                        data_lines.append(line[5:].lstrip())
                if not name or not data_lines:
                    continue

                data = json.loads("\n".join(data_lines))
                on_event(name, data)

                if name == "done":
                    finished = data
                elif name == "error":
                    raise RuntimeError(data["error"])

        if finished is None:
            raise RuntimeError("Stream ended before the turn finished.")
        return finished["history"]


def print_event(name, data):
    if name == "text_delta":
        print(data["delta"], end="", flush=True)
    elif name == "search_results":
        print(f"\n[{data['label']}] {[p['product']['name'] for p in data['results']]}")
    elif name == "looks":
        for look in data["looks"]:
            print(f"\n[look] {look['label']}: {[i['itemId'] for i in look['items']]}")


history = stream_agent_turn([], "I need something for a summer wedding, menswear.", print_event)
history = stream_agent_turn(history, "Show me trousers to match", print_event)
```

## Non-streaming mode

Send `"stream": false` to get one JSON body after the turn finishes. You receive the same `output` and `history`, plus the screen content collected into arrays: `searches`, `looks`, `questions` and `suggestions`. This mode has two limitations:

- `cart_action` events are not returned. Always stream when you send a `cart`.
- `suggestions` holds only chips from `suggest_replies`, not chips attached to product rows or looks.

A run that fails in this mode returns `500` with `Internal server error.` in `error`.
