Monoize
API 使用指南

流式输出

三种协议的 SSE 事件格式、结束标志、心跳和流中错误。

在请求体中设置 "stream": true,网关会用 SSE(Server-Sent Events)逐段返回结果。三种协议的事件格式不同。

通用规则

  • 响应头为 Content-Type: text/event-stream。
  • 认证失败、请求 JSON 无法解析、模型不在密钥的允许列表中时,网关直接返回 JSON 错误和对应的 HTTP 状态码。
  • 通过这些检查后,网关立即返回 HTTP 200 并开始推送。之后发生的错误(例如模型不存在、没有可用上游、上游断开)以错误事件的形式出现在流中。客户端必须检查流中的错误事件,不能只看 HTTP 状态码。
  • 长时间没有数据时,网关每 15 秒发送一次心跳,防止代理或负载均衡断开连接。

Chat Completions

每个事件只有 data: 行,没有 event: 行。流以 data: [DONE] 结束。

data: {"id":"chatcmpl_...","object":"chat.completion.chunk","created":1791730171,"model":"deepseek-v4-pro","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"chatcmpl_...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"reasoning_details":[{"type":"reasoning.text","text":"用户在打招呼。"}]},"finish_reason":null}]}

data: {"id":"chatcmpl_...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"你好"},"finish_reason":null}]}

data: {"id":"chatcmpl_...","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"chatcmpl_...","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":20,"total_tokens":32}}

data: [DONE]
  • 文本在 choices[0].delta.content 中。
  • 推理在 choices[0].delta.reasoning_details[] 中。见思维链(推理)。
  • 工具调用在 choices[0].delta.tool_calls[] 中,参数按片段拼接。见工具调用。
  • 结束块的 delta 为空,finish_reason 为 stop、length、tool_calls 或 content_filter。
  • 上游报告用量时,网关在结束块之后、[DONE] 之前发送一个 choices 为空数组的用量块。你不需要设置 stream_options.include_usage。
  • 心跳是 SSE 注释行 : heartbeat。标准 SSE 解析器会忽略它。

用 OpenAI SDK 读取流:

stream = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "写一首两行的诗。"}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print("\n用量:", chunk.usage.total_tokens)

流中的错误事件:

data: {"error":{"code":"model_not_found","message":"Model not found: no-such","param":null,"type":"invalid_request_error"}}

data: [DONE]

Responses

每个事件有 event: 行和 data: 行,data 中的 type 与事件名相同。流以 data: [DONE] 结束。

常见事件顺序:

  1. response.created、response.in_progress
  2. response.output_item.added:开始一个输出项(推理、消息或函数调用)。
  3. 内容事件:
    • response.output_text.delta:回复文本片段。
    • response.reasoning_summary_text.delta:推理摘要片段。
    • response.reasoning_text.delta:原始推理文本片段。
    • response.function_call_arguments.delta:函数参数片段。
  4. 对应的 .done 事件和 response.output_item.done。
  5. 终止事件:response.completed、response.incomplete 或 response.failed。
  6. data: [DONE]
event: response.output_text.delta
data: {"type":"response.output_text.delta","item_id":"msg_...","output_index":1,"content_index":0,"delta":"你好","sequence_number":7}

用 OpenAI SDK 读取流:

stream = client.responses.create(model="gpt-5.5", input="写一首两行的诗。", stream=True)
for event in stream:
    if event.type == "response.output_text.delta":
        print(event.delta, end="", flush=True)
    elif event.type == "response.completed":
        print("\n用量:", event.response.usage.total_tokens)
    elif event.type in ("response.failed", "error"):
        print("\n失败:", event)

流中的错误先发送 event: error,再发送 event: response.failed,最后是 data: [DONE]:

event: error
data: {"type":"error","code":"model_not_found","message":"Model not found: no-such","param":null,"sequence_number":1}

event: response.failed
data: {"type":"response.failed","response":{"status":"failed","error":{"code":"model_not_found","message":"Model not found: no-such"}}}

data: [DONE]

客户端应以 response.failed 作为失败的终止信号。

Messages

事件格式与 Anthropic 官方相同。每个事件有 event: 行,流以 message_stop 结束,没有 [DONE]。

event: message_start
data: {"type":"message_start","message":{"id":"msg_...","role":"assistant","content":[],"usage":{"input_tokens":12,"output_tokens":1}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"text_delta","text":"你好"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"end_turn","stop_sequence":null},"usage":{"output_tokens":20}}

event: message_stop
data: {"type":"message_stop"}
  • content_block_delta 的 delta.type 可以是 text_delta、thinking_delta、signature_delta 或 input_json_delta。
  • 块按 index 从 0 开始依次出现,一个块结束后才开始下一个块。
  • 心跳是 event: ping 事件,数据为 {"type":"ping"}。
  • 流中的错误是 event: error 事件,之后连接关闭。

用 Anthropic SDK 读取流:

with client.messages.stream(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "写一首两行的诗。"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

连接中断

  • 流已经开始后,网关不会切换到另一个上游。上游中途失败时,你会收到错误事件。
  • 上游长时间不发数据时,网关按超时处理,并发送错误终止事件。网关不会把截断的回复报告为成功。
  • 客户端断开连接不会取消上游请求。上游正常完成时,网关照常计费,日志中的状态为 client_gone。需要停止生成时,不要依赖断开连接来节省费用。

本页目录