Monoize
API 使用指南

Responses API

用 OpenAI Responses 协议调用 LynShen,包括多轮对话、推理摘要、内置图片工具和 WebSocket。

Responses 是 OpenAI 的新一代协议。Codex 使用它。需要推理摘要或内置图片生成工具时,也推荐使用它。网关可以把 Responses 请求转发给任意聊天模型,包括 Claude 和国内模型。

  • 端点:POST https://api.lynshen.org/v1/responses
  • 认证:Authorization: Bearer sk-... 或 x-api-key: sk-...

基本请求

curl https://api.lynshen.org/v1/responses \
  -H "Authorization: Bearer $LYNSHEN_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.5",
    "instructions": "你是一名简洁的技术助手。",
    "input": "解释什么是向量数据库。"
  }'

output_text 是 SDK 提供的便捷属性。它拼接 output 中所有 message 项的文本。

常用请求字段

字段说明
model必填。
input字符串,或由消息项和工具结果项组成的数组。
instructions系统指令。
stream设为 true 时返回命名事件的 SSE 流。
max_output_tokens输出上限。
reasoning{ "effort": "high", "summary": "auto" }。见思维链(推理)。
tools、tool_choice、parallel_tool_calls函数工具和内置工具。
text.format结构化输出。
include例如 ["reasoning.encrypted_content"]。网关原样转发,不会自动添加。
store默认 true。设为 false 时,网关不保存本次历史。
previous_response_id续接网关保存的上一轮历史。见下文。

输入格式

input 数组中的消息项可以包含文本和图片:

{
  "model": "gpt-5.5",
  "input": [
    {
      "role": "user",
      "content": [
        { "type": "input_text", "text": "这张图里有什么?" },
        { "type": "input_image", "image_url": "data:image/png;base64,iVBORw0KGgo..." }
      ]
    }
  ]
}

image_url 可以是 HTTPS 地址,也可以是 Base64 数据 URL。

响应

{
  "id": "resp_monoize_4e90e472fa904b7bbd7dd1b28939a162",
  "object": "response",
  "status": "completed",
  "model": "gpt-5.5",
  "output": [
    {
      "type": "reasoning",
      "id": "rs_...",
      "summary": [{ "type": "summary_text", "text": "**规划** 先给出定义,再举例。" }],
      "encrypted_content": "mz2...."
    },
    {
      "type": "message",
      "id": "msg_...",
      "role": "assistant",
      "status": "completed",
      "content": [{ "type": "output_text", "text": "向量数据库……", "annotations": [] }]
    }
  ],
  "usage": {
    "input_tokens": 12,
    "input_tokens_details": { "cached_tokens": 0 },
    "output_tokens": 20,
    "output_tokens_details": { "reasoning_tokens": 8 },
    "total_tokens": 32
  }
}

output 是有序数组。常见的项类型有:

类型内容
message回复文本,位于 content[].text。
reasoning推理摘要(summary[])、原始推理文本(content[])和加密推理(encrypted_content)。
function_call函数调用,包含 call_id、name、arguments。
image_generation_call内置图片工具的结果,result 是 Base64 图片。

status 为 completed、incomplete 或 failed。incomplete 时,查看 incomplete_details。

多轮对话

Responses 有两种续接方式。

方式一:发送完整历史(推荐)

把上一轮 output 中的所有项追加到下一轮的 input 中,再追加新的用户消息。这种方式不依赖网关的会话状态,适合所有场景。

history = [{"role": "user", "content": "我叫小林。"}]
first = client.responses.create(
    model="gpt-5.5",
    input=history,
    include=["reasoning.encrypted_content"],
    store=False,
)
history += first.output  # 原样追加上一轮的全部输出项
history.append({"role": "user", "content": "我叫什么名字?"})
second = client.responses.create(model="gpt-5.5", input=history, store=False)

把 reasoning 项原样传回,模型可以在下一轮沿用推理状态。不要修改 encrypted_content 的值。网关发出的加密推理以 mz2. 开头,网关在转发前把它还原成上游的原始值。

下一轮换成另一个模型时,也可以直接发送这段历史。加密推理只发给产生它的那一类上游,其他上游不会收到它。

方式二:previous_response_id

网关保存每次 store 不为 false 的请求历史。下一轮只发送新输入,并把上一轮的 id 填入 previous_response_id:

{
  "model": "gpt-5.5",
  "previous_response_id": "resp_monoize_4e90e472fa904b7bbd7dd1b28939a162",
  "input": [{ "role": "user", "content": "我叫什么名字?" }]
}

使用前了解这些限制:

  • 历史只属于创建它的 API 密钥。
  • 历史保存在单个网关进程的内存中。历史在 30 分钟后失效,进程重启或缓存淘汰也会清除它。
  • 历史不可用时,网关返回 HTTP 400 previous_response_not_found,param 为 previous_response_id。收到这个错误时,改用方式一重新发送完整历史。
  • 不要同时使用 conversation 和 previous_response_id。

内置图片生成工具

在 tools 中加入 image_generation,模型可以在回复中生成图片:

{
  "model": "gpt-5.5",
  "input": "画一座黎明时分的灯塔。",
  "tools": [{ "type": "image_generation", "size": "1024x1024", "quality": "low" }]
}

图片位于 output 中 type 为 image_generation_call 的项,result 字段是 Base64 数据。这个工具只适用于支持它的上游模型。详见图片生成。

WebSocket 传输

GET /v1/responses 可以升级为 WebSocket 连接。Codex 等客户端用它减少每轮的连接开销。

  • 认证方式与 HTTP 相同。
  • 每条客户端消息是一个 JSON 对象,类型为 response.create。网关用同一套流水线处理它,并把每个 Responses 流事件作为一条 WebSocket 文本消息发回。
  • 同一连接上一次只处理一个回复。连接内的续接状态在连接关闭时删除。
  • 每个连接默认最多 128 轮。超过上限时,网关发送 websocket_connection_limit_reached 错误并关闭连接。客户端应新建连接重试。

GET /v1/codex/responses 与 GET /v1/responses 行为相同。普通应用使用 HTTP 流式输出即可。

压缩会话

POST /v1/responses/compact 把一段长会话压缩成更短的输入。网关把请求转发给上游的同名端点。只有支持这个端点的上游模型可以使用它。Codex 在上下文接近上限时会调用它。

本页目录