思维链(推理)
开启和调节模型推理,在界面中展示推理内容,并在多轮对话中正确回传。
推理(reasoning)是模型在给出回答之前生成的思考过程,也叫思维链或“思考”。本页说明如何开启推理、如何读取和展示推理内容,以及多轮对话中要回传哪些字段。
三种推理内容
上游模型返回的推理内容分为三种。网关在三种协议之间转换它们,但不会凭空生成上游没有返回的内容。
| 内容 | 说明 | 常见来源 |
|---|---|---|
| 原始推理文本 | 模型完整的思考文字。 | DeepSeek、GLM、Kimi、Qwen、MiniMax 等模型 |
| 推理摘要 | 模型对思考过程的概括。 | GPT 系列(需要显式请求)、Claude(新模型默认返回摘要式思考) |
| 加密推理 | 不可读的签名或密文,用于下一轮恢复推理状态。 | GPT 系列的 encrypted_content、Claude 的 signature |
能否看到推理文字,取决于模型本身:
- GPT 系列只在请求了摘要时返回可读的推理摘要。没有请求摘要时,只返回加密推理。
- Claude 系列在开启 thinking 后返回思考文字和签名。
- 返回原始推理文本的模型,在思考模式下总会返回推理文字。
开启和调节推理
推理强度
| 协议 | 写法 |
|---|---|
| Chat Completions | "reasoning_effort": "high" |
| Responses | "reasoning": { "effort": "high" } |
| Messages(新 Claude 模型) | "thinking": { "type": "adaptive" },加 "output_config": { "effort": "high" } |
| Messages(旧写法) | "thinking": { "type": "enabled", "budget_tokens": 16384 } |
可用级别为 none、minimal、low、medium、high、xhigh、max。不传推理参数时,上游使用模型默认行为。例如,Claude 模型默认不开启 thinking。
网关把强度转换成上游的格式:
| 上游 | 网关发送的内容 |
|---|---|
| Chat 类上游(DeepSeek 等) | reasoning_effort: <级别> |
| Responses 类上游(GPT 系列) | reasoning: { "effort": <级别> } |
| Claude 新模型(Opus 4.6、Sonnet 4.6 及以后) | thinking: { "type": "adaptive" } 和 output_config: { "effort": <级别> } |
| Claude 旧模型(Opus 4.5、Sonnet 4.5 及以前) | thinking: { "type": "enabled", "budget_tokens": N } |
Claude 旧模型的 budget_tokens 映射:minimal 和 low 为 1024,medium 为 4096,high 为 16384,xhigh 和 max 为 32000。
Messages 的旧写法 budget_tokens 只对 Claude 旧模型有意义。把它发往非 Claude 模型时,网关不设置推理强度,上游使用默认值。要控制非 Claude 模型的强度,请使用 adaptive 加 output_config.effort。
请求推理摘要(GPT 系列)
GPT 系列需要显式请求摘要。网关不会自动添加摘要设置。
| 协议 | 写法 |
|---|---|
| Responses | "reasoning": { "effort": "high", "summary": "auto" } |
| Chat Completions | "reasoning_effort": "high",加 "reasoning": { "summary": "auto" } |
summary 可以是 auto、concise 或 detailed。Messages 协议没有摘要参数。用 Claude Code 等 Messages 客户端调用 GPT 模型时,通常只能看到内容为空的思考块。
OpenAI Python SDK 的 chat.completions.create 没有 reasoning 参数。请用 extra_body 传入:
completion = client.chat.completions.create(
model="gpt-5.5",
messages=[{"role": "user", "content": "9.11 和 9.9 哪个大?"}],
reasoning_effort="high",
extra_body={"reasoning": {"summary": "auto"}},
)关闭推理
| 协议 | 写法 |
|---|---|
| Chat Completions | "reasoning_effort": "none" |
| Responses | "reasoning": { "effort": "none" } |
| Messages | "thinking": { "type": "disabled" } |
none 发往 Claude 时,网关发送 thinking: { "type": "disabled" }。Fable、Mythos、Opus 5.5 及以后的模型不允许关闭推理,网关改为发送最低强度 low。只能推理的模型无法关闭推理。
读取推理内容
Chat Completions
非流式响应中,推理位于 choices[0].message:
{
"role": "assistant",
"content": "你好!",
"reasoning": "用户在打招呼,我简短回应。",
"reasoning_details": [
{ "type": "reasoning.summary", "summary": "用户在打招呼,我简短回应。" },
{ "type": "reasoning.encrypted", "data": "mz2...." }
]
}| 字段 | 内容 |
|---|---|
reasoning_content | 原始推理文本。上游使用这个字段时出现,例如 DeepSeek。 |
reasoning | 可读推理文本。有原始文本时取原始文本,否则取摘要。 |
reasoning_details[] | 结构化推理,按顺序排列。type 为 reasoning.text(字段 text)、reasoning.summary(字段 summary)或 reasoning.encrypted(字段 data)。 |
流式响应中,推理片段位于 choices[0].delta.reasoning_details[],格式与上表相同。如果管理员为某个模型开启了 reasoning_content 兼容输出,同一个 delta 中还会有 reasoning_content。
很多客户端(例如基于 AI SDK @ai-sdk/openai-compatible 的工具)在流式响应中只读取 delta.reasoning_content 或 delta.reasoning。这些客户端可能显示不出 LynShen 默认的 reasoning_details 推理。遇到这种情况,改用 Messages 或 Responses 协议,或联系客服为该模型开启 reasoning_content 兼容输出。
data: {"choices":[{"index":0,"delta":{"reasoning_details":[{"type":"reasoning.text","text":"用户在打招呼。"}]}}]}下面的类从 message 或 delta 中读取可读推理。它先检查 reasoning_content 和 reasoning,再检查 reasoning_details。开启了 reasoning_content 兼容输出时,同一段推理会同时出现在两种字段中。所以一旦读到标量字段,它就忽略后续的 reasoning_details 文本,避免重复显示。
class ReasoningReader:
"""读取 Chat Completions 的可读推理。每个回复使用一个新实例。"""
def __init__(self):
self.scalar_seen = False
def read(self, part) -> str:
"""part 是 choices[0].message 或 choices[0].delta(OpenAI SDK 对象)。"""
for name in ("reasoning_content", "reasoning"):
value = getattr(part, name, None)
if isinstance(value, str) and value:
self.scalar_seen = True
return value
if self.scalar_seen:
return ""
pieces = []
for detail in getattr(part, "reasoning_details", None) or []:
if detail.get("type") == "reasoning.text":
pieces.append(detail.get("text", ""))
elif detail.get("type") == "reasoning.summary":
pieces.append(detail.get("summary", ""))
return "".join(pieces)非流式读取:ReasoningReader().read(completion.choices[0].message)。
流式读取并分开显示推理和回答:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.lynshen.org/v1", api_key=os.environ["LYNSHEN_API_KEY"])
stream = client.chat.completions.create(
model="deepseek-v4-pro",
messages=[{"role": "user", "content": "9.11 和 9.9 哪个大?"}],
reasoning_effort="high",
stream=True,
)
reader = ReasoningReader()
in_reasoning = False
for chunk in stream:
if not chunk.choices:
continue
delta = chunk.choices[0].delta
thought = reader.read(delta)
if thought:
if not in_reasoning:
print("【思考过程】")
in_reasoning = True
print(thought, end="", flush=True)
if delta.content:
if in_reasoning:
print("\n【回答】")
in_reasoning = False
print(delta.content, end="", flush=True)Responses
推理位于 output 中 type 为 reasoning 的项:
{
"type": "reasoning",
"id": "rs_...",
"summary": [{ "type": "summary_text", "text": "**规划** 先比较整数部分。" }],
"content": [{ "type": "reasoning_text", "text": "9.11 的整数部分是 9……" }],
"encrypted_content": "mz2...."
}summary[].text是推理摘要。content[].text是原始推理文本。上游返回原始推理时出现,例如通过 Responses 协议调用 DeepSeek。encrypted_content是加密推理。
流式事件:
| 事件 | 内容 |
|---|---|
response.reasoning_summary_text.delta | 摘要片段,字段 delta。 |
response.reasoning_summary_part.added | 开始一段新的摘要。可以在这里换行。 |
response.reasoning_text.delta | 原始推理文本片段,字段 delta。 |
stream = client.responses.create(
model="gpt-5.5",
input="9.11 和 9.9 哪个大?",
reasoning={"effort": "high", "summary": "auto"},
stream=True,
)
for event in stream:
if event.type in ("response.reasoning_summary_text.delta", "response.reasoning_text.delta"):
print(event.delta, end="", flush=True) # 显示在“思考过程”区域
elif event.type == "response.reasoning_summary_part.added":
print()
elif event.type == "response.output_text.delta":
print(event.delta, end="", flush=True) # 显示在“回答”区域Messages
推理是 content 中的 thinking 块:
{ "type": "thinking", "thinking": "用户在比较两个小数……", "signature": "mz2...." }上游没有返回可读文本时,thinking 为空字符串,只有 signature。流式响应中,content_block_start 的块类型为 thinking,文本通过 thinking_delta 到达,签名通过 signature_delta 到达:
import os
import anthropic
client = anthropic.Anthropic(base_url="https://api.lynshen.org", api_key=os.environ["LYNSHEN_API_KEY"])
with client.messages.stream(
model="claude-sonnet-5",
max_tokens=16000,
thinking={"type": "adaptive"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "9.11 和 9.9 哪个大?"}],
) as stream:
for event in stream:
if event.type == "content_block_delta" and event.delta.type == "thinking_delta":
print(event.delta.thinking, end="", flush=True) # 思考过程
elif event.type == "content_block_delta" and event.delta.type == "text_delta":
print(event.delta.text, end="", flush=True) # 回答thinking.display 可以设为 summarized 或 omitted。设为 omitted 时,模型不返回思考文字,只返回签名。网关原样转发这个值。
在界面中展示推理
- 把推理和回答放在两个区域。推理区域默认折叠,标题写“思考过程”。
- 流式输出时,先显示推理区域。收到第一段回答文本后,自动折叠推理区域。
- 推理文本可能包含 Markdown,例如摘要中的
**标题**。用与回答相同的 Markdown 渲染器显示它。 - 不要显示
reasoning.encrypted、encrypted_content或signature。它们不可读。
一个最小的 HTML 结构:
<details class="reasoning" open>
<summary>思考过程</summary>
<div id="reasoning-text"></div>
</details>
<div id="answer-text"></div>多轮对话中回传推理
下一轮请求要把上一轮的推理原样发回。这在工具调用中尤其重要。
| 协议 | 回传什么 |
|---|---|
| Chat Completions | 把助手消息原样放回 messages,保留 reasoning_details、reasoning_content。 |
| Responses | 把上一轮 output 中的 reasoning 项原样放回 input。无状态调用 GPT 模型时,加 "include": ["reasoning.encrypted_content"],上游才返回加密推理。 |
| Messages | 把 thinking 和 redacted_thinking 块原样放回助手消息,不要修改 signature。 |
网关发出的加密推理和签名以 mz2. 开头。网关在转发前把它还原成上游的原始值。如果你在对话中途换了模型,网关不会把旧模型的加密推理发给新模型。
缺少或修改了签名时,Claude 可能返回 HTTP 400,错误码可能是 thinking_signature_invalid。
用量和计费
推理消耗的 Token 计入输出 Token。
| 协议 | 推理 Token 字段 |
|---|---|
| Chat Completions | usage.completion_tokens_details.reasoning_tokens |
| Responses | usage.output_tokens_details.reasoning_tokens |
| Messages | 包含在 usage.output_tokens 中 |
模型设置了单独的推理价格时,推理 Token 按推理价格计费。否则按输出价格计费。上游不报告推理 Token 时,字段为 0。
常见问题
| 现象 | 原因和处理 |
|---|---|
| GPT 模型没有推理文字 | 请求中加 summary: "auto"。 |
| Claude 没有 thinking 块 | 设置推理强度或 thinking。Claude 默认不思考。 |
开启推理后返回 400,提示 temperature | Claude 开启推理时,temperature 必须为 1 或不传。 |
流式 Chat 中找不到 reasoning_content | 流式推理默认在 delta.reasoning_details[] 中。使用上面的 ReasoningReader 读取。 |
| 关闭推理不生效 | 部分模型只能推理,无法关闭。 |