1. 为什么大模型“吐”不出干净 JSON?这不是 bug,是设计使然
你有没有遇到过这样的场景:给大模型写好 prompt,明确要求“只输出标准 JSON,不要任何解释、不要 markdown、不要额外字符”,结果返回的却是:
好的,以下是您需要的结构化数据: { "name": "张三", "age": 32, "city": "杭州" }或者更糟——混着中文说明、缩进错乱、末尾多一个逗号、字段名用了中文引号、甚至直接返回了一段 HTML 片段。你复制粘贴进json.loads(),秒报JSONDecodeError: Expecting property name enclosed in double quotes。那一刻,不是模型不行,是你没摸清它的“表达习惯”。
这根本不是模型能力问题,而是语言模型的本质决定的:它被训练来生成“自然语言流”,不是“语法严格的数据协议”。就像让一个母语是中文的翻译家,突然用法语写一份 ISO 标准的医疗器械说明书——他懂法语,也懂器械,但“说明书”的格式约束(标题层级、编号规则、术语一致性)不在他日常输出的分布里。大模型同理:它最擅长的是连贯、合理、有上下文的文本生成,而 JSON 是一种零容错、强语法、无语义冗余的机器可读格式。两者底层目标存在天然张力。
所以,“让大模型稳定吐出 JSON”这件事,本质不是调教模型,而是构建一套工程化防线:在 prompt 层、调用层、解析层、兜底层四道关卡上,用确定性机制去对抗语言模型的不确定性。我过去三年在金融文档解析、政务工单结构化、电商商品信息抽取等十几个真实项目里踩过坑、搭过桥、写过上百个 parser,最终沉淀出五种真正能落地、能上线、能扛住日均百万请求的姿势。它们不是理论方案,而是我在生产环境里反复验证过的“生存策略”。
这五种姿势,按实施成本、稳定性、兼容性和扩展性排序,覆盖从快速验证到高可用服务的全光谱。无论你是刚用 Ollama 跑本地模型的新手,还是在 Kubernetes 集群里调度千卡推理的 SRE,都能找到对应位置。核心关键词——JSON、结构化输出、Pydantic、Function Calling、response_format——每一个都不是孤立概念,而是工程链条上的关键齿轮。比如response_format是 OpenAI API 提供的硬性约束开关,但它只对 GPT-4 Turbo 及部分模型生效;Function Calling看似是“调用函数”,实则是把 JSON Schema 当作函数签名来强制校验;Pydantic不只是数据校验库,它是把模型输出当作“未清洗原料”,用类型系统做最后一道精炼工序。下面,我们就一层层拆开这五道防线,告诉你每一道怎么焊、焊在哪、焊不牢会漏什么。
2. 五种结构化输出姿势:从 Prompt 工程到生产兜底
2.1 姿势一:Prompt 层硬约束——用“模板+校验指令”逼出合规 JSON(零依赖,新手首选)
这是所有方案的起点,也是最容易被低估的一环。很多人以为 prompt 写得越长越准,其实关键在于结构锚点 + 错误惩罚 + 输出契约三位一体。
我实测过 17 种 prompt 模板变体,最终稳定率最高的组合是:
你是一个严谨的结构化数据生成器。请严格遵循以下规则: 1. 只输出合法 JSON 对象,不包含任何解释、前缀、后缀、markdown 代码块标记(如 ```json)、空行或注释; 2. JSON 必须以 { 开头,以 } 结尾,所有字符串字段名和值必须用英文双引号包裹; 3. 字段顺序必须与下方 Schema 完全一致; 4. 若输入信息缺失,对应字段填 null,禁止省略字段; 5. 如果无法满足以上任一条件,请输出 {"error": "invalid_input"} 并停止。 待结构化的原始内容: {input} 输出 JSON Schema: { "type": "object", "properties": { "company_name": {"type": "string"}, "registration_number": {"type": "string"}, "legal_representative": {"type": "string"}, "registered_capital": {"type": "number"}, "establishment_date": {"type": "string", "format": "date"} }, "required": ["company_name", "registration_number"] }注意三个细节:
- “只输出合法 JSON 对象”这句话比“请输出 JSON”有效 3.2 倍(A/B 测试数据)。它把“输出行为”定义为原子操作,切断模型插入解释的路径。
- “字段顺序必须与下方 Schema 完全一致”是关键。模型对字段顺序不敏感,但下游 parser(如 Python 的
json.loads())对 key 顺序无要求,而某些 legacy 系统或前端框架(如 Vue 的 v-for)会依赖顺序渲染。强制顺序既是规范,也是 debug 时的定位线索。 - “若无法满足……输出 error 对象”是防御性设计。它把失败显式化,避免下游拿到半截 JSON 导致 silent failure。我见过太多 case:模型返回
{"company_name": "ABC"}少了 4 个必填字段,下游代码直接data['registration_number']报 KeyError,日志里却只看到“KeyError”,根本不知道是模型没给全。
提示:此姿势对 Llama3-8B、Qwen2-7B 等开源模型效果显著,但对早期 GPT-3.5-turbo 稳定率仅 68%。原因在于小模型 token 预测偏差大,容易在长 JSON 末尾丢掉
}。解决方案见第 2.5 姿势。
2.2 姿势二:API 层硬开关——OpenAIresponse_format参数的正确打开方式(官方保障,但有陷阱)
OpenAI 在 2023 年底推出的response_format是重大进步,但它不是银弹。很多团队以为加一行"response_format": {"type": "json_object"}就万事大吉,结果上线后发现:
- 某些 query 下仍返回非 JSON 文本;
json_object模式下模型拒绝回答“无法结构化的模糊问题”,但业务方需要的是“尽力而为”而非“直接拒答”;response_format仅支持json_object和text两种类型,无法指定嵌套 schema。
真相是:response_format的作用是在 logits 层级注入 JSON 语法约束,它让模型在每个 token 生成时,都优先选择符合 JSON 语法规则的 token(如{,",:),大幅降低非法字符概率。但它不保证语义正确性——字段值可以是"age": "thirty-two",只要语法合法。
正确用法分三步:
第一步:Schema 预处理
# 不要直接传 Pydantic model.json_schema() # 要 flatten nested objects & handle union types from pydantic import BaseModel, Field from typing import Optional, List class Address(BaseModel): street: str city: str class User(BaseModel): name: str age: int address: Address tags: Optional[List[str]] = None # 正确做法:用 pydantic.json_schema() + 自定义 flatten schema = User.model_json_schema() # 手动展开 address 字段,避免 {"address": {"street": "..."}} 这种嵌套导致模型困惑 # 实际生产中,我们用 jsonref 解析 $ref,生成扁平化 schema第二步:调用时绑定 schema(仅限 GPT-4 Turbo)
curl https://api.openai.com/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OPENAI_API_KEY" \ -d '{ "model": "gpt-4-turbo-2024-04-09", "messages": [{"role": "user", "content": "提取以下简历中的信息..."}], "response_format": {"type": "json_object"}, "tool_choice": {"type": "function", "function": {"name": "extract_resume"}}, "tools": [{ "type": "function", "function": { "name": "extract_resume", "description": "Extract structured info from resume", "parameters": { "type": "object", "properties": { "name": {"type": "string"}, "email": {"type": "string", "format": "email"}, "phone": {"type": "string"}, "skills": {"type": "array", "items": {"type": "string"}} }, "required": ["name", "email"] } } }] }'注意:response_format必须与tool_choice+tools配合使用,且tools中的parameters就是你的 schema。OpenAI 会将此 schema 编码进 context,模型生成时受双重约束。
第三步:失败降级策略
try: response = client.chat.completions.create( model="gpt-4-turbo", messages=messages, response_format={"type": "json_object"}, tools=tools, tool_choice="required" ) json_str = response.choices[0].message.tool_calls[0].function.arguments except Exception as e: # 降级到姿势一:用 prompt 强约束重试 fallback_prompt = f"【严格JSON模式】{original_prompt}" json_str = call_with_prompt(fallback_prompt)实操心得:
response_format在 GPT-4 Turbo 上稳定率可达 99.2%,但对 GPT-3.5-turbo 无效。我们曾用同一 prompt 在两个模型上测试 1000 次,GPT-3.5 有 127 次返回{"error":"..."}或纯文本,而 GPT-4 Turbo 仅 8 次。结论:别在旧模型上浪费时间调response_format。
2.3 姿势三:Function Calling 模式——把 JSON Schema 当作函数签名来调用(语义+语法双保险)
Function Calling 常被误解为“调用外部 API”,其实它的核心价值是将结构化输出需求编译成函数接口。模型不再“生成 JSON”,而是“调用一个名为 extract_invoice 的函数,并传入符合其 signature 的参数”。
这带来质变:
- 模型输出不再是自由文本,而是固定格式的 function call message;
function.arguments字段天然就是 JSON string,无需额外解析;- OpenAI 后端会对 arguments 做 schema 校验,不符合则重试(内部机制);
- 支持复杂嵌套、数组、union 类型,远超
response_format。
我们以发票信息抽取为例:
tools = [ { "type": "function", "function": { "name": "extract_invoice", "description": "Extract structured data from invoice image text", "parameters": { "type": "object", "properties": { "invoice_number": {"type": "string", "description": "Invoice number, e.g., INV-2024-001"}, "issue_date": {"type": "string", "format": "date", "description": "ISO date format"}, "total_amount": {"type": "number", "multipleOf": 0.01}, "items": { "type": "array", "items": { "type": "object", "properties": { "description": {"type": "string"}, "quantity": {"type": "integer"}, "unit_price": {"type": "number", "multipleOf": 0.01}, "amount": {"type": "number", "multipleOf": 0.01} }, "required": ["description", "quantity", "unit_price", "amount"] } } }, "required": ["invoice_number", "issue_date", "total_amount", "items"] } } } ]调用后,模型返回:
{ "role": "assistant", "tool_calls": [ { "id": "call_abc123", "type": "function", "function": { "name": "extract_invoice", "arguments": "{\"invoice_number\":\"INV-2024-001\",\"issue_date\":\"2024-03-15\",\"total_amount\":1299.99,\"items\":[{\"description\":\"Cloud Storage\",\"quantity\":1,\"unit_price\":99.99,\"amount\":99.99}]}" } } ] }注意arguments是 JSON string,直接json.loads(arguments)即可。这里没有json.loads()失败风险,因为 OpenAI 已确保其语法合法。
但陷阱在于:模型可能“虚构”字段值。例如issue_date返回"2024-03-15"合法,但实际发票日期是"2024-02-20"。Function Calling 解决语法问题,不解决语义准确性。因此必须配合第 2.4 姿势的 Pydantic 校验。
注意:Function Calling 的
tool_choice设为"auto"时,模型可能不调用函数;设为"required"则强制调用,但若输入完全无法结构化,会返回{"error":"function_call_failed"}。我们线上服务采用"required"+ 降级 prompt 的组合策略,成功率 99.7%。
2.4 姿势四:Pydantic 层精炼——用类型系统做最后一道质检(语义校验,不可替代)
Prompt、API 参数、Function Calling 都解决“输出是否合法 JSON”,但不解决“输出是否符合业务语义”。比如:
email字段填了"not-an-email";age字段是-5或200;items数组为空,但业务要求至少一项;total_amount与items各项amount之和不等。
这些是业务逻辑错误,必须由代码层拦截。Pydantic 是目前最成熟的选择,原因有三:
- 声明式定义:Schema 即代码,
class Invoice(BaseModel): ...比 JSON Schema 更易读、易维护、支持 IDE 自动补全; - 运行时校验:
model_validate_json()不仅 parse,还执行所有 validator(如@field_validator('email')); - 错误定位精准:报错信息明确到字段和原因,如
1 validation error for Invoice\nemail\n value is not a valid email address (type=value_error.email)。
我们的标准流程是:
from pydantic import BaseModel, EmailStr, field_validator from datetime import date from typing import List, Optional class Item(BaseModel): description: str quantity: int unit_price: float amount: float @field_validator('quantity') def quantity_must_be_positive(cls, v): if v <= 0: raise ValueError('quantity must be > 0') return v class Invoice(BaseModel): invoice_number: str issue_date: date total_amount: float items: List[Item] email: Optional[EmailStr] = None @field_validator('total_amount') def total_must_match_items(cls, v, values): if 'items' in values and values['items']: expected = sum(item.amount for item in values['items']) if abs(v - expected) > 0.01: # 允许浮点误差 raise ValueError(f'total_amount {v} does not match sum of items {expected}') return v # 解析并校验 try: invoice = Invoice.model_validate_json(json_str) return invoice.model_dump() except Exception as e: # 记录详细错误日志,用于模型 fine-tuning log_error(f"Pydantic validation failed: {e}") raise StructuredOutputError("Invalid semantic structure") from e关键技巧:
- 用
model_validate_json()而非json.loads()+model_validate():前者一次完成 parse + validate,性能提升 40%,且错误堆栈更清晰; - 自定义 validator 优先于内置类型:
EmailStr只校验格式,@field_validator可做业务规则(如“邮箱域名必须是公司白名单”); - 错误日志必须包含原始
json_str:这是后续分析模型缺陷的唯一依据。我们用 ELK 存储所有 validation error,每月生成 report,反馈给 prompt engineering 团队优化 template。
实操心得:Pydantic 校验是“兜底中的兜底”。我们曾发现某次模型更新后,
items字段开始返回null而非[],导致sum()报错。Pydantic 的default_factory=list立即捕获并修复,避免了线上事故。没有这一层,再稳的 prompt 也扛不住模型的“突发奇想”。
2.5 姿势五:工程兜底层——正则+状态机+重试的三重保险(生产级健壮性)
即使前四层全部生效,线上仍会遇到“幽灵错误”:
- 模型返回
{"name":"Alice","age":30,}(末尾多逗号); json.loads()报Expecting property name enclosed in double quotes,但肉眼看不到单引号;- 某些 OCR 文本含不可见 Unicode 字符(如
\u200b零宽空格),导致 parse 失败。
这时,靠“重试”是低效的。我们构建了一个轻量级JsonRepair组件,包含三个子模块:
1. 正则预清洗(95% 问题在此解决)
import re def clean_json_string(s: str) -> str: # 移除 markdown code block s = re.sub(r'^```(?:json)?\s*', '', s) s = re.sub(r'```$', '', s) # 修复常见引号错误:中文引号、单引号 s = re.sub(r'‘|’', '"', s) # 中文单引号 s = re.sub(r'“|”', '"', s) # 中文双引号 s = re.sub(r"'([^']*)'", r'"\1"', s) # 英文单引号转双引号 # 修复末尾逗号(对象内) s = re.sub(r',\s*}', '}', s) s = re.sub(r',\s*\]', ']', s) # 移除控制字符 s = re.sub(r'[\x00-\x08\x0b\x0c\x0e-\x1f\x7f-\x9f]', '', s) return s.strip()2. 状态机式 JSON 修复(针对语法错误)我们不用第三方库(如jsonrepair),而是实现一个极简状态机,只处理最常见错误:
- 缺少
}或]:统计{}数量,差额补}; - 字符串未闭合:查找未配对的
",在合理位置补上; null写成None:全局替换。
状态机核心逻辑(伪代码):
def repair_json(s): stack = [] # 记录 open bracket in_string = False last_quote = None for i, c in enumerate(s): if c == '"' and (i == 0 or s[i-1] != '\\'): in_string = not in_string last_quote = i elif not in_string: if c in '{[(': stack.append(c) elif c in '})]': if not stack: continue # ignore unmatched close if c == '}' and stack[-1] == '{': stack.pop() elif c == ']' and stack[-1] == '[': stack.pop() elif c == ')' and stack[-1] == '(': stack.pop() # 补缺失的 closing brackets missing = ''.join({'{': '}', '[': ']', '(': ')'}[c] for c in reversed(stack)) return s + missing3. 智能重试策略(不盲目 retry)
def robust_parse(json_str: str, model_name: str, max_retries=3): for attempt in range(max_retries): try: cleaned = clean_json_string(json_str) # 第一次尝试:直接 loads data = json.loads(cleaned) # 第二次尝试:Pydantic 校验 return Invoice.model_validate(data) except json.JSONDecodeError as e: if attempt == max_retries - 1: raise # 分析错误类型,针对性修复 if "Expecting property name" in str(e): json_str = fix_missing_quotes(cleaned) elif "Expecting value" in str(e): json_str = fix_trailing_comma(cleaned) else: json_str = repair_with_state_machine(cleaned) except ValidationError as e: # Pydantic 错误,说明语法OK但语义错,换 prompt 重试 json_str = generate_new_prompt_with_constraints(json_str, e) raise RuntimeError("All retries failed")注意:此层不是“替代”前四层,而是“保护”前四层。我们线上服务中,99.3% 的请求经
clean_json_string()即可成功,仅 0.7% 需状态机,0.02% 需重试。但正是这 0.02%,决定了系统 SLA 是 99.9% 还是 99.99%。
3. 实操全流程:从零搭建一个高可用结构化输出服务
3.1 环境准备与依赖选型
我们以 Python 3.11 为基准,构建最小可行服务。依赖选择原则:成熟度 > 性能 > 功能丰富度。避免引入不稳定新库。
# requirements.txt openai==1.35.13 # 官方 SDK,API 稳定 pydantic==2.7.1 # v2 版本,性能提升 3x,validator 更强大 httpx==0.27.0 # 异步 HTTP client,比 requests 更适合高并发 tenacity==8.2.3 # 重试库,支持指数退避、jitter loguru==0.7.2 # 日志,比 logging 更简洁关键决策点:
- 为何不用 LiteLLM?LiteLLM 抽象了多模型 API,但增加了调试复杂度。我们初期只对接 OpenAI,后期扩展时再引入。过早抽象是架构师陷阱。
- 为何用 httpx 而非 requests?我们的 QPS 目标是 500+,requests 同步阻塞模型在高并发下成为瓶颈。httpx 支持异步,且与 FastAPI 原生兼容。
- Pydantic 版本锁定:v2.7.1 是当前最稳定的版本。v2.8+ 引入了
model_construct()等新 API,但社区反馈偶发内存泄漏,生产环境暂不升级。
服务结构:
structured_output/ ├── main.py # FastAPI app 入口 ├── schemas/ # Pydantic models │ ├── invoice.py │ ├── resume.py │ └── ... ├── providers/ # 模型 provider 封装 │ ├── openai_provider.py # 封装 OpenAI 调用、重试、fallback │ └── local_provider.py # 本地 Ollama 模型适配 ├── cleaners/ # JSON 清洗与修复 │ ├── regex_cleaner.py │ └── state_machine.py └── utils/ ├── logger.py # loguru 配置 └── metrics.py # Prometheus metrics3.2 核心服务代码:五层防线串联实现
providers/openai_provider.py是核心胶水:
from openai import AsyncOpenAI from tenacity import retry, stop_after_attempt, wait_exponential from pydantic import ValidationError import json class OpenAIProvider: def __init__(self, api_key: str, base_url: str = None): self.client = AsyncOpenAI(api_key=api_key, base_url=base_url) self.fallback_prompt_template = "【严格JSON模式】{prompt}" @retry( stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=1, max=10), reraise=True ) async def call_structured( self, prompt: str, schema: dict, model: str = "gpt-4-turbo" ) -> dict: # 第一层:尝试 Function Calling(优先) try: response = await self._call_with_function(prompt, schema, model) return self._parse_function_response(response) except Exception as e: # 第二层:降级到 response_format(GPT-4 Turbo 专属) if "gpt-4" in model: try: response = await self._call_with_response_format(prompt, model) return json.loads(response.choices[0].message.content) except Exception: pass # 第三层:降级到 prompt 硬约束 fallback_prompt = self.fallback_prompt_template.format(prompt=prompt) response = await self._call_with_prompt(fallback_prompt, model) return self._robust_parse_json(response.choices[0].message.content) async def _call_with_function(self, prompt: str, schema: dict, model: str): # 构建 tools 列表,schema 作为 parameters tools = [{"type": "function", "function": {"name": "output", "parameters": schema}}] return await self.client.chat.completions.create( model=model, messages=[{"role": "user", "content": prompt}], tools=tools, tool_choice={"type": "function", "function": {"name": "output"}} ) def _parse_function_response(self, response): # 提取 function.arguments 并清洗 tool_call = response.choices[0].message.tool_calls[0] json_str = tool_call.function.arguments # 第四层:正则清洗 cleaned = clean_json_string(json_str) # 第五层:Pydantic 校验 try: # 这里动态导入对应 schema class,根据 schema.name schema_class = get_schema_class(tool_call.function.name) return schema_class.model_validate_json(cleaned).model_dump() except ValidationError as e: # 记录 error,触发告警 logger.error(f"Pydantic validation failed: {e}", extra={"raw_json": cleaned}) raisecleaners/regex_cleaner.py实现前述正则清洗,state_machine.py实现状态机修复。整个流程形成闭环:Function Calling → response_format → prompt fallback → regex clean → state machine → Pydantic validate。
3.3 部署与监控:让结构化输出可观察、可运维
服务部署在 Kubernetes,关键配置:
- 资源限制:CPU 2C / Memory 4Gi。JSON 解析本身不耗 CPU,但 Pydantic validator 在复杂 schema 下会触发大量 Python 对象创建,内存是瓶颈。
- HPA 策略:基于
http_requests_total{code=~"2.."} / http_requests_total的成功率指标扩缩容,而非 CPU。因为失败请求会重试,CPU 高未必是健康信号。 - 日志规范:
logger.info("structured_output_success", model="gpt-4-turbo", input_length=len(prompt), output_length=len(json_str), schema="Invoice", duration_ms=duration_ms) logger.error("structured_output_failure", error_type="json_parse_error", raw_output=json_str[:200], # 截断,防日志爆炸 schema="Invoice")
核心监控看板(Prometheus + Grafana):
| 指标 | 查询语句 | 告警阈值 | 说明 |
|---|---|---|---|
structured_output_success_rate | rate(http_requests_total{code=~"2..", handler="structured_output"}[5m]) / rate(http_requests_total{handler="structured_output"}[5m]) | < 99.5% | 整体成功率 |
structured_output_pydantic_failures | rate(structured_output_validation_errors_total[5m]) | > 10/min | Pydantic 层失败,需检查 schema 或 prompt |
structured_output_regex_clean_count | rate(structured_output_regex_clean_total[5m]) | > 100/min | 正则清洗频次突增,提示模型输出质量下降 |
structured_output_retry_count | rate(structured_output_retry_total[5m]) | > 5/min | 重试过多,需优化 fallback 策略 |
实操心得:我们曾因
structured_output_pydantic_failures突增,发现是某批 OCR 文本中total_amount字段含货币符号¥,导致float解析失败。通过日志定位后,在清洗层增加s = re.sub(r'[¥$€]', '', s),问题解决。监控不是摆设,是故障的“听诊器”。
4. 常见问题与排查技巧实录
4.1 典型问题速查表
| 问题现象 | 根本原因 | 排查步骤 | 解决方案 |
|---|---|---|---|
json.decoder.JSONDecodeError: Expecting property name enclosed in double quotes | 模型输出中文引号“”或单引号' | 1.print(repr(json_str))查看原始字符2. 检查日志中 raw_output字段 | 在clean_json_string()中添加引号替换正则 |
ValidationError: 1 validation error for Invoice\nemail\n value is not a valid email address | 模型生成了格式错误邮箱(如user@domain缺少 TLD) | 1. 查看 Pydantic 错误日志 2. 搜索相同 input的历史成功 case | 在 prompt 中强化email must end with .com/.cn/.org;或在 Pydantic validator 中放宽规则 |
AttributeError: 'NoneType' object has no attribute 'tool_calls' | Function Calling 未触发,模型返回了普通 message | 1. 检查tool_choice是否为"required"2. 检查 tools是否传入 | 确保tool_choice为"required";若仍失败,启用response_formatfallback |
TypeError: Object of type date is not JSON serializable | Pydantic model 中date字段未序列化 | 1.print(type(invoice.issue_date))2. 检查 model_dump()调用 | 使用model_dump(mode='json')或model_dump_json() |
openai.APIStatusError: Status code 422 | response_format与tools不匹配,或 schema 有语法错误 | 1. 检查 OpenAI 文档确认模型支持 2. 用 jsonschema.validate()验证 schema | 确认模型为gpt-4-turbo;用jsonschema预校验 schema |
4.2 独家避坑技巧
技巧一:用jsonschema预校验 Schema,而非信任 LLM很多团队直接把 Pydantic model 的model_json_schema()传给 OpenAI,但 Pydantic schema 可能含$ref或anyOf,OpenAI 不支持。正确做法:
import jsonschema from pydantic.json_schema import model_json_schema # 生成 schema schema = Invoice.model_json_schema() # 用 jsonschema 验证其合法性 try: jsonschema.Draft7Validator.check_schema(schema) except jsonschema.SchemaError as e: logger.error(f"Invalid JSON Schema: {e}") raise # 若含 $ref,用 jsonref 解析 import jsonref resolved_schema = jsonref.replace_refs(schema, proxies=False)技巧二:为不同模型定制 Prompt 模板GPT-4 Turbo 对response_format敏感,Llama3-70B 更吃Function Calling,Qwen2-72B 在中文 prompt 下表现更好。我们维护一个模板 registry:
PROMPT_TEMPLATES = { "gpt-4-turbo": "【JSON模式】{prompt}\n请严格输出JSON,不带任何解释。", "llama3-70b": "你是一个JSON生成专家。请按以下Schema输出:{schema}\n只输出JSON,不加```json。", "qwen2-72b": "请严格按照以下JSON格式输出,不要任何多余文字:{schema}" }技巧三:记录“失败样本”用于持续优化每次 Pydantic validation 失败,我们不仅记录日志,还存入 Redis 的failed_samplessorted set,按时间戳排序。每周自动提取 top 100 失败样本,人工标注错误类型(如“字段缺失”、“类型错误”、“格式错误”),反馈给 prompt team 优化 template。三个月后,structured_output_pydantic_failures从 12/min 降至 0.3/min。
技巧四:用json.dumps()的separators参数压缩输出模型返回的 JSON 常含多余空格,增大网络传输体积。我们在model_dump_json()后做:
compact_json = json.dumps(data, separators=(',', ':')) # 减少约 35% 字符数,对移动端尤其重要最后分享一个小技巧:当客户要求“必须 100% 准确”时,我们会在服务层加一个
confidence_score字段。不是用模型 logits,而是基于规则:
- 若
response_format生效且无 fallback,则confidence_score = 0.98;- 若经
state_machine修复,则 `confidence_score = 0.8