☰
Qwen3.8-Flash-Next 模型解析与实战部署指南:混合注意力架构、N-gram 嵌入与长上下文推理
2026/9/30 5:53:04 网站建设 项目流程
  • 人工智能
  • 基础模型
  • 大模型
  • 多模态

【免费下载链接】Qwen3.8-Flash-Next

项目地址:https://ai.gitcode.com/hf_mirrors/Qwen/Qwen3.8-Flash-Next
点击查看免费下载

本篇技术指南以仓库根目录 README.md 为主体,围绕开源权重模型Qwen3.8-Flash-Next展开,系统讲解其面向 Qwen4 的混合注意力架构(Gated DeltaNet + Qwen Sparse Attention)、Gated Residual、N-gram Embedding 与定制训练配方等核心创新,并结合 config.json、generation_config.json、chat_template.jinja 等仓库配置,完整给出从服务部署、OpenAI 兼容 API 调用(文本/图像/视频)、思考模式控制到超长上下文扩展的端到端实战方案。读完本文,你将掌握该模型 125B 参数的内部结构、各配置文件的关键字段含义,以及在生产环境中配置 1M 上下文与多模态推理的具体做法。

一、仓库与模型定位:一个实验性架构预览

Qwen3.8-Flash-Next是 Qwen 系列中一个以"架构创新"为核心的开源权重版本。仓库本身是一个标准 Hugging Face Transformers 格式的模型仓库,包含:

  • 模型权重:131 个分片,model-00001-of-00131.safetensors至model-00131-of-00131.safetensors,总大小约 360 GB(见 model.safetensors.index.json 中metadata.total_size = 359999963128);
  • 配置与预处理文件:config.json、generation_config.json、tokenizer_config.json、preprocessor_config.json、video_preprocessor_config.json;
  • 对话模板:chat_template.jinja 与 tokenizer 内嵌的聊天模板,二者共同决定推理时的提示词拼装方式。

根据 README.md,这批产物与 Hugging Face Transformers、vLLM、SGLang、TokenSpeed 等主流推理框架兼容;而Qwen3.8-Flash则是基于该开源版本的官方 API 版本,默认提供 1M 上下文与官方内置工具,二者是"开源底座"与"生产化服务"的关系。

值得强调的是,README 明确指出这是一个实验性预览(experimental preview),其架构将是后续 Qwen4 的基础——理解它的动机是"不仅追求规模,更追求效率":在参数规模与上下文窗口不断增长的背景下,架构层面的创新才是可持续进步的方向。

二、四大核心架构创新(Highlights)

README 将 Qwen3.8-Flash-Next 的首发亮点归纳为四点,这四点构成了理解整个 config.json 的钥匙。

2.1 混合注意力:Gated DeltaNet + Qwen Sparse Attention(QSA)

  • Gated DeltaNet(线性注意力):承担绝大部分层(36 层)的序列混合职责,以 O(1) 的线性复杂度维护状态,适合超长上下文下的低成本处理。
  • Qwen Sparse Attention(QSA):与早期的 Gated DeltaNet + Gated Attention 配对不同,新版将稀疏注意力端升级为 QSA——不再逐 token 挑选,而是以微块(micro-block)级别进行选择性处理。这能显著降低长上下文延迟,尤其适合以 Agent 为主导的现实负载。

在 config.json 中,该结构直接体现在layer_types字段上:48 层中依次交替排列linear_attention与full_attention,即"3 层线性注意力 + 1 层全注意力"的周期重复(共 12 组)。这与 README 中Hidden Layout: 12 × (3 × (Gated DeltaNet → MoE) → 1 × (Qwen Sparse Attention → MoE))的描述完全对应。

2.2 Gated Residual:带门控的残差流

深层 LLM 训练之所以可控,依赖的是带归一化的残差流。Gated Residual 在此基础上,通过**逐元素、数据依赖的读门(read gate)与逐分支的标量写门(write gate)**来调制加宽后的残差流信息。其收益是:层间表达更细粒度、训练稳定性保持、推理开销增量很小。

在 config.json 中对应字段为hc_count: 4(分支数 4)与hc_lowrank: 320(瓶颈秩 320),并在 model.safetensors.index.json 中可看到hyper_connection_mixer、attn_hyper_connection等权重分组(含block_inject_weight、input_mix_weight_down/up、hc_norm),即为该门控机制的实际参数载体。

2.3 N-gram Embedding:参数扩展的新轴

词嵌入提供了一条"低计算成本、易于 offload"的参数扩展路径,相比 MoE 更适合显存受限的加速器。实现方式是以短 n-gram 作为索引来扩展参数规模而不牺牲质量。仓库配置中:

  • ngram_vocab_size_base: 20000000(2000 万 n-gram 词表,README 标注为 layer 2 处的 bigrams/trigrams);
  • ngram_size: 3、make_ngram_vocab_size_divisible_by: 128、split_ngram_parts: 128;
  • ple_layer_ids: [2](仅在 layer 2 引入)、ple_embed_dim: 2560、ple_conv_kernel_size: 4。

README 的模型总参数构成即为:125B 总参数 = 6B 激活参数 + 51B n-gram 嵌入 + 4B MTP,可见 n-gram 嵌入占据了大量"非激活"参数。

2.4 定制训练配方:Muon + AdamW 双优化器

README 说明:Muon 与 AdamW 优化器被应用于特定的权重类别以最大化效率;在重拟合缩放定律指导下,取消了传统的 batch size warmup,直接以目标 batch size 启动训练,从而显著减少优化器步数,并支持更大的学习率以稳健收敛。这是"预训练 & 后训练"两阶段产物得以成形的重要工程前提。

三、Model Overview:从 README 到 config.json 的逐项对照

README 给出的模型规格非常详细,下表将其与 config.json 的实际字段一一对应,便于你在部署时核查配置:

维度README 描述config.json 对应字段/值
类型Causal LM with Vision Encoderarchitectures: Qwen4ExpForConditionalGeneration,model_type: qwen4_exp
训练阶段Pre-training & Post-training—
激活/总参数125B 总、6B 激活、51B n-gram、4B MTP见权重索引与ngram_vocab_size_base等
隐藏维度2560hidden_size: 2560
Token 嵌入248320 (Padded)vocab_size: 248320
层数48num_hidden_layers: 48
Hidden Layout12 × (3×(DeltaNet→MoE) → 1×(QSA→MoE))layer_types: 36 个linear_attention+ 12 个full_attention
DeltaNet 头V 48 / QK 16,头维 128linear_num_value_heads: 48、linear_num_key_heads: 16、linear_key_head_dim: 128、linear_value_head_dim: 128、linear_conv_kernel_dim: 4
QSA 头Q 24 / KV 2,头维 256,RoPE 维 64num_attention_heads: 24、num_key_value_heads: 2、head_dim: 256、partial_rotary_factor: 0.25
IndexerMQA(4 查询头 + 1 共享键头),头维 128,预算 512 块 / 2048 tokenindexer_n_heads: 4、indexer_kv_heads: 1、indexer_head_dim: 128、indexer_budget: 2048、indexer_compress_ratio: 4
MoE512 专家,10 路由 + 1 共享,中间维 640num_experts: 512、num_experts_per_tok: 10、moe_intermediate_size: 640、shared_expert_intermediate_size: 640
Gated Residual4 分支,瓶颈秩 320hc_count: 4、hc_lowrank: 320
LM 输出248320 (Padded)lm_head.weight存在于权重索引中
MTP1 层,多步训练mtp_num_hidden_layers: 1、mtp.hybrid: true、mtp.layer_types: [full_attention]
上下文长度原生 262,144,可扩展至 1,000,000max_position_embeddings: 262144、rope_parameters.rope_type: default

几个值得注意的细节:

  • 位置编码:config.json 中rope_parameters默认rope_type: "default"、rope_theta: 10000000、mrope_interleaved: true、mrope_section: [11, 11, 10]——这是多模态 MRoPE 配置,视觉 token 与文本 token 使用不同的旋转维度段。
  • 生成默认参数:generation_config.json 中do_sample: true、temperature: 1.0、top_k: 20、top_p: 0.95,与 README 推荐的思考模式采样参数一致;eos_token_id为[248046, 248044](<|im_end|>与<|endoftext|>)。
  • 对话与思维标签:tokenizer_config.json 中<think>/</think>(id 248068/248069)与工具调用标签<tool_call>/<tool_response>均为普通 token,说明思维链与工具调用是模型语言能力的一部分;<|image_pad|>(248056)与<|video_pad|>(248057)则对应视觉占位符,与 config.json 的image_token_id/video_token_id一致。

四、性能基准解读(Benchmark Results)

README 用两张对照表给出评测数据(数据为仓库声明,非本文实测)。阅读时请注意其评测口径:

语言能力(对照 Qwen3.8-27B、Qwen3.7-Plus、DeepSeek-V4-Flash-0731、Claude-Opus-4.6 (Max)):

  • 在 125B 总参数 / 6B 激活的前提下,多项 Agentic Coding 指标取得该行最优,如 DeepSWE 1.1(58.7)、SWE-bench Pro(62.5)、SWE-bench Multilingual(81.0);
  • Agent 类任务表现突出:CoWorkBench 73.9、JobBench 55.7、Toolathlon Verified(Pass@1)73.5;Agents' Last Exam 的 Pass@1 为 24.3、Score 51.2;
  • 通用能力方面:IFBench 81.3、GPQA Diamond 91.7、LiveCodeBench v6 91.9 均为行内最优;HLE 35.9、NL2Repo-Bench 48.1 略低于行内最高。

视觉语言能力(对照 Qwen3.8-27B、Qwen3.7-Plus、Claude-Opus-4.6 (Max)):多模态 Agent 场景(ClawEval-MM、RecreationBench、AndroidWorld、OSWorld 2.0、Vision2Web)与通用多模态(ERQA、LVBench、RealWorldQA、MathVision、CharXiv)多项取得最优或并列最优。

必须注意的评测条件(README 脚注原文信息):

  • DeepSWE 1.1 与 SWE-bench Multilingual 使用 Claude Code / mini-SWE-agent harness,temp=1.0、top_p=0.95、256K 上下文;
  • SWE-bench Pro 除 Claude-Opus-4.6 使用官方公布分数外,其余模型均用 Claude Code harness 复评;
  • NL2Repo-Bench 为避免 reward hacking,禁用了访问特定仓库的 Bash 命令(pip download、pip install、git clone);
  • CoWorkBench 与 RecreationBench 为内部自研基准;HLE 由 GPT-4o 判分;行内最优以加粗显示。

这些细节提示我们:不同模型、不同 harness、不同采样设置下的分数不可直接横比,引用时应保留脚注条件。

五、快速上手:部署与 API 调用(Quickstart)

5.1 服务端部署(Serving)

README 明确给出建议:推理效率与吞吐在不同框架间差异显著,生产负载或高吞吐场景强烈推荐使用专用服务引擎(SGLang、KTransformers、vLLM),并优先使用各框架最新版本以保证性能与兼容性。相关部署手册(Cookbook / Recipe)由各框架官方提供。该模型也支持 Hugging Face Transformers 直接加载推理。

5.2 API 调用的通用准备

Chat Completions API 可被绝大多数推理框架使用。以 OpenAI 兼容客户端为例,先安装并配置环境变量:

pip install -U openai # Set the following accordingly export OPENAI_BASE_URL="http://localhost:8000/v1" export OPENAI_API_KEY="EMPTY"

模型名使用Qwen/Qwen3.8-Flash-Next。

5.3 思考模式与采样参数(重要)

Qwen3.8-Flash-Next默认以思考模式运行,会在最终回复前生成以<think>\n...</think>\n\n标记的思考内容。README 推荐两套采样参数:

模式temperaturetop_ptop_kmin_ppresence_penaltyrepetition_penalty
思考模式(Thinking)1.00.95200.00.01.0
直答模式(Instruct / Non-Thinking)0.70.80200.01.51.0

注意:各推理框架对采样参数的支持范围不尽相同。同时,README 提醒:在多轮 Agent 任务中,降低推理强度(reasoning effort)不一定缩短总耗时——单轮响应虽快,但分析不足可能导致失败重试,反而增加总延迟与 token 消耗。

模型通过三个参数控制思考行为:enable_thinking、preserve_thinking、reasoning_effort。这三个参数也正是 chat_template.jinja 中模板渲染所读取的核心开关:

  • enable_thinking(默认 true):控制是否生成思考块;当显式置为 false 时,模板在生成提示词末尾仍会注入空思考块<think>\n\n</think>\n\n,以保证推理引擎行为一致;
  • reasoning_effort(默认xhigh,可选xhigh/medium/low):模板据此注入对应的推理指令文本(xhigh 要求"仔细思考、校验关键假设、权衡替代方案";low 要求"思考简短聚焦、直接给出结论");
  • preserve_thinking(默认 true):决定历史消息中的思考块是否保留(详见下文 5.7)。

5.4 文本输入 + 流式输出(含 Usage)

from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [ {"role": "user", "content": "Write a Python function to merge two sorted linked lists."}, ] completion = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=messages, extra_body={ "chat_template_kwargs": { "enable_thinking": True, # on by default "preserve_thinking": True, # on by default }, }, reasoning_effort="xhigh", # xhigh by default; supported levels are xhigh, medium, and low stream=True, stream_options={"include_usage": True}, ) reasoning_content = "" answer_content = "" is_answering = False print("\n" + "=" * 20 + "Reasoning" + "=" * 20 + "\n") for chunk in completion: if not chunk.choices: print("\nUsage:") print(chunk.usage) continue delta = chunk.choices[0].delta if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None: if not is_answering: print(delta.reasoning_content, end="", flush=True) reasoning_content += delta.reasoning_content elif hasattr(delta, "reasoning") and delta.reasoning is not None: if not is_answering: print(delta.reasoning, end="", flush=True) reasoning_content += delta.reasoning if hasattr(delta, "content") and delta.content: if not is_answering: print("\n" + "=" * 20 + "Answer" + "=" * 20 + "\n") is_answering = True print(delta.content, end="", flush=True) answer_content += delta.content messages.append({ "role": "assistant", "content": answer_content, "reasoning_content": reasoning_content, "reasoning": reasoning_content, })

这段代码的要点:流式 chunk 中思考内容可能通过delta.reasoning_content或delta.reasoning两个字段之一携带(兼容不同框架),最终把思考与答案拼回 assistant 消息回传,以维持多轮对话的上下文连续性。

5.5 图像输入(Image Input)

from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [ { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "<your-image-url>" } }, { "type": "text", "text": "The centres of the four illustrated circles are in the corners of the square. The two big circles touch each other and also the two little circles. With which factor do you have to multiply the radii of the little circles to obtain the radius of the big circles?\nChoices:\n(A) $\\frac{2}{9}$\n(B) $\\sqrt{5}$\n(C) $0.8 \\cdot \\pi$\n(D) 2.5\n(E) $1+\\sqrt{2}$" } ] } ] chat_response = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=messages, ) print("Chat response:", chat_response)

图像在提示词中会经由 chat_template.jinja 的render_content宏被替换为<|vision_start|><|image_pad|><|vision_end|>视觉占位符序列,视觉特征则由 Qwen3VL 处理器完成提取。

5.6 视频输入(Video Input)

from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [ { "role": "user", "content": [ { "type": "video_url", "video_url": { "url": "<your-video-url>" } }, { "type": "text", "text": "How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?" } ] } ] chat_response = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=messages, ) # When vLLM is launched with `--media-io-kwargs '{"video": {"num_frames": -1}}'`, # video frame sampling can be configured via `extra_body` (e.g., by setting `fps`). # This feature is currently supported only in vLLM. # # By default, `fps=2` and `do_sample_frames=True`. # With `do_sample_frames=True`, you can customize the `fps` value to set your desired video sampling rate. # chat_response = client.chat.completions.create( # model="Qwen/Qwen3.8-Flash-Next", # messages=messages, # extra_body={ # "mm_processor_kwargs": {"fps": 2, "do_sample_frames": True}, # }, # ) print("Chat response:", chat_response)

视频默认以fps=2抽帧、do_sample_frames=True;如需自定义采样率,可在 vLLM 以--media-io-kwargs '{"video": {"num_frames": -1}}'启动后,通过mm_processor_kwargs里的fps覆盖。视频处理对应的预处理配置见 video_preprocessor_config.json(processor_class: Qwen3VLProcessor、video_processor_type: Qwen3VLVideoProcessor)。

5.7 直答模式与思考保留控制

关闭思考(Instruct 模式)——模型默认先思考后回答,可通过参数直接得到不思考的回答:

from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [ { "role": "user", "content": [ { "type": "image_url", "image_url": { "url": "<your-image-url>" } }, { "type": "text", "text": "Where is this?" } ] } ] chat_response = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=messages, temperature=0.7, top_p=0.8, presence_penalty=1.5, extra_body={ "top_k": 20, "chat_template_kwargs": {"enable_thinking": False}, }, ) print("Chat response:", chat_response)

若使用 Qwen Cloud 的 API,除更换model外,应直接传"enable_thinking": False,而非包在chat_template_kwargs中。

关闭思考保留(Disable Preserved Thinking)——默认情况下,模型会保留所有历史消息中的思考块,形成完整的推理轨迹,这保证了上下文连续性,尤其适合需要决策一致性与减少冗余推理的 Agent 场景,同时改善了 KV cache 利用率。若只想保留最近一条用户消息的思考块,可将preserve_thinking置为 False:

from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [...] chat_response = client.chat.completions.create( model="Qwen/Qwen3.8-Flash-Next", messages=messages, extra_body={ "chat_template_kwargs": {"preserve_thinking": False}, }, ) print("Chat response:", chat_response)

若使用 Qwen Cloud 的 API,应直接传"preserve_thinking": False,而非包在chat_template_kwargs中。

该开关的实现逻辑可在 chat_template.jinja 中找到:模板遍历历史消息时,只有满足preserve_thinking未显式关闭、或该消息位于最后一次用户查询之后时,才会在 assistant 消息中恢复<think>...</think>块。

六、最佳实践(Best Practices)

6.1 采样参数与防重复

复用上文的"思考模式 / 直答模式"两套参数即可。对支持的框架,可将presence_penalty在0~2之间调整以抑制无休止重复;但注意较高的取值偶尔会引发语言混杂与轻微性能下降。

6.2 充足的输出长度(Agent 任务关键)

为了在 Agent 任务中发挥最佳性能,建议在 1M 上下文内,为"内部推理"与"最终输出"分别配置独立的 token 上限:

  • Reasoning Content(思考内容):最大输出 262,144 tokens;
  • Final Response(最终回答):最大输出 131,072 tokens。

这样既为复杂推理留足空间,又保证最终交付物有充分篇幅。

6.3 超长文本处理:YaRN RoPE 缩放

模型原生支持 262,144 tokens;当输入+输出总长度超过该上限时,README 推荐采用YaRN这类 RoPE 缩放技术(vLLM、SGLang、TokenSpeed 均已支持)。开启方式有两种:

方式一:修改模型配置文件。在 config.json 的text_config.rope_parameters中改为:

{ "mrope_interleaved": true, "mrope_section": [ 11, 11, 10 ], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144 }

方式二:启动参数覆盖(无需改文件)。

vLLM:

VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000

SGLang:

SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... --json-model-override-args '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --context-length 1000000

TokenSpeed:

TOKENSPEED_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 tokenspeed serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000

重要注意事项(README 原文要点):

  • 主流开源框架实现的是静态 YaRN,缩放因子不随输入长度变化,可能对较短文本的性能有影响;
  • 因此仅在确实需要处理长上下文时才修改rope_parameters;
  • 建议按实际场景调整factor:例如若应用典型上下文为 524,288 tokens,则把factor设为 2.0 更合适。

6.4 长视频理解优化

为了兼顾纯文本与图像的推理效率,发布版 video_preprocessor_config.json 中的size.longest_edge被保守地配置为 25165824。README 建议:在需要小时级长视频的高帧率采样时,将longest_edge调至469,762,048(对应约 224k 视频 token)以获得更优性能,例如:

{"longest_edge": 469762048, "shortest_edge": 4096}

也可通过引擎启动参数覆盖默认值(实现细节参见 vLLM / SGLang 相关改动)。

6.5 引用(Citation)

若你的工作参考了本模型,README 提供了两条 BibTeX 记录:一条对应架构设计技术报告《On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability》,另一条对应本模型的技术博客《Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency》,可按需引用。

七、总结

Qwen3.8-Flash-Next 是理解"下一代高效 LLM 架构"的绝佳样例:Gated DeltaNet 线性注意力与 QSA 稀疏注意力的混合布局(config.json 中 36 层线性 + 12 层稀疏交替)、Gated Residual 门控残差(hc_count: 4)、2000 万词表的 N-gram 嵌入(ngram_vocab_size_base: 20000000)以及Muon/AdamW 定制训练配方,共同支撑起 125B 总参、6B 激活的高性价比形态。在工程侧,通过enable_thinking/preserve_thinking/reasoning_effort三个参数(其语义与 chat_template.jinja 的模板逻辑一一对应)即可精细控制思考行为,配合 YaRN 配置可将原生 262,144 的上下文扩展到 1M,再结合图像、视频的多模态输入能力,完全具备承载长程 Agent 任务的生产条件。建议读者在部署时结合本仓库的配置文件逐项核对,并优先采用最新版本的专用推理框架以获得最佳吞吐与兼容性。

  • 人工智能
  • 基础模型
  • 大模型
  • 多模态

【免费下载链接】Qwen3.8-Flash-Next

项目地址:https://ai.gitcode.com/hf_mirrors/Qwen/Qwen3.8-Flash-Next
点击查看免费下载

创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询