- SAST
- 应用安全
- 静态分析
- 开发工具
- 代码质量
【免费下载链接】semgrep
Lightweight static analysis for many languages. Find bug variants with patterns that look like source code.
本篇技术指南聚焦 Semgrep 仓库中一处小而关键的工程决策:在 cli/src/semdep/external/parsy/ 目录下,Semgrep 将 Python 解析器组合库 parsy 以 vendor 形式内嵌,并为其加入"增量行号与列号追踪"能力,以显著提升依赖锁文件(lockfile)解析速度。读完本文,你将理解offset / line / column三元组索引的改造原理、字符串输入下逐字符增量更新的实现细节、非字符串输入的特殊约定,以及行号信息如何被下游 SCA(软件组成分析)解析器用于错误报告与依赖定位,并掌握这一改造与上游合并计划的来龙去脉。
一、背景:为什么 Semgrep 需要内嵌(Vendor)一个第三方解析库
Semgrep 的依赖解析子系统semdep(源码位于 cli/src/semdep/)负责扫描各类生态系统的lockfile与manifest文件,例如poetry.lock、Pipfile.lock、yarn.lock、go.mod、Gemfile.lock等。从仓库的目录结构可以看到,cli/src/semdep/parsers/ 下按生态拆分了大量解析器(pipfile.py、poetry.py、yarn.py、go_mod.py、gradle.py、mix.py、swiftpm.py、gem.py、pom_tree.py、requirements.py等),它们共同使用 cli/src/semdep/parsers/util.py 提供的通用工具函数。
这些解析器并不依赖项目自研的解析框架,而是构建在parsy——一个经典的 Python 解析器组合子(parser combinator)库之上。parsy 的核心思想是"通过组合小解析器来构造复杂解析器",仓库中的 cli/src/semdep/external/parsy/README.rst 明确写道:Parsy 是一个"组合小解析器成为复杂、更大解析器"的文本解析库,属于 LL(infinity) 文法的单子式解析器组合子库。
VENDOR_README.md(即本次讨论的关联文档)点明了内嵌的根本动机:
Parsy is vendored in order to add incremental line and column tracking functionality, which speeds up our lockfile parsers significantly.
即:内嵌 parsy 的目的是添加"增量式行号与列号追踪"功能,从而显著加速 Semgrep 的锁文件解析器。由于这一改动尚未被上游接受,Semgrep 选择将带补丁的 parsy 直接放入自己的源码树(即 vendor),而不是通过 PyPI 依赖原版。类似的 vendoring 策略在仓库中还有先例,例如 cli/src/semdep/external/packaging/ 同样内嵌了packaging库(含specifiers.py、version.py、tags.py等模块)。
从源码结构看,cli/src/semdep/external/parsy/ 一共包含 5 个文件:
| 文件 | 作用 |
|---|---|
VENDOR_README.md | 说明 vendoring 原因、改动设计与上游合并计划 |
__init__.py | 改造后的 parsy 核心实现(解析器、组合子、Position 等) |
__init__.pyi | Semgrep 自写的类型桩(type stubs),约束解析器输入类型 |
version.py | 版本号,当前为"2.0"(cli/src/semdep/external/parsy/version.py) |
README.rst | 上游 parsy 的原始项目说明文档 |
二、核心改造:从"单一整数索引"到"offset / line / column 三元组"
原版 parsy 在解析字符串时,用一个整数索引(integer index)记录"流中已消费了多少个元素"。Semgrep 的定制版本将其替换为三个整数组成的索引三元组,分别表示 offset(偏移)、line(行号)、column(列号)。VENDOR_README.md 对此给出了精确的定义:
The edits replace the integer index into a stream with a triple of integers representing offset, line and column information.
offset is what index originally was: how many individual elements of the stream have been consumed.
即offset就是原索引的含义:流的单个元素被消费的个数;line与column则是新增的行列信息。这一改造在 cli/src/semdep/external/parsy/init.py 中落地为不可变数据类:
@dataclass(frozen=True) class Position: offset: int line: int column: intPosition贯穿整个解析流程:Parser的包装函数签名变为接收(stream, index: Position);Result中代表解析位置、错误最远位置的字段也从整数升级为Position:
@dataclass(frozen=True) class Result: status: bool index: Position value: Any furthest: Position expected: FrozenSet[str] @staticmethod def success(index, value): return Result(True, index, value, Position(-1, -1, -1), frozenset()) @staticmethod def failure(index, expected): return Result(False, Position(-1, -1, -1), None, index, frozenset([expected]))这里可以观察到两个细节:成功结果用Position(-1, -1, -1)占位furthest(失败位置字段),失败结果则用Position(-1, -1, -1)占位index,而把真正的失败位置放进furthest。Result.aggregate负责在alt(选择)等组合子中"合并"多个候选解析器的失败信息,保留最远失败位置并合并expected集合,这正是 parsy 能输出"expected ... at line:column"式错误信息的基础。
对应地,ParseError的line_info()方法会按输入类型决定展示方式:
def line_info(self): if isinstance(self.stream, str): return f"{self.index.line}:{self.index.column}" else: return str(self.index.offset)即字符串输入报错时显示line:column,否则退化为仅显示offset。
三、增量更新机制:make_index_update 与逐字符行进
三元组索引只是数据结构层面的一半,另一半在于如何高效地更新它。VENDOR_README.md 描述了更新策略:
When parser input is a string line and column are updated incrementally with each character of the input string.
即在字符串输入下,line与column会随输入串的每一个字符被增量地更新。核心实现是make_index_update工厂函数(cli/src/semdep/external/parsy/init.py):
def make_index_update(consumed: str) -> Callable[[Position], Position]: slen = len(consumed) line_count = consumed.count("\n") last_nl = consumed.rfind("\n") return lambda index: Position( offset=index.offset + slen, line=index.line + line_count, column=slen - (last_nl + 1) if last_nl >= 0 else index.column + slen, )其数学含义非常清晰:
offset:直接累加已消费字符串的长度slen;line:累加已消费字符串中换行符\n的数量line_count;column:分两种情况——若已消费串中包含换行(last_nl >= 0),则列号重置为最后一个换行之后剩余的字符数slen - (last_nl + 1);若不包含换行,则列号直接累加slen。
这样,每次消费一段字符都只需做常数级的字符串统计(count/rfind),无需回溯整个已解析前缀重新计算行列。这也是"incremental(增量式)"命名的由来:与"每次需要行列号时都重新扫描输入串"的做法相比,增量更新将锁文件解析的行列计算开销摊薄到每次消费中,从而显著提升解析性能——这正是 VENDOR_README.md 强调的加速效果来源。
该工厂函数被三类基础解析器复用:
1.string解析器——精确匹配一段字符串后,用index_update(index)更新位置:
def string(expected_string: str, transform: Callable[[str], str] = noop) -> Parser: slen = len(expected_string) transformed_s = transform(expected_string) index_update = make_index_update(expected_string) @Parser def string_parser(stream, index): if transform(stream[index.offset : index.offset + slen]) == transformed_s: return Result.success(index_update(index), expected_string) else: return Result.failure(index, expected_string) return string_parser2.regex解析器——正则匹配成功后,对"实际匹配到的子串"做增量更新:
def regex_parser(stream, index): match = exp.match(stream, index.offset) if match: index = ( make_index_update(stream[match.start() : match.end()])(index) if isinstance(stream, str) else Position(match.end(), -1, -1) ) return Result.success(index, match.group(*group))注意正则匹配的长度在匹配前是未知的,因此这里用stream[match.start():match.end()]取出真实匹配串再交给make_index_update计算行列增量。
3.test_item/test_char解析器——逐字符测试时按单字符更新:
if isinstance(stream, str): index = make_index_update(item)(index) else: index = Position(index.offset + 1, index.line, index.column)由于make_index_update只处理字符串,非字符串路径需要单独分支,这正好引出下一节的约定。
四、非字符串输入的特殊约定:line / column 恒为 -1
VENDOR_README.md 还规定了一条重要约定:
When parser input is not a string (which can never happen with our type stubs) line and column are both set to -1 and ignored.
即当解析器的输入不是字符串(例如字节流或 token 列表)时,line与column一律设置为-1并被忽略。文档特别补充说明:在 Semgrep 的类型桩约束下,这种情况实际上永远不会发生——cli/src/semdep/external/parsy/init.pyi 中的parse、parse_partial等签名都把stream限定为str,从静态类型层面保证了调用方只传入字符串。
这个约定在parse_partial的入口处落地:
def parse_partial(self, stream: str | bytes | list): result = self( stream, Position(0, 0, 0) if isinstance(stream, str) else Position(0, -1, -1), ) ...- 字符串输入:起点为
Position(0, 0, 0)(offset、line、column 均从 0 开始); - 非字符串输入:起点为
Position(0, -1, -1),即 offset 正常累加,而 line / column 保持-1表示"不可用"。
前文展示的regex_parser与test_item_parser中,非字符串分支分别构造Position(match.end(), -1, -1)与Position(index.offset + 1, index.line, index.column),正是为了维持这一约定。而ParseError.line_info()中isinstance(self.stream, str)的分支判断,则确保错误信息对非字符串输入回退为纯 offset 展示。
五、行号信息的消费端:错误报告与依赖定位
改造的最终目的,是让下游解析器在产出依赖(dependency)和报错时都能拿到精确的行号。这一能力在 cli/src/semdep/parsers/util.py 中被系统性消费。
5.1 零基到一基的转换:line_number
parsy 的行列是**零基(zero-indexed)**的,而编辑器和开发者习惯一基(1-indexed)编号。util.py 通过line_info组合子做了显式转换:
# parsy line and column numbers are zero indexed, but editors are generally 1 indexed # so we add one to the line number account for this, and discard the column number line_number = line_info.map(lambda t: t[0] + 1)其中line_info是 vendored parsy 提供的原语解析器,返回当前(line, column)二元组(见 cli/src/semdep/external/parsy/init.py):
line_info = Parser(lambda _, index: Result.success(index, (index.line, index.column)))5.2 mark_line:为每条解析结果标注行号
mark_line在运行目标解析器p之前先取当前行号,产出一对(line, result):
def mark_line(p: Parser[A]) -> Parser[tuple[int, A]]: """ Returns a parser which gets the current line number, runs [p] and then produces a pair of the line number and the result of [p] """ return line_number.bind(lambda line: p.bind(lambda x: success((line, x))))它是解析器文档与锁文件条目行号信息的通用来源。以 cli/src/semdep/parsers/poetry.py 为例,poetry_dep用mark_line包裹每个[[package]]块的键值对解析,并把行号存入ValueLineWrapper:
poetry_dep = mark_line( string("[[package]]\n") >> key_value_list.map( lambda x: { key_val[0]: ValueLineWrapper(line_number, key_val[1]) for line_number, key_val in x } ) )最终,parse_poetry将dep["name"].line_number写入FoundDependency.line_number字段,把"依赖出现在 poetry.lock 的哪一行"带进输出:
output.append( FoundDependency( package=dep["name"].value.lower(), version=dep["version"].value, ... line_number=dep["name"].line_number, lockfile_path=Fpath(str(lockfile_path)), manifest_path=Fpath(str(manifest_path)) if manifest_path else None, ... ) )cli/src/semdep/parsers/pipfile.py 也采用同一模式:用mark_line(key_value).sep_by(new_lines)解析Pipfile清单的键值对,并用key_value_list标注行号;而对Pipfile.lock(JSON 格式)则复用 util.py 中基于mark()的 JSON 解析器,直接将dep_json.line_number写入FoundDependency。
5.3 错误报告:ParseError 到 DependencyParserError
在 util.py 的parse_dependency_file中,捕获ParseError后会把零基的line / column转为一基,并连同出错行原文一并封装进DependencyParserError:
except ParseError as e: # These are zero indexed but most editors are one indexed line, col = e.index.line, e.index.column text_lines = text.splitlines() + ( ["<trailing newline>"] if text.endswith("\n") else [] ) error_str = parse_error_to_str(e) if line < len(text_lines): offending_line = text_lines[line] return DependencyParserError( out.Fpath(str(file_to_parse.path)), file_to_parse.parser_name, error_str, line + 1, col + 1, offending_line, )这段代码完整利用了改造带来的行列信息:不仅能告诉用户"哪一行哪一列出错",还能直接把出错行原文摘出来放进错误对象,极大提升锁文件解析失败时的可诊断性。若行列信息缺失(例如解析器在文件开头就失败),也会落入专门的分支给出针对性提示。
5.4 带行号的 JSON 解析器
值得一提的还有 util.py 中从 parsy 官方示例改编而来的 JSON 解析器json_doc。它用mark()组合子包裹每个 JSON 值,在解析值前后分别取得(line, column)位置,从而让Pipfile.lock等 JSON 锁文件中的每个依赖都能带上行号:
become( json_value, alt( quoted_str.mark().map(JSON.make), number.mark().map(JSON.make), json_object.mark().map(JSON.make), array.mark().map(JSON.make), true.mark().map(JSON.make), false.mark().map(JSON.make), null.mark().map(JSON.make), ), ) json_doc = whitespace >> json_valueJSON.make取marked[0][0] + 1(即起始行号 + 1)存入JSON.line_number,之后pipfile.py将其透传到FoundDependency.line_number。
六、类型桩:如何在不改运行时的前提下获得静态类型安全
VENDOR_README.md 提到"which can never happen with our type stubs",指的就是 cli/src/semdep/external/parsy/init.pyi。这份类型桩由 Semgrep 自行编写,核心做法是把Parser声明为泛型类Parser(Generic[T]),并在parse、parse_partial、__call__等签名中把输入流限定为str,从而在 mypy 静态检查层面杜绝"非字符串输入"路径。
有趣的是,这种泛型桩与运行时实现之间需要一个小技巧。util.py 开头的模块注释解释得很清楚:
In Python, type annotations do not do anything, but they are still expressions that get evaluated. The runtime class Parser, as implemented by parsy, takes no parameters, but our type stubs for parsy give this class a generic type variable parameter, so we can enforce staticly that parsers are combined in sensible ways. As a result, the expression Parser[int] is perfectly fine for Mypy, but causes a runtime error. Thankfully, "Parser[int]" is a perfectly acceptable type annotation for Mypy, and evaluates immediately to string, causing no runtime errors.
即:运行时Parser不接受类型参数,而桩中它是泛型,因此源码里不能写Parser[int]这种会被求值、从而触发运行时错误的表达式;解决办法是配合from __future__ import annotations使用字符串形式的注解(如Parser[str]、Parser[A]),让注解在运行时立即求值为字符串,mypy 却仍能正确解析其类型含义。这也是 cli/src/semdep/parsers/ 中几乎所有解析器类型注解都写成字符串的原因。类型桩还额外提供了line_info: Parser[Pos]、index: Parser[int]等组合子的签名,与运行时实现一一对应。
七、与上游的关系:向后不兼容的改动与合并计划
VENDOR_README.md 的最后一段交代了这一改动与上游 parsy 的关系,这是理解"为什么选择 vendoring 而非提 PR 等合入"的关键:
Matthew is planning to merge this change upstream into parsy, but it's currently backwards incompatible, so we're vendoring until that's solved, the change is merged, and parsy is released.
也就是说:
- Semgrep 团队成员 Matthew 计划把这一改动合入上游 parsy;
- 但改动目前是向后不兼容的(把公开的索引语义从整数改为三元组
Position,会破坏依赖原 API 的用户代码); - 因此在兼容性问题解决、改动合入并发布新版 parsy 之前,Semgrep 选择持续 vendoring这份定制实现。
配套的版本信息可以在 cli/src/semdep/external/parsy/version.py 中看到:__version__ = "2.0"。这一版本号既标识了 vendored 实现所对应的 parsy 基线版本,也方便后续跟踪替换为上游正式版本时的 diff。对想要替换回上游依赖的开发者而言,对比__init__.py中Position与make_index_update的实现,即可精确识别出 Semgrep 的增量行列补丁范围。
八、总结与延伸阅读
Semgrep 对 parsy 的 vendoring 是一次典型的"带着私有补丁的第三方依赖管理"实践:以最小侵入(仅替换索引表示 + 增量行列更新)换取锁文件解析的显著加速,同时用类型桩守住类型安全底线,用明确的文档记录改动语义与上游合并计划,为将来回迁上游版本铺好了路。其增量更新思路(用count("\n")/rfind("\n")做常数级行列推进)也可以作为其他文本解析器性能优化的通用参考。
若想继续深入,可以在仓库中按以下路径追踪完整链路:
- 改造文档:cli/src/semdep/external/parsy/VENDOR_README.md
- 改造实现:
Position、Result、make_index_update、string、regex、test_item、line_info等均位于 cli/src/semdep/external/parsy/init.py - 类型桩:cli/src/semdep/external/parsy/init.pyi
- 行号消费工具(
line_number、mark_line、json_doc、parse_dependency_file的错误处理):cli/src/semdep/parsers/util.py - 典型消费方:cli/src/semdep/parsers/poetry.py、cli/src/semdep/parsers/pipfile.py
- 其余基于 parsy 的解析器:cli/src/semdep/parsers/yarn.py、cli/src/semdep/parsers/go_mod.py、cli/src/semdep/parsers/gradle.py、cli/src/semdep/parsers/mix.py、cli/src/semdep/parsers/swiftpm.py、cli/src/semdep/parsers/gem.py、cli/src/semdep/parsers/pom_tree.py、cli/src/semdep/parsers/requirements.py
- SAST
- 应用安全
- 静态分析
- 开发工具
- 代码质量
【免费下载链接】semgrep
Lightweight static analysis for many languages. Find bug variants with patterns that look like source code.
相关推荐
Flutter News Toolkit数据管理:缓存策略与离线功能实现教程
Flutter News Toolkit数据管理:缓存策略与离线功能实现教程 Flutter News Toolkit是由Google和Very Good Ve
3 步跑通 Hermes Agent:自进化 AI 代理安装配置完全指南
3 步跑通 Hermes Agent:自进化 AI 代理安装配置完全指南 本文带你装好 Hermes Agent——一个能从经验里沉淀技能、越用越懂你的 AI
AI Agent人工智能AI 应用工具调用Agent 记忆交互助手RAG任务调度MCP 服务Puppeteer HTTPRequest.redirectChain() 深入解析:如何追踪与检测页面重定向链路
Puppeteer HTTPRequest.redirectChain 深入解析:如何追踪与检测页面重定向链路 导读 重定向(Redirect)是 Web 中最
浏览器控制测试网页爬虫开发工具
创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考