Skip to content

Windows 解压包含 UTF-8 路径的归档时因多字节代码页转换失败 #516

Description

@yizhinailong

环境

  • mcpp:2026.8.10.3
  • 操作系统:Microsoft Windows Server 2025,10.0.26100
  • GitHub Runner:windows-2025-vs2026
  • mcpp 选择的工具链:LLVM/Clang 20.1.7
  • MCPP_INDEX_MIRROR=GLOBAL
  • 复现仓库:mcpplibs/mcpp-index PR 留言版 | 使用mcpp工具构建的项目(库/工具/应用) #260
  • CI merge SHA:d98961cbdc2ba9de8298f5dc474370f9d5ba128c

归档信息

下载地址:

https://github.com/yhirose/cpp-httplib/archive/refs/tags/v0.53.1.tar.gz

SHA-256:

185af9587e270de9a3bfee234c6740f02e82265da33c7a41f97e02ee42f979d2

归档中包含以下非 ASCII 路径:

cpp-httplib-0.53.1/test/www/日本語Dir/
cpp-httplib-0.53.1/test/www/日本語Dir/meson.build
cpp-httplib-0.53.1/test/www/日本語Dir/日本語File.txt

复现步骤

在 mcpp-index PR #260 的代码中执行:

mcpp index update
mcpp test -p httplib-brotli

以下测试也会发生相同错误:

mcpp test -p httplib-tls
mcpp test -p httplib-zstd

实际结果

三个测试均在下载 compat.httplib 后立即失败:

Downloading compat.httplib v0.53.1
error: internal: unhandled exception: No mapping for the Unicode character exists in the target multi-byte code page.

错误发生在 feature 依赖下载和源码编译之前。

预期结果

mcpp 应当能够在 Windows 上正确解压包含 UTF-8 文件名的归档,并继续构建包。

如果确实无法解压,至少应当报告失败的归档条目,而不是抛出未处理的内部异常。

排查证据

  • 同一个包在 Linux 和 macOS 上可以通过下载和解压阶段。
  • Brotli、TLS 和 Zstd 三个测试都在相同位置失败。
  • 同一个 Windows CI job 后续能够成功下载和构建 compat.openssl、compat.zstd,因此不是可选依赖导致的问题。
  • cpp-httplib 归档中只有上述三个路径包含非 ASCII 字符。
  • 错误发生在 Downloading compat.httplib 之后、编译开始之前,符合归档解压阶段失败的特征。

综合以上信息,问题很可能是 Windows 解压逻辑将 UTF-8 归档路径通过当前 ANSI/多字节代码页进行转换,导致无法表示日文路径。

建议修复方向

Windows 归档解压过程应完整使用 Unicode 路径:

  1. 将归档中的 UTF-8 文件名明确转换为 UTF-16。
  2. 使用宽字符 Windows 文件系统 API。
  3. 避免通过当前系统 ANSI 代码页转换路径。
  4. 解压失败时输出归档 URL、目标路径和具体失败条目。

Activity

  1. self-assigned this
    on Aug 27, 2026
  2. Sunrisepeak commented on Aug 27, 2026

    @Sunrisepeak
    Member

    已在 mcpp 2026.8.27.2 发布(PR #517),四个平台的产物已在 GitHub 与 GitCode 两个
    host 上齐备,xim-pkgindex 的 latest 也已指向它。

    结论:不是解压问题,解压是对的

    抛异常的是 mcpp 自己,在 src/modgraph/scanner.cppm:238 的
    dir.filename().string()。MSVC 的 path::string() 走 WideCharToMultiByte(ACP),
    遇到当前代码页拼不出的字符就抛 std::system_error。

    你给的排查方向是合理的,但结论指错了方向,而这一点有硬证据:
    ERROR_NO_UNICODE_TRANSLATION 的前提是宽名里存在 ACP 拼不出的字符。如果 xlings
    真的按 ANSI 写坏了名字、盘上落的是 mojibake,那些字符逐个都在 CP1252 里,mcpp
    反而不会抛。mcpp 抛了,恰恰证明文件是以正确的 UTF-16 名字落盘的。

    这也是 #230 的同一处漏网:#231 加固了三个窄化站点,漏掉了同一个 walk 循环里
    早一行执行的 is_excluded_walk_dir —— 所以加固 path_matches_glob 对目录名从来无效。

    触发条件比看上去宽:include_dirs = { "*" } 的字面前缀为空,会从解压根无界递归遍历
    整棵上游源码树
    ;在 mcpplibs/mcpp-index 891b2f7 上量,130 个 recipe 有 103 个含这样的
    glob。cpp-httplib 只是第一个带非 ASCII 路径的上游。

    修好之后的行为

    test/www/日本語Dir/ 这类条目会被跳过并报告一次,构建继续:

    warning: 'C:/.../test/www' contains names this system's active code page cannot represent
      impact: those files take no part in the build
    

    这正是你提的第 4 条("至少应当报告失败的条目")。注意 chcp 改的是控制台代码页,
    对进程 ACP 无效。测试数据/文档被跳过是无害的;源文件则需要改名,或换一台代码页覆盖得了的机器。

    验证

    同一个 runner、同一个测试、同一个 ACP 上的红→绿:

    • 不含修复的 commit:FAILED ... it throws std::system_error with description "No mapping for the Unicode character exists in the target multi-byte code page."
      —— 与你报的逐字相同
    • 含修复的 commit:[ OK ] Scanner.GlobWalkSurvivesNamesTheCodePageCannotSpell (40 ms)
      —— 是 OK 不是 SKIPPED,即它确实在非 UTF-8 ACP 的 runner 上跑了

    mcpp-index PR #260 把 pin 升到 2026.8.27.2 后,httplib / httplib-tls /
    httplib-zstd 三个用例应当在 Windows 上转绿。

    后续:#518 记了两件本次范围之外的事——项目根目录含非 ACP 字符时构建会先以
    "没有源文件"失败(指错方向的诊断),以及嵌 activeCodePage=UTF-8 清单的可行性。

  3. Sunrisepeak commented on Aug 27, 2026

    @Sunrisepeak
    Member

    使用最新版本mcpp, mcpp-index 已合入

  4. added a commit that references this issue on Sep 1, 2026
    8442434
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions