feat: TS 构建管线、Spine 抓取器与 wallpapers/ 唯一真相来源

把项目从「手写 dist/」改成「wallpapers/ 是唯一真相来源,dist/ 由 pnpm build 生成」,
并补上配套的类型、门禁与抓取器。一次提交落地整条管线,因为拆开会留下不能构建的中间态。

- src/:运行时与模拟器源码(TS,strict),编译到 build/ 再拷进各分发
- tools/:build / dev / check-{syntax,paths,dist},以及抓取器与回归门禁 tools/checks/
  (.scratch/ 下那批一次性脚本移入 tools/checks/ 并入库为长期门禁)
- wallpapers/:七档壁纸的源数据 + README.md(id/音频/预设的完整规范)
- docs/adr/0005-0008:构建管线与分发拓扑、模拟器契约、自包含 sim、每骨架资源布局
- .gitignore:排除 .scratch/ 的参考资料副本(上游 spine 整仓克隆 ~1.2 GB、
  抓取侦查数据 ~680 MB)与调试转储;这些是本地调查材料,补偿会让仓库无法克隆
- 归一化 .gitignore/CONTEXT.md 行尾(工作区 CRLF、索引 LF 造成的整文件假 diff)

同时修掉三档卡住构建的未完工壁纸:
- kv45 的 meta.json 里 id 还是抓取期场景名 scene_main,经 downloader promote 正名为 kv45
- shajin / zhigengniao_juheye 的 meta.json 误用了骨架描述文件(name/spine/animations/pages)、
  且都缺 preset.template.json;现按规范重建:骨架沉到 spines/<名>/(spine-ts 按 atlas 所在
  目录解析贴图页)、补上元数据与单骨架预设,并清掉 zhigengniao 骨架里指向作者机的绝对路径
- 顺带 promote 已在 sources.yml 里的 kv46(月升之前,与兽共舞)

pnpm check 五道门全绿:10 个分发 / 7 档壁纸 / 141 处引用自包含。
This commit is contained in:
Shuery committed 2026-10-02 01:27:02 +08:00
1 parent b8eee05d78
commit 3f11426964
297 files changed
+216627 -1926

No files matched your search

+248
View File
@@ -0,0 +1,248 @@
# `tools/downloader/` — 米哈游活动页 Spine 抓取器
从 `wallpapers/sources.yml` 列出的活动页里把 **Spine 骨架 + 贴图 + 场景装配信息** 抓下来,
落到 staging(`_out/`)。**每个有内容的场景各产出一个「可直接搬走的壁纸目录」**——用户挑中的那份
整个移到 `wallpapers/<游戏id>/` 就能用(也可以让 `promote` 代劳)。
它不是构建输入:`wallpapers/` 只放**最终要发布**的资源,抓取脚本与临时下载都留在 `tools/`。
```bash
python -m tools.downloader fetch --page kv45 # 抓一页(结尾自动 verify,有硬伤就非 0 退出)
python -m tools.downloader fetch # 抓 sources.yml 里的全部页面
python -m tools.downloader fetch --interactive # 逐页问:这一页推荐哪个场景(不影响落盘范围)
python -m tools.downloader select --page nico-tea # 只做选择,写 selection.yml
python -m tools.downloader verify # 只自检产物
python -m tools.downloader fetch --offline # 只用 _cache/ + 已落盘的产物,不发任何请求(重跑验证)
python -m tools.downloader promote --page kv45/scene_ava --game hsr --id scene_ava \
--name 挥掷千星的筹码 --title '【崩坏:星穹铁道】挥掷千星的筹码' --description '[b]挥掷千星的筹码[/b]'
node tools/checks/check-downloader.mts # 门:编译 + 夹具绿 + 四处破坏必红
```
依赖:Python 3.10+ 与 **PyYAML**(`requirements.txt`)。不需要浏览器——这些页面免登录即可拿到
全量入口 bundle,骨架数据就在里面。
---
## 一、产物形状
```
tools/downloader/
├── _cache/<页面>/<脚本名> # 抓下来的 bundle 原文(文本调试缓存,可重跑零请求)
├── _out/<游戏>/<页面>/
│ ├── page.json # 页面级溯源:来源 URL / 入口脚本 / 场景清单 / 每个场景目录的摘要
│ └── <场景id>/ # ← 一个场景 = 一个可直接搬走的壁纸目录
│ ├── meta.json # id / name / title / description / game / page / scene / source / audio
│ ├── preset.template.json # 运行时配置(背景图 + sceneConfig.parts)
│ ├── scene.json # 场景侧车:part 列表 + 每个 part 的世界变换
│ ├── spines/<骨架名>/
│ │ ├── <骨架名>.json # 骨架(skeleton.images 已归一化为空串)
│ │ ├── <骨架名>.atlas # 第一行 = 贴图页名,与落盘文件名逐字一致
│ │ ├── <贴图页…> # 与 atlas **同居**(布局 A,见 ADR 0008)
│ │ └── meta.json # spine 版本 / 动画 / 皮肤 / 来源 URL
│ ├── scene/<场景图> # 几何平面(geometry + material.diffuse)贴的图
│ └── audios/ # 音源(当前恒为空:页面的 BGM 还没抓)
└── selection.yml # 机器所有的「哪一页推荐哪个场景」(**入库**)
```
场景目录名 = 场景 id 转义成合法壁纸 id(只允许 `[a-z0-9_-]`,首字符必须是字母数字):
`scene_main` / `P1` / `loading` 原样,中文场景 id(back-moon 的 `动画预览`)折成 `scene`。
**真实 id 不丢**——它在同目录的 `meta.json` 与 `scene.json` 里。
`_cache/` 与 `_out/` 在 `.gitignore` 里;`selection.yml` 入库。
### 场景内去重、跨场景不去重
* **场景内按 id 去重**:同一个场景里同一具骨架被引用多次只存一份(`ctc_rewards` 在
back-moon 的 `scene_content` 里出现 4 次、`gc_win` 在 `scene_gacha` 里 6 次,都只落一份)。
`scene.json` 里那些 part 仍然一条不少——只有**文件**去重。
* **跨场景不去重**:`scene_main/` 与 `scene_ava/` 会各存一份用到的骨架。`_out/` 是可随时
重生成的 staging,体积换"每个目录都能单独搬走"。总体积记在 `page.json` 的 `totalBytes` 里。
### `usable`:不是每个场景目录都能搬走
各页的 `scene_ui` 是一层**场景渲染目标**:它的平面用 `drawScene` modifier 把别的场景渲染成纹理
(diffuse 名就是场景 id),本身没有任何资源文件。这种场景照样落盘("页面上有几个场景"这件事在
产物里是完整的),但它的预设 `parts` 是空的:
```json
{ "usable": false, "reason": "场景里只有运行时纹理(drawScene / 贴图缓冲),没有可搬走的资源" }
```
`promote` 拒绝搬这种目录;`page.json` 的 `sceneDirs[]` 里逐条标了 `usable`。
## 二、场景选择:只决定"推荐哪个"
`fetch` **落盘全部有内容的场景**(至少有一具骨架或一块贴图平面)。纯色平面组成不了壁纸
(`sceneConfig.parts` 会是空的),完全没有 part 的场景也没有内容——这两类不落盘,原因记在
`page.json` 的 `skippedScenes` 里。
推荐场景由 `selection.yml` / `sources.yml` / 默认值算出(优先级写死,避免两份文件打架):
1. `wallpapers/sources.yml` 里这一页写了 `scene:` / `spines:` → 用它(**人的意志最高**)
2. `tools/downloader/selection.yml` 里有这一页的记录 → 用它(上次交互的结果)
3. 都没有 → **骨架最多的那个场景**(`--interactive` 时才问人)
默认规则与旧项目的判断一致(get-memory 会推荐 `P1`、back-moon 会推荐 `scene_main`)。
推荐值写进 `page.json` 的 `chosenScene` / `chosenDir`,是 `promote` 不给 `--scene` 时的默认值。
> **这份选择不再裁剪产物。** 以前它决定"只落哪几个骨架",现在产物是"一个场景一个自包含的
> 壁纸目录"——砍掉几具骨架会做出一个缺件的坏场景。要精简就在场景目录里的
> `preset.template.json` 上改(那是纯数据,构建期才读)。
> `sources.yml` 的 `spines:` 字段同理,只剩记录语义。
## 三、`scene.json` 里有什么
页面用自研引擎(three.js 系)描述场景:`sceneList` 是场景数组,每个场景是一棵树,节点有三类负载:
| kind | 含义 | 需要资源? |
| --- | --- | --- |
| `spine` | `spine:{id:"main_nike"}`,一具骨架 = 一个视觉元件 | 是(骨架 + atlas + 贴图页) |
| `image` | `geometry` + `material.uniforms.diffuse`,贴图平面 | 是(`scene/<名>.<ext>`) |
| `solid` | `defines.USE_TEXTURE == 0` 的**纯色平面**(diffuse 常写 `DEFAULT`) | 否 |
每个 part 带 `position` / `scale`(**世界变换**)与 `localPosition` / `localScale`,合成语义照抄引擎:
```
world_position = parent_position + parent_scale ⊙ local_position
world_scale = parent_scale ⊙ local_scale
```
旋转暂不参与合成(旧项目同样如此),只在 `stats.rotatedNodes` 里报出数量——不假装算了。
绘制层级 = 树序遍历序(`order`)+ 节点自身的 `renderOrder`。
带 `"runtime": true` 的 `image` part 是**没有独立文件**的平面:`drawScene` 的场景渲染目标、
`cacheContainer` 的贴图缓冲,或 diffuse 指向同场景骨架缓存。它们不进预设,也不算缺资源。
**用户口中的 "scene / geometric" 就是这里的两类节点**:`sceneList` 的场景容器与 `geometry` 平面。
Spine 官方没有这两个数据概念(官方 7 种附件类型里没有 geometric;`scene` 在 spine-webgl 里指的是
`SceneRenderer` 这个渲染器)。所以"支持场景动画"= 运行期把多具骨架 + 若干贴图平面按这套摆放合成。
## 四、`preset.template.json` 怎么生成的
* `sceneConfig.parts`:`scene.json` 里每个 `spine` / `image` part 一条;纯色平面与运行时纹理不写
(运行时画不了它们,写进去只是死配置)。路径一律写成 `./spines/<名>/…`、`./scene/<名>.<ext>`——
**正斜杠**,因为预设会被整体搬走,`str(Path)` 的反斜杠在别的机器上是错的。
* `backgroundImage`:构建期它是**必填且必须真实存在**的。自动挑法 = 场景里**面积最大**的贴图平面
(它决定 `document.body` 的底图与取景用的宽高比,挑到一块小按钮会让整幅画的比例全错);
一块贴图平面都没有时退到第一具骨架的 atlas 第一行声明的贴图页。
* 写法与校验在同一个模块 `layout.py` 里(`build_preset` / `missing_preset_paths`):
产物形状一旦改,校验规则必须同时改,分成两个文件迟早漂移。写完之后**当场**把每条路径在磁盘上
核一遍,不留到构建期才炸。
## 五、页面侧的坑(都踩过,别再踩)
- **描述表的 `src` 表达式不能用一条 `.+?` 通吃**:kv45 的 bundle 里有个编译后的模板片段写着
`{src:e.activeIcon,alt:""}})`,`\{src:(.+?),id:"…",type:"image"\}` 会从那里一路吃 **8.8 万字符**
去够后面的 `,id:"…",type:"image"}`,把夹在中间的真表项(`{src:$w,id:"loading_dt1",…}`)整个吞掉。
现在分成两种形状各匹配各的:`_DESCRIPTOR_INLINE`(`Object.values(Object.assign({…}))[0]`)与
`_DESCRIPTOR_SIMPLE`(不含 `,{}` 的表达式)。
- **资源引用有四种写法**,缺一种就会把真资源当"不存在":
1. `X = a.p + "images/x.png"`(`_ASSET_LITERAL`)
2. `{src:a(38458),id:"x"}`:模块直接导出字符串——小图被 webpack 内联成
`e.exports="data:image/png;base64,…"`(`_MODULE_STRING`)。get-memory 的 `loading_moutain_a`、
`a01_lizi`、start-ndkl 的 `loading_start_1` 都栽在这里(少了它们,平面会报"资源表里没有 URL")。
3. `{src:$w,id:"x"}`:`$w="data:image/png;base64,…"`(`_STRING_ASSIGN`,kv45 的 `loading_dt1`)。
4. `Object.values(Object.assign({"<源路径>":"data:…"}))[0]`(`_DATA_IN_ASSIGN`)。
- **骨架数据一律内联在 bundle 里**,6 个页面里 `.atlas` / `.skel` 的网络请求数是 **0**。
只有 hsr 的骨架 json 走网络(页面根目录 `<hash>.json`),而且它在 **webpack 异步 chunk**
里(`258.ccb0954b.js`)——入口 HTML 根本没列它,得读 `.u=` 的 chunk 名映射再排队抓。
这也是 hsr 的 `scene_ava` **首次离线抓不到**的原因:那两个 json 从没进过 `_cache/`,
得先联网抓一次(见第七节的离线边界)。
- **三种内联家族**都要认:A `Object.values(Object.assign({"…/spine/<N>.json":{…}}))[0]`(JS 对象字面量,
键不带引号,用 `jslit` 规范化);B 匿名模块 + 配对表 `{atlas:fn(id),json:fn(id)}`,
骨架是 `JSON.parse('…')`(**单引号** JS 字符串,要先按 JS 语义还原);C hsr 的 `spineSetting`。
- **公共路径变量名逐页不同**:`n.p` / `a.p` / `t.p`。写死 `n.p` 会让半个页面的资源表全空。
- **同一逻辑名可能有两个候选**(引擎的桌面/移动两套预载表)。判据:数组式描述表是基准集,
字典式表是移动端覆盖(引擎里是 `desktop() || base.forEach(e => override[e.id] && …)`),桌面取基准集。
- **atlas 有两种书写风格**:`size:498,330` 与 `size: 256, 256`,解析器两种都要吃。
- **`.atlas` 的页名要读 atlas 自己声明的**(`atlas.py` 解析出来的第一行/页行),不要按 `<stem>_N` 猜:
多页是 `_2.png`,也有完全不同的名字。落盘用逻辑名,atlas 第一行不用改。
- **`skeleton.images` 要归一化**:作者目录(`../images/`)搬进分发后一定指错,写空串即可
(页名相对 atlas 所在目录解析),原值记在 `meta.json` 的 `originalImages` 里。
- **跨平台路径**:别拿 `str(Path)` 当映射键(Windows 反斜杠 vs 预设里的正斜杠),一律
`Path.as_posix()`。
## 六、多版本 Spine
`meta.json` 逐具骨架记录 `skeleton.spine` 原文。事实基线(详见 `.scratch/spine-versions/REPORT.md`):
- 版本串是**编辑器版本**;`4.0-from-4.1-from-4.2-from-4.3.23` 这类 `-from-` 串表示
**数据是 4.0 格式**(降级导出),运行库只认开头那个 `major.minor`。
- 运行库**不校验**这个串;spine-ts 4.2 能读 4.0/4.1/4.2 的数据。**4.3 数据会被静默丢掉全部约束**
(4.3 把约束并进 `root.constraints`),所以 `verify` 对 `≥4.3` 的骨架直接报红。
- 目前 7 页共 170+ 具骨架全是 ≤4.2 格式,**一套 4.2 运行库就够**;真出现 4.3 数据再谈多套 UMD 共存
(官方没有 `spine-version` 属性,只能各包一层别名函数避免 `window.spine` 互相覆盖)。
- 重分发运行库要带 Spine Runtimes License Agreement(Exhibit A)+ 版权声明,不得删各文件头。
## 七、`verify` 检查什么
逐场景目录核一遍:`meta.json` / `preset.template.json` / `scene.json` 都在;`meta.json` 的
`id` 等于目录名、字符集合法、`name`/`title`/`description` 非空、`audio.choices` 指向的文件存在;
预设里每条 `./…` 路径在磁盘上找得到;侧车引用的骨架有目录与文件、atlas 声明的贴图页逐字落盘、
几何平面的图存在;纯色与运行时纹理不计入缺失;骨架版本在运行库可读范围内。
页面级还会抓"旧形状的残留"(页面根的 `scene.json` / `spine/` / `scene/`)与"不是场景目录的目录"。
**退出码非 0 = 有硬伤**。
它不评判画面对不对——那是 `tools/checks/verify-scene-player.mts` 的事(多骨架 + 贴图平面真的合成出来)。
### 离线能跑到哪一步
`fetch --offline` 只读 `_cache/`(bundle 文本)与**已经落盘的 `_out/`**。所以:
* **重跑**永远是安全的:产物存在就跳过,不发任何请求。
* **首次**抓一个"从没抓过的场景"需要联网——它的骨架 json 与贴图页都还没有本地副本。
hsr 的 `scene_ava`(`zhigengniao_juheye` / `shajin`)就是这样:先 `fetch --page kv45`(联网)
一次,之后再 `--offline` 就完全绿。缺什么会**逐条报出来**(不会甩一条 traceback 就走)。
## 八、离线预览(`_out/` 能看,`_cache/` 不能)
| 目录 | 内容 | 能否预览 |
| --- | --- | --- |
| `_cache/<页面>/` | 原始 bundle **文本**(引用仍是线上绝对地址、贴图不在里面) | ❌ 按设计就是文本调试缓存 |
| `_out/<游戏>/<页面>/<场景>/` | 真实文件树(骨架 / atlas / 贴图页 / 场景图 / 侧车 / 预设) | ✅ **可离线预览** |
```bash
node tools/preview.mts # 列出 _out 里可预览的场景
node tools/preview.mts ys/nico-tea # 页面 = 取它的推荐场景
node tools/preview.mts ys/nico-tea/scene_main # 点名场景
node tools/preview.mts hsr/kv45/scene_ava --port 8199
node tools/preview.mts ys/nico-tea --no-serve # 只组装到 tools/.cache/preview/
```
`tools/preview.mts` 把「构建出的运行时 + staging 的资产 + 由 `scene.json` 生成的预设」组装成
一个独立目录再起静态服务;**页面只读本地文件**。页面里的调试面是 `window.__sceneDebug`
(`parts` / `loaded` / `errors` / `framing` / `rotatedSkipped`),排查"少加载了一件"直接看它。
两点说明:
- **要预览"原页面"而不是我们的产物**,得走整站镜像(`.scratch/page-mirror/spec.md`,尚未实现);
`_cache/` 不能满足这个需求——它只有文本,没有资产树,也没有改写引用。
- `tools/checks/verify-scene-player.mts` 里另有一份"从 scene.json 生成预设"的代码,
**故意不与 preview 共用**:门必须独立于被验证对象,共用一份就变成自己验自己。
## 九、`promote`:搬进 `wallpapers/`
`fetch` 已经产出终态,所以 promote 退化成 **拷贝 + 校验 + 写 meta**:
```bash
python -m tools.downloader promote --page kv45/scene_ava --game hsr --id scene_ava \
--name '挥掷千星的筹码' --title '【崩坏:星穹铁道】挥掷千星的筹码' --description '[b]挥掷千星的筹码[/b]'
```
* `--page` 认三种写法:页面 id(`kv45`)、`页面/场景`(`kv45/scene_ava`)、场景目录的路径。
给页面 id 时用 `--scene` 点名,或让它取 `page.json` 的 `chosenScene`。
* `--name` / `--title` / `--description` 可选:不给就沿用场景目录 `meta.json` 里的
(那里已经是"页面名(场景id)"的形状,够用但通常要改成正式文案)。
* `--cover` 给逻辑名时覆盖 `backgroundImage`。
* **目标目录里的 `meta.json` 的 `id` 会被改写成 `--id`**——所以壁纸 id 可以跟场景目录名不同,
但目录名与 `meta.json.id` 必须一致(构建期会校验)。
* 搬之前会再核一遍预设里的每条路径在**目标目录**上存在;搬完不用手工改任何路径。
也可以完全不用 promote:**整个场景目录拷到 `wallpapers/<游戏id>/<壁纸id>/` 就完事**
(前提是目录名 = `meta.json.id`,且那个目录 `usable`)。两种做法等价,promote 只是顺手改名与校验。
## 十、还没做的
- **音源**:页面 BGM 是另一条链,`audios/` 现在是空的,`meta.json` 的 `audio.choices` 也是空的。
- **预览图**:没有抓,也没有生成。
- **`scene_ui` 这类纯渲染目标场景**:落盘但不可搬走(运行时不会 `drawScene`)。
+11
View File
@@ -0,0 +1,11 @@
"""米哈游活动页的 Spine 抓取器(Python,只依赖 stdlib + PyYAML)。
它不是构建输入:`wallpapers/` 只放最终要发布的资源,抓取脚本与临时下载都留在这里。
产物先落 `_out/<游戏>/<页面>/`,经 `--verify` 自检后再由人 promote 进 `wallpapers/`。
子命令与用法见同目录 README.md。
"""
__all__ = ["__version__"]
__version__ = "0.1.0"
+582
View File
@@ -0,0 +1,582 @@
"""命令行入口:`python -m tools.downloader <fetch|select|verify|promote>`。
设计原则(见 README.md):
* **只落 staging**(`_out/`),不碰 `wallpapers/`——promote 是后续单独一步。
* **一页 = 全部场景**:每个有内容的场景各产出一个「可直接搬走的壁纸目录」
(`_out/<游戏>/<页面>/<场景id>/`),不再只落被选中的那一个。形状与校验见 `layout.py`。
* **可重跑**:文本走 `_cache/`,产物存在就跳过;`--offline` 下不发起任何请求。
* **抓完就自检**:`fetch` 结尾自动跑 `verify`,有硬伤就非 0 退出。
"""
from __future__ import annotations
import argparse
import datetime as dt
import json
import shutil
import sys
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
import yaml
from . import layout as layout_mod
from . import promote as promote_mod
from . import scene as scene_mod
from . import selection as selection_mod
from . import verify as verify_mod
from .sites import mihoyo
HERE = Path(__file__).resolve().parent
ROOT = HERE.parents[1]
SOURCES = ROOT / "wallpapers" / "sources.yml"
CACHE = HERE / "_cache"
OUT = HERE / "_out"
SELECTION = HERE / "selection.yml"
# 旧形状(一页只落一个场景)留在页面根上的东西。新形状里它们是页面根的污染:
# `scene.json` 归到每个场景目录里、`spine/` 改名 `spines/` 并下沉到场景目录。
_LEGACY_ENTRIES = ("scene.json", "spine", "scene")
# "看着像贴图平面、其实没有独立文件"的 modifier:
# cacheContainer —— 渲染进贴图缓冲
# drawScene —— 把**另一个场景**渲染成纹理(diffuse 名就是场景 id,如各页的 scene_ui)
_RUNTIME_MODIFIERS = ("cacheContainer", "drawScene")
@dataclass
class PageEntry:
game: str
id: str
name: str
url: str
scene: str | None = None
spines: list[str] = field(default_factory=list)
class SourcesError(RuntimeError):
"""`sources.yml` 形状不对。"""
def load_sources(path: Path = SOURCES) -> list[PageEntry]:
"""读 `sources.yml`:``<游戏>: [{id, name, url, scene?, spines?}, …]``。"""
if not path.exists():
raise SourcesError(f"找不到来源清单:{path}")
raw = yaml.safe_load(path.read_text(encoding="utf-8"))
if not isinstance(raw, dict):
raise SourcesError(f"{path} 顶层必须是「游戏 → 页面列表」的映射")
entries: list[PageEntry] = []
for game, pages in raw.items():
if not isinstance(pages, list):
raise SourcesError(f"{path} 里 {game} 必须是列表")
for i, page in enumerate(pages):
where = f"{path} 的 {game}[{i}]"
if not isinstance(page, dict):
raise SourcesError(f"{where} 不是映射(YAML 里少写了一个 `- `?)")
missing = [k for k in ("id", "name", "url") if not page.get(k)]
if missing:
raise SourcesError(f"{where} 缺少字段:{', '.join(missing)}")
spines = page.get("spines") or []
if not isinstance(spines, list):
raise SourcesError(f"{where} 的 spines 必须是列表")
entries.append(
PageEntry(
game=str(game),
id=str(page["id"]),
name=str(page["name"]),
url=str(page["url"]),
scene=str(page["scene"]) if page.get("scene") else None,
spines=[str(s) for s in spines],
)
)
return entries
def _download(url: str, dest: Path, *, force: bool) -> int:
"""下载一个资源到 dest(已存在则跳过)。返回落盘字节数。"""
if dest.exists() and not force:
return 0
dest.parent.mkdir(parents=True, exist_ok=True)
data = mihoyo.http_get(url, None)
dest.write_bytes(data)
return len(data)
def _write_spine(scene_dir: Path, entry: PageEntry, data: mihoyo.PageData, name: str,
*, force: bool, offline: bool, report: list[str]) -> int:
"""落一具骨架到 `<场景目录>/spines/<名>/`:json + atlas + 贴图页 + meta.json。
页与 atlas **同居**(布局 A)——spine-ts 按 `<atlas 目录>/<页名>` 解析页,页跑到别处就读不到。
"""
asset = data.spines.get(name)
spine_dir = scene_dir / layout_mod.SPINE_DIR / name
if asset is None:
report.append(f"场景引用了骨架 {name},但页面里没有它的数据(已跳过)")
return 0
try:
payload = mihoyo.load_spine_json(data, asset, cache_dir=CACHE, offline=offline)
except mihoyo.HttpError as exc:
# 离线缺缓存 / 网络失败:报出来,别让一条 traceback 把整页的抓取带崩。
report.append(f"骨架 {name} 的数据取不到:{exc}")
return 0
original_images = (payload.get("skeleton") or {}).get("images")
normalized = dict(payload)
skeleton = dict(normalized.get("skeleton") or {})
# 页名按 atlas 所在目录解析:作者目录("../images/" 之类)搬进分发后一定指错。
skeleton["images"] = ""
normalized["skeleton"] = skeleton
spine_dir.mkdir(parents=True, exist_ok=True)
written = 0
json_file = spine_dir / f"{name}.json"
if force or not json_file.exists():
json_file.write_text(json.dumps(normalized, ensure_ascii=False, separators=(",", ":")), encoding="utf-8")
written += json_file.stat().st_size
atlas_file = spine_dir / f"{name}.atlas"
if force or not atlas_file.exists():
atlas_file.write_text(asset.atlas_text, encoding="utf-8")
written += atlas_file.stat().st_size
# 页名读 atlas 自己声明的页(`asset.pages` 就是 atlas 解析出来的),不按 `<stem>_N` 猜。
for page_name in asset.pages:
stem = page_name.rsplit(".", 1)[0]
rel = data.images.get(stem)
if rel is None:
report.append(f"骨架 {name} 的贴图页 {page_name} 在资源表里找不到 URL")
continue
try:
written += _download(data.asset_url(rel), spine_dir / page_name, force=force)
except mihoyo.HttpError as exc:
report.append(f"骨架 {name} 的贴图页 {page_name} 下载失败:{exc}")
meta = asset.as_meta(entry.url)
meta["originalImages"] = original_images
meta["fetchedAt"] = dt.datetime.now(dt.timezone.utc).isoformat(timespec="seconds")
(spine_dir / "meta.json").write_text(json.dumps(meta, ensure_ascii=False, indent=2), encoding="utf-8")
return written
def _write_images(scene_dir: Path, data: mihoyo.PageData, parts: list[Any], *,
runtime_names: set[str], force: bool, report: list[str],
notes: list[str]) -> tuple[int, set[str]]:
"""落几何平面用的场景图到 `<场景目录>/scene/`。
没有 URL 的平面**一律算运行时纹理**(不下载、不进预设,在 `scene.json` 里标 `runtime`):
资源表就是页面向网络索取资源的完整清单,名字不在表里说明页面自己也不去网上取它。
四种来源:`drawScene` 的场景渲染目标、`cacheContainer` 的贴图缓冲、diffuse 指向同场景骨架
缓存,以及运行时生成的纹理。其中只有最后一种会记一条 note——前三种是页面的正常结构。
"""
written = 0
runtime: set[str] = set()
for part in parts:
if part.kind != "image":
continue
image_id = part.id
rel = data.images.get(image_id)
if rel is None:
runtime.add(image_id)
if not any(m in part.modifiers for m in _RUNTIME_MODIFIERS) and image_id not in runtime_names:
notes.append(f"[{scene_dir.name}] 平面 {image_id} 在资源表里没有 URL,按运行时纹理处理")
continue
ext = mihoyo.image_ext(rel)
try:
written += _download(
data.asset_url(rel), scene_dir / layout_mod.SCENE_DIR / f"{image_id}{ext}", force=force
)
except mihoyo.HttpError as exc:
report.append(f"几何平面 {image_id} 下载失败:{exc}")
return written, runtime
def _scene_dir_names(scenes: list[scene_mod.Scene]) -> dict[str, str]:
"""场景 id → 目录名;同时消掉转义后可能出现的重名(大小写不敏感)。"""
used: set[str] = set()
out: dict[str, str] = {}
for scene in scenes:
base = layout_mod.scene_dir_name(scene.id)
name, index = base, 2
while name.lower() in used:
name = f"{base}_{index}"
index += 1
used.add(name.lower())
out[scene.id] = name
return out
def _scene_has_content(scene: scene_mod.Scene) -> bool:
"""这个场景有没有**内容**(骨架或贴图平面)。
只有纯色平面的场景(back-moon 的 `动画预览`)与完全没有 part 的场景(`effect_DofBlur`)
不算——它们连一张图都不需要,落出来的目录必然是空的。
注意"有内容"不等于"能搬走":各页的 `scene_ui` 全是 `drawScene` 渲染目标,一落地就是
没有 part 的目录(`usable: false`),它进 `_out` 只是为了让"页面上有几个场景"这件事
在产物里是完整的。
"""
return any(p.kind in ("spine", "image") for p in scene.parts)
def _write_scene(page_dir: Path, entry: PageEntry, data: mihoyo.PageData, scene: scene_mod.Scene, *,
dir_name: str, force: bool, offline: bool, report: list[str],
notes: list[str], fetched_at: str) -> dict[str, Any]:
"""把一个场景落成「可直接搬走的壁纸目录」。返回它的摘要(写进 page.json)。"""
scene_dir = page_dir / dir_name
(scene_dir / layout_mod.AUDIO_DIR).mkdir(parents=True, exist_ok=True)
before = len(report)
written = 0
kept_spines = scene.spine_ids # 场景内的全部骨架,按出现序去重(同一骨架被引用多次只存一份)
for name in kept_spines:
written += _write_spine(scene_dir, entry, data, name, force=force, offline=offline, report=report)
image_bytes, runtime_images = _write_images(
scene_dir, data, scene.parts, runtime_names=set(kept_spines), force=force,
report=report, notes=notes,
)
written += image_bytes
payload = scene.as_dict(page=entry.id, game=entry.game)
payload["parts"] = []
for part in scene.parts:
item = part.as_dict()
if part.kind == "image" and part.id in runtime_images:
item["runtime"] = True # 由骨架 / 别的场景渲染出来,没有独立文件
payload["parts"].append(item)
written += layout_mod.write_json(scene_dir / "scene.json", payload)
preset, preset_notes = layout_mod.build_preset(payload, scene_dir)
notes.extend(f"[{dir_name}] {note}" for note in preset_notes)
for rel in layout_mod.missing_preset_paths(preset, scene_dir):
report.append(f"[{dir_name}] preset.template.json 引用的 {rel} 不在磁盘上")
written += layout_mod.write_json(scene_dir / "preset.template.json", preset)
# 该落盘却没落进预设 = 硬伤(上面的 report 已经写了原因),别让目录看起来是好的。
planned = [p for p in scene.parts if p.kind == "spine" or (p.kind == "image" and p.id not in runtime_images)]
usable = bool(preset["sceneConfig"]["parts"])
reason = ""
if not usable:
reason = "场景里只有运行时纹理(drawScene / 贴图缓冲),没有可搬走的资源"
if not usable and planned:
report.append(f"[{dir_name}] 有 {len(planned)} 件该落盘的 part 却没进预设(见上面的缺失报告)")
meta = layout_mod.build_meta(
wallpaper_id=dir_name,
name=f"{entry.name}({scene.id})",
title=f"{entry.name}({scene.id})",
description=f"{entry.name} · 场景 {scene.id}\n来源:{entry.url}",
game=entry.game,
page=entry.id,
scene=scene.id,
source=entry.url,
fetched_at=fetched_at,
)
meta["usable"] = usable
if reason:
meta["reason"] = reason
written += layout_mod.write_json(scene_dir / "meta.json", meta)
return {
"dir": dir_name,
"scene": scene.id,
"usable": usable,
"reason": reason or None,
"spines": len(kept_spines),
"images": len({p.id for p in scene.parts if p.kind == "image"}),
"runtimeImages": sorted(runtime_images),
"solids": sum(1 for p in scene.parts if p.kind == "solid"),
"parts": len(preset["sceneConfig"]["parts"]),
"bytes": written,
"problems": len(report) - before,
}
def cmd_fetch(args: argparse.Namespace) -> int:
entries = load_sources()
if args.page:
wanted = set(args.page)
entries = [e for e in entries if e.id in wanted]
missing = wanted - {e.id for e in entries}
if missing:
print(f"来源清单里没有这些页面:{', '.join(sorted(missing))}", file=sys.stderr)
return 2
if not entries:
print("没有要抓的页面。", file=sys.stderr)
return 2
selection = selection_mod.load_selection(SELECTION)
problems: list[str] = []
total_bytes = 0
for entry in entries:
print(f"\n=== {entry.game}/{entry.id} {entry.name}")
def fail(message: str, _id: str = entry.id) -> None:
"""抓取期的硬伤当场打印——攒到最后再报会让人以为"这页没东西"。"""
problems.append(f"{_id}:{message}")
print(f" ✗ {message}")
try:
data = mihoyo.fetch_page(entry.id, entry.game, entry.url, cache_dir=CACHE, offline=args.offline)
except mihoyo.HttpError as exc:
fail(f"抓取失败 {exc}")
continue
scenes = data.scenes
if not scenes:
fail("页面里一个场景都没有")
continue
# 默认场景仍然按"骨架最多"算,但它现在只用于「哪一档是推荐的」——
# 全部有内容的场景都会落盘,选择不再裁剪产物(选择记录也只剩这个用途)。
fallback = scene_mod.pick_default_scene(scenes)
choice = selection_mod.resolve(
entry.id,
scenes,
sources_entry={"scene": entry.scene, "spines": entry.spines},
selection=selection,
default_scene=fallback.id if fallback else None,
default_spines=fallback.spine_ids if fallback else [],
interactive=args.interactive,
)
chosen = choice.scene or (fallback.id if fallback else None)
page_dir = OUT / entry.game / entry.id
page_dir.mkdir(parents=True, exist_ok=True)
for legacy in _LEGACY_ENTRIES:
stale = page_dir / legacy
if not stale.exists():
continue
# 旧形状的残留:留着会被 verify 当成形状不对的场景目录。
if stale.is_dir():
shutil.rmtree(stale, ignore_errors=True)
else:
stale.unlink()
print(f" · 清掉旧形状的 {legacy}")
names = _scene_dir_names(scenes)
wanted_scenes = [s for s in scenes if _scene_has_content(s)]
wanted_ids = {s.id for s in wanted_scenes}
skipped = [
{"scene": s.id, "reason": "只有纯色平面或完全没有 part,不需要任何资源"}
for s in scenes
if s.id not in wanted_ids
]
if not wanted_scenes:
fail("所有场景都不需要资源(没有骨架、也没有贴图平面)")
continue
print(f" 场景 {len(wanted_scenes)}/{len(scenes)} 个有内容,逐个落盘:")
fetched_at = dt.datetime.now(dt.timezone.utc).isoformat(timespec="seconds")
summaries: list[dict[str, Any]] = []
notes: list[str] = []
for scene in wanted_scenes:
before = len(problems)
summary = _write_scene(
page_dir, entry, data, scene,
dir_name=names[scene.id], force=args.force, offline=args.offline,
report=problems, notes=notes, fetched_at=fetched_at,
)
summaries.append(summary)
total_bytes += int(summary["bytes"])
for message in problems[before:]:
print(f" ✗ {message}")
mark = "✓" if summary["usable"] else "○"
print(
f" {mark} {summary['dir']}/(原 id {summary['scene']}):骨架 {summary['spines']}、"
f"贴图平面 {summary['images']}、纯色 {summary['solids']}、预设 part {summary['parts']},"
f"{int(summary['bytes']) / 1024:.1f} KB"
+ (f"(不可搬走:{summary['reason']})" if not summary["usable"] else "")
)
for note in notes:
print(f" · {note}")
# 推荐的场景要真的落了盘——推荐到一个被跳过的场景会让 promote 的默认值指向空气。
usable_dirs = {s["dir"] for s in summaries if s["usable"]}
chosen_dir = names.get(chosen, "") if chosen else ""
if chosen_dir not in usable_dirs:
chosen = next((s["scene"] for s in summaries if s["usable"]), None)
chosen_dir = names.get(chosen, "") if chosen else ""
page_payload = {
"id": entry.id,
"game": entry.game,
"name": entry.name,
"url": entry.url,
"site": data.site.id,
"entry": data.entry_url,
"bundles": sorted(data.bundles),
"scenes": [
{
"id": s.id,
"dir": names.get(s.id) if s.id in wanted_ids else None,
"spines": len(s.spine_ids),
"images": len({p.id for p in s.parts if p.kind == "image"}),
"solids": sum(1 for p in s.parts if p.kind == "solid"),
}
for s in scenes
],
"chosenScene": chosen,
"chosenDir": chosen_dir or None,
"sceneDirs": summaries,
"skippedScenes": skipped,
"totalBytes": sum(int(s["bytes"]) for s in summaries),
"warnings": data.warnings + [w for w in problems if w.startswith(entry.id)] + notes,
"fetchedAt": fetched_at,
}
layout_mod.write_json(page_dir / "page.json", page_payload)
for warning in data.warnings:
print(f" ! {warning}")
usable_count = sum(1 for s in summaries if s["usable"])
print(
f" 页面合计 {page_payload['totalBytes'] / 1024 / 1024:.2f} MB → "
f"{len(summaries)} 个场景目录(其中 {usable_count} 个可直接搬走)"
)
selection[entry.id] = choice.as_dict()
selection_mod.save_selection(SELECTION, selection)
print(f"\n合计新增 {total_bytes / 1024 / 1024:.2f} MB;选择记录 → {SELECTION.relative_to(ROOT)}")
print("\n=== verify")
found, stats = verify_mod.verify_all(OUT, pages=[e.id for e in entries])
for problem in found:
print(f" ✗ {problem}")
print(
f" 页面 {stats['pages']}、场景目录 {stats['sceneDirs']}、骨架 {stats['spines']}、"
f"几何平面 {stats['images']}、体积 {stats['bytes'] / 1024 / 1024:.2f} MB"
)
if stats["versions"]:
versions = "、".join(f"{v}×{n}" for v, n in sorted(stats["versions"].items()))
print(f" 骨架版本:{versions}")
if found:
print(f"\n{len(found)} 处硬伤,未通过。", file=sys.stderr)
return 1
return 0
def cmd_select(args: argparse.Namespace) -> int:
"""只记录「推荐哪个场景」——产物不再按选择裁剪(fetch 落全部有内容的场景)。"""
entries = load_sources()
if args.page:
entries = [e for e in entries if e.id in set(args.page)]
selection = selection_mod.load_selection(SELECTION)
for entry in entries:
try:
data = mihoyo.fetch_page(entry.id, entry.game, entry.url, cache_dir=CACHE, offline=args.offline)
except mihoyo.HttpError as exc:
print(f"{entry.id}:抓取失败 {exc}", file=sys.stderr)
return 1
fallback = scene_mod.pick_default_scene(data.scenes)
choice = selection_mod.resolve(
entry.id,
data.scenes,
sources_entry={"scene": entry.scene, "spines": entry.spines},
selection=selection,
default_scene=fallback.id if fallback else None,
default_spines=fallback.spine_ids if fallback else [],
interactive=True,
)
selection[entry.id] = choice.as_dict()
print(f"[{entry.id}] 已记录推荐场景:{choice.scene}")
selection_mod.save_selection(SELECTION, selection)
print(f"\n选择记录 → {SELECTION.relative_to(ROOT)}")
return 0
def cmd_promote(args: argparse.Namespace) -> int:
"""把 staging 的一个**场景目录**转成发布形状(`wallpapers/<游戏>/<壁纸id>/`)。"""
staged = Path(args.page)
if not staged.is_absolute():
candidate = OUT / args.page
staged = candidate if candidate.exists() else Path(args.page)
try:
result = promote_mod.promote_page(
staged,
ROOT / "wallpapers",
game=args.game,
wallpaper_id=args.id,
scene=args.scene,
name=args.name,
title=args.title,
description=args.description,
cover=args.cover,
force=args.force,
)
except (FileNotFoundError, FileExistsError) as exc:
print(str(exc), file=sys.stderr)
return 2
print(
f"promote {result.scene_dir.name} → {result.target.relative_to(ROOT)}\n"
f" 预设 part {result.parts}(骨架 {result.spines}、贴图平面 {result.images})"
f",跳过纯色平面 {result.solids_skipped},落盘 {result.bytes / 1024 / 1024:.1f} MB"
)
print(" 下一步:pnpm build(或 pnpm dev)让构建把它烘焙进 preset.js")
return 0
def cmd_verify(args: argparse.Namespace) -> int:
root = Path(args.root).resolve() if args.root else OUT
found, stats = verify_mod.verify_all(root, pages=args.page)
for problem in found:
print(f"✗ {problem}")
print(
f"页面 {stats['pages']}、场景目录 {stats['sceneDirs']}、骨架 {stats['spines']}、"
f"几何平面 {stats['images']}、体积 {stats['bytes'] / 1024 / 1024:.2f} MB"
)
if stats["versions"]:
versions = "、".join(f"{v}×{n}" for v, n in sorted(stats["versions"].items()))
print(f"骨架版本:{versions}")
return 1 if found else 0
def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(prog="python -m tools.downloader", description="米哈游活动页 Spine 抓取器")
sub = parser.add_subparsers(dest="command", required=True)
def common(p: argparse.ArgumentParser) -> None:
p.add_argument("--page", action="append", help="只处理指定页面 id(可重复)")
fetch = sub.add_parser("fetch", help="抓取并落 staging:每个有内容的场景一个壁纸目录(结尾自动 verify)")
common(fetch)
fetch.add_argument("--interactive", action="store_true", help="逐页交互式选「推荐哪个场景」")
fetch.add_argument("--offline", action="store_true", help="只用 _cache/,不发起任何请求")
fetch.add_argument("--force", action="store_true", help="重下已存在的产物")
fetch.set_defaults(func=cmd_fetch)
pick = sub.add_parser("select", help="只做交互式选择(推荐场景),写入 selection.yml")
common(pick)
pick.add_argument("--offline", action="store_true", help="只用 _cache/,不发起任何请求")
pick.set_defaults(func=cmd_select)
promote = sub.add_parser("promote", help="把一个场景目录转成 wallpapers/<游戏>/<壁纸id>/")
promote.add_argument("--page", required=True,
help="页面 id(如 kv45)、「页面/场景」(如 kv45/scene_ava)或场景目录的路径")
promote.add_argument("--game", required=True, help="游戏 id(wallpapers/<游戏>/)")
promote.add_argument("--id", required=True, help="壁纸 id(全局唯一,与 WE combo 的 value 一致)")
promote.add_argument("--scene", help="页面目录下要 promote 的场景(默认取 page.json 的 chosenScene)")
promote.add_argument("--name", help="显示名(默认沿用场景目录 meta.json 的)")
promote.add_argument("--title", help="创意工坊标题(默认沿用场景目录 meta.json 的)")
promote.add_argument("--description", help="WE 的 BBCode 文案(默认沿用场景目录 meta.json 的)")
promote.add_argument("--cover", help="封面图逻辑名(scene/<名>.*),写进 backgroundImage")
promote.add_argument("--force", action="store_true", help="目标目录已存在时覆盖")
promote.set_defaults(func=cmd_promote)
check = sub.add_parser("verify", help="自检 _out/ 的引用闭包与文件齐全")
common(check)
check.add_argument("--root", help="检查别的产物目录(默认 tools/downloader/_out)")
check.set_defaults(func=cmd_verify)
return parser
def main(argv: list[str] | None = None) -> int:
args = build_parser().parse_args(argv)
try:
return int(args.func(args))
except SourcesError as exc:
print(f"sources.yml 有问题:{exc}", file=sys.stderr)
return 2
if __name__ == "__main__":
raise SystemExit(main())
+76
View File
@@ -0,0 +1,76 @@
"""Spine atlas 文本的解析与页名规范化。
.atlas 是文本格式:**第一行是贴图页文件名**,接着 `size:` / `filter:` / `format:` / `scale:`
等页面头,之后才是各区域的 `bounds` / `offsets` / `rotate`。换页就是再来一行页文件名。
两种书写风格都要吃(同一批页面里都存在):
size:498,330 ← 紧凑式(nico-tea / zhidong-wonder / kv45)
size: 256, 256 ← 带空格式(get-memory)
抓取期我们只关心两件事:**这具骨架需要哪些贴图页**,以及**页名能不能原样落盘**
(运行时按第一行去请求贴图页,所以文件名必须与第一行逐字一致)。
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
__all__ = ["AtlasInfo", "parse", "page_names", "retarget_pages"]
_PAGE_EXT = re.compile(r"\.(png|webp|jpg|jpeg)$", re.IGNORECASE)
_PAGE_HEAD = re.compile(r"^(size|format|filter|repeat|pma|scale)\s*:", re.IGNORECASE)
_REGION_ATTR = re.compile(
r"^(bounds|offsets|rotate|xy|orig|index|split|pad|width|height)\s*:", re.IGNORECASE
)
@dataclass
class AtlasInfo:
"""一具骨架的 atlas 摘要。"""
pages: list[str] = field(default_factory=list)
regions: list[str] = field(default_factory=list)
region_pages: list[str | None] = field(default_factory=list)
@property
def multi_page(self) -> bool:
return len(self.pages) > 1
def parse(text: str) -> AtlasInfo:
"""解析 atlas 文本。无法识别的行按区域名处理(atlas 里非属性行就是区域名)。"""
info = AtlasInfo()
current: str | None = None
for raw in text.splitlines():
line = raw.strip()
if not line:
current = None
continue
if ":" not in line and _PAGE_EXT.search(line):
current = line
info.pages.append(line)
continue
if _PAGE_HEAD.match(line) or _REGION_ATTR.match(line):
continue
info.regions.append(line)
info.region_pages.append(current)
return info
def page_names(text: str) -> list[str]:
"""只要贴图页名(顺序即 atlas 里声明的顺序)。"""
return parse(text).pages
def retarget_pages(text: str, mapping: dict[str, str]) -> str:
"""把 atlas 里的页名按 mapping 改写(只改整行的页名行,不碰区域名)。"""
out: list[str] = []
for raw in text.splitlines():
line = raw.strip()
if line and ":" not in line and _PAGE_EXT.search(line) and line in mapping:
out.append(mapping[line])
else:
out.append(raw)
tail = "\n" if text.endswith("\n") else ""
return "\n".join(out) + tail
+217
View File
@@ -0,0 +1,217 @@
"""JS 字面量 → Python 对象。
米哈游活动页把骨架数据与场景树以 **JS 对象字面量** 内联在 bundle 里,不是合法 JSON:
{skeleton:{hash:"KR8Ibf8DXTI",spine:"4.2.43",x:-110.93},bones:[{name:"root"}],scaleX:.6487}
键不带引号、小数可以省略前导 0、布尔写作 ``!0`` / ``!1``、末尾可以有逗号。
这里只做**词法层**的规范化(把字面量改写成 JSON 文本),再交给 ``json.loads``——
不 ``eval``、不执行页面代码。
唯一被容忍的"语义"是未知裸标识符(例如 ``undefined``):默认替换为 ``null`` 并计数,
``strict=True`` 时抛错。骨架数据里出现别的裸标识符说明提取边界错了,值得知道。
"""
from __future__ import annotations
import json
import re
from typing import Any
__all__ = ["loads", "match_literal", "first_object_value", "unescape", "JsLitError"]
# 按优先级排列的词法单元。字符串/数字/``!0`` 必须排在 ``other`` 之前。
_TOKEN = re.compile(
r"""
(?P<ws>\s+)
| (?P<dquote>"(?:[^"\\]|\\.)*")
| (?P<squote>'(?:[^'\\]|\\.)*')
| (?P<template>`(?:[^`\\]|\\.)*`)
| (?P<negnot>!0|!1)
| (?P<num>-?(?:\d+\.\d*|\.\d+|\d+)(?:[eE][+-]?\d+)?)
| (?P<ident>[A-Za-z_$][\w$]*)
| (?P<other>.)
""",
re.VERBOSE | re.DOTALL,
)
_LITERAL_IDENTS = {"true": "true", "false": "false", "null": "null"}
_NULL_IDENTS = {"undefined", "NaN", "Infinity", "void"}
_ESCAPES = {"n": "\n", "t": "\t", "r": "\r", "b": "\b", "f": "\f", "v": "\v", "0": "\0"}
class JsLitError(ValueError):
"""字面量无法规范化成 JSON。"""
def unescape(body: str) -> str:
"""还原 JS 字符串体里的转义(``\\n`` / ``\\uXXXX`` / ``\\'`` / ``\\\\`` …)。
Family B 的骨架是 ``e.exports=JSON.parse('…')``:外层是**单引号** JS 字符串,
必须先按 JS 语义还原,再交给 ``json.loads``。
"""
return _unescape(body, "'")
def _unescape(body: str, quote: str) -> str:
"""还原 JS 字符串体(不含两端引号)里的转义。"""
out: list[str] = []
i = 0
while i < len(body):
ch = body[i]
if ch != "\\":
out.append(ch)
i += 1
continue
nxt = body[i + 1] if i + 1 < len(body) else ""
if nxt == "u" and i + 6 <= len(body):
out.append(chr(int(body[i + 2 : i + 6], 16)))
i += 6
elif nxt == "x" and i + 4 <= len(body):
out.append(chr(int(body[i + 2 : i + 4], 16)))
i += 4
elif nxt == "\n": # 行延续
i += 2
else:
out.append(_ESCAPES.get(nxt, nxt))
i += 2
return "".join(out)
def _num_key(value: str) -> str:
"""把数字字面量还原成 JS 当键时用的字符串(``0`` → ``"0"``、``.5`` → ``"0.5"``)。"""
try:
as_float = float(value)
except ValueError:
return value
if as_float.is_integer():
return str(int(as_float))
return repr(as_float)
def _normalize(text: str, *, strict: bool) -> tuple[str, int]:
out: list[str] = []
unknown = 0
pos = 0
for m in _TOKEN.finditer(text):
kind = m.lastgroup
raw = m.group()
if kind == "ws" or kind == "other":
out.append(raw)
pos = m.end()
continue
if kind in ("dquote", "squote", "template"):
body = raw[1:-1]
quote = raw[0]
out.append(json.dumps(_unescape(body, quote), ensure_ascii=False))
elif kind == "negnot":
out.append("true" if raw == "!0" else "false")
elif kind == "num":
value = raw
if value.startswith("-."):
value = "-0" + value[1:]
elif value.startswith("."):
value = "0" + value
if value.endswith("."):
value += "0"
# 数字也能当键(动画名就叫 "0" 的骨架真实存在:`animations:{0:{…}}`)。
# JS 会把数字字面量转成字符串当键,所以这里要按同一语义还原。
if re.match(r"\s*:", text[m.end() :]):
out.append(json.dumps(_num_key(value), ensure_ascii=False))
else:
out.append(value)
elif kind == "ident":
# 后面(跳过空白)跟冒号的标识符是键,否则是值。
tail = text[m.end() :]
is_key = bool(re.match(r"\s*:", tail))
if is_key:
out.append(json.dumps(raw, ensure_ascii=False))
elif raw in _LITERAL_IDENTS:
out.append(_LITERAL_IDENTS[raw])
elif raw in _NULL_IDENTS:
out.append("null")
unknown += 1
if strict:
raise JsLitError(f"未知标识符 {raw!r} @ {m.start()}")
else:
out.append("null")
unknown += 1
if strict:
raise JsLitError(f"未知标识符 {raw!r} @ {m.start()}")
pos = m.end()
if pos < len(text):
out.append(text[pos:])
normalized = "".join(out)
# 尾逗号:`,}` / `,]`
normalized = re.sub(r",(\s*[}\]])", r"\1", normalized)
return normalized, unknown
def loads(text: str, *, strict: bool = False) -> Any:
"""把 JS 字面量文本解析成 Python 对象。"""
normalized, _ = _normalize(text, strict=strict)
try:
return json.loads(normalized)
except json.JSONDecodeError as exc:
head = text[max(0, exc.pos - 120) : exc.pos + 120].replace("\n", " ")
raise JsLitError(f"字面量规范化后仍不是 JSON:{exc.msg} @ {exc.pos}\n…{head}…") from exc
def match_literal(text: str, start: int) -> int:
"""返回 ``text[start]`` 处那个配平字面量的**闭括号下标**;找不到返回 -1。
与 JS 侧同名的辅助函数一一对应:跳过字符串/模板串/注释,按开闭括号配平。
它是所有提取器的地基——正则数不清嵌套括号,只有这个能。
"""
if start >= len(text):
return -1
open_ch = text[start]
close_ch = {"{": "}", "[": "]", "(": ")"}.get(open_ch)
if close_ch is None:
return -1
depth = 0
i = start
while i < len(text):
ch = text[i]
if ch in "\"'`":
quote = ch
i += 1
while i < len(text):
if text[i] == "\\":
i += 2
elif text[i] == quote:
break
else:
i += 1
i += 1
continue
if ch == "/" and i + 1 < len(text) and text[i + 1] == "/":
while i < len(text) and text[i] != "\n":
i += 1
continue
if ch == "/" and i + 1 < len(text) and text[i + 1] == "*":
i += 2
while i + 1 < len(text) and not (text[i] == "*" and text[i + 1] == "/"):
i += 1
i += 2
continue
if ch == open_ch:
depth += 1
elif ch == close_ch:
depth -= 1
if depth == 0:
return i
i += 1
return -1
def first_object_value(text: str, brace_index: int) -> Any:
"""``Object.values(Object.assign({k: v}))[0]`` 的取值语义:解析 ``{...}`` 并返回第一个值。"""
end = match_literal(text, brace_index)
if end < 0:
raise JsLitError(f"未配平的对象字面量 @ {brace_index}")
obj = loads(text[brace_index : end + 1])
if not isinstance(obj, dict) or not obj:
raise JsLitError(f"期望非空对象字面量 @ {brace_index}")
return next(iter(obj.values()))
+289
View File
@@ -0,0 +1,289 @@
"""把一个场景组装成「可直接搬走的壁纸目录」。
抓取期产出的**终止形状**就是 `wallpapers/<游戏>/<壁纸id>/` 的同形拷贝
(布局 A,见 `docs/adr/0008-asset-layout-per-skeleton.md`):
<页面>/
├── page.json ← 页面级溯源
└── <场景id>/ ← 一个场景 = 一个壁纸目录
├── meta.json
├── preset.template.json
├── scene.json
├── spines/<骨架名>/ ← <名>.json + <名>.atlas + 贴图页**同居**
├── scene/<场景图>
└── audios/
目录名固定 `spines` / `scene` / `audios`——`tools/lib/generate.ts` 的 `copyWallpaperAssets`
按这三个名字搬文件,改名要两边一起改。
这个模块只做两件事:**写**(`build_preset` / `build_meta`)与**查**(`missing_preset_paths`)。
写与查放在一起是刻意的:产物形状一旦改,校验规则必须同时改,分成两个文件迟早漂移。
"""
from __future__ import annotations
import json
import re
from pathlib import Path
from typing import Any
__all__ = [
"AUDIO_DIR",
"IMAGE_EXT",
"SCENE_DIR",
"SPINE_DIR",
"build_meta",
"build_preset",
"image_file",
"missing_preset_paths",
"preset_paths",
"scene_dir_name",
"spine_page_file",
"write_json",
]
SPINE_DIR = "spines"
SCENE_DIR = "scene"
AUDIO_DIR = "audios"
IMAGE_EXT = (".png", ".webp", ".jpg", ".jpeg", ".gif")
# 壁纸 id 的字符集(与 tools/lib/vault.ts 的校验逐字一致):它会进 WE 的 combo value。
_UNSAFE = re.compile(r"[^A-Za-z0-9_-]+")
def scene_dir_name(scene_id: str) -> str:
"""场景 id → 目录名。
页面里的场景 id 多数本来就合法(`scene_main` / `P1` / `loading`),但也有中文的
(back-moon 的 `动画预览`)。目录名一旦成为壁纸 id 就必须是 `^[a-z0-9][a-z0-9_-]*$`,
所以这里统一转义:非法字符折成一个 `_`,首字符不是字母数字时补前缀。
**真实 id 不会被丢掉**——它写在同目录的 `meta.json` / `scene.json` 里。
"""
name = _UNSAFE.sub("_", str(scene_id).strip()).strip("_-")
if not name:
return "scene"
if not re.match(r"[A-Za-z0-9]", name):
return f"scene_{name}"
return name
def image_file(scene_dir: Path, name: str) -> Path | None:
"""场景图落盘名:`scene/<逻辑名>.<ext>`(扩展名由资源路径决定,见 mihoyo.image_ext)。"""
for ext in IMAGE_EXT:
candidate = scene_dir / SCENE_DIR / f"{name}{ext}"
if candidate.is_file():
return candidate
return None
def spine_page_file(scene_dir: Path, spine_id: str) -> Path | None:
"""一具骨架的第一张贴图页。
页名**读 atlas 自己声明的第一行**,不按 `<stem>_N` 猜——去 hash 之后页名未必是
`<骨架名>.png`(多页是 `_2.png`,也有完全不同的名字)。这条路径只用于"整个场景一张
贴图平面都没有"时的背景兜底(`wallpapers/README.md` 允许背景指向骨架贴图页)。
"""
atlas = scene_dir / SPINE_DIR / spine_id / f"{spine_id}.atlas"
if not atlas.is_file():
return None
for line in atlas.read_text(encoding="utf-8").splitlines():
page = line.strip().strip('"')
if page.lower().endswith(IMAGE_EXT):
candidate = atlas.parent / Path(page).name
if candidate.is_file():
return candidate
return None
def _rel(scene_dir: Path, path: Path) -> str:
"""磁盘路径 → 预设里的 `./…` 相对路径(**一律正斜杠**,见"跨平台路径"那条坑)。"""
return "./" + path.relative_to(scene_dir).as_posix()
def _part_common(part: dict[str, Any]) -> dict[str, Any]:
"""part 的公共字段(世界变换 + 绘制层级 + 页面指定的动画/皮肤/时间缩放)。"""
common: dict[str, Any] = {
"kind": part.get("kind"),
"id": part["id"],
"order": part.get("order", 0),
"position": part.get("position", [0, 0, 0]),
"scale": part.get("scale", [1, 1, 1]),
}
if part.get("renderOrder"):
common["renderOrder"] = part["renderOrder"]
if part.get("geometrySize"):
common["width"], common["height"] = part["geometrySize"]
if part.get("geometryCenter"):
common["center"] = part["geometryCenter"]
if part.get("rotation") and any(abs(float(v)) > 1e-9 for v in part["rotation"]):
common["rotation"] = part["rotation"]
# 页面指定的动画 / 皮肤:不抄就会去播骨架的第一个动画(常是入场动画 in,姿态不同)。
if part.get("animation"):
common["animation"] = part["animation"]
if part.get("skin"):
common["skin"] = part["skin"]
if part.get("timeScale") is not None:
common["timeScale"] = part["timeScale"]
return common
def build_preset(
scene: dict[str, Any],
scene_dir: Path,
*,
cover: str | None = None,
) -> tuple[dict[str, Any], list[str]]:
"""由 `scene.json` 的内容 + 磁盘上的场景目录,生成 `preset.template.json`。
返回(预设, 说明列表)。说明是"这个场景里没能进预设的东西"——纯色平面、缺文件的 part,
它们不是错误(`verify` 会独立判红),但用户该知道少了几件。
`cover` 给的是**逻辑名**(`scene/<名>.<ext>`);不给就自动挑一张背景图,
因为构建期 `backgroundImage` 是必填且必须真实存在(见 `tools/lib/vault.ts`)。
"""
notes: list[str] = []
parts: list[dict[str, Any]] = []
solids = 0
for part in scene.get("parts") or []:
kind = part.get("kind")
if kind == "solid":
# 纯色平面运行时这一轮画不了(没有贴图),写进预设只是死配置。
solids += 1
continue
if kind not in ("spine", "image"):
notes.append(f"未知的 part 类型 {kind!r}(id={part.get('id')}),已跳过")
continue
common = _part_common(part)
if kind == "spine":
spine_id = str(part["id"])
spine_dir = scene_dir / SPINE_DIR / spine_id
if not (spine_dir / f"{spine_id}.json").is_file():
notes.append(f"骨架 {spine_id} 没有落到 spines/{spine_id}/,未写进预设")
continue
common["jsonUrl"] = f"./{SPINE_DIR}/{spine_id}/{spine_id}.json"
common["atlasUrl"] = f"./{SPINE_DIR}/{spine_id}/{spine_id}.atlas"
else:
if part.get("runtime"):
# 运行时缓冲(cacheContainer / diffuse 指向同场景骨架缓存):由骨架渲染进贴图
# 缓冲,没有独立文件,不该进预设也不该报缺资源(`scene.json` 里标了 runtime)。
continue
source = image_file(scene_dir, str(part["id"]))
if source is None:
notes.append(f"贴图平面 {part['id']} 没有独立文件,未写进预设")
continue
common["image"] = _rel(scene_dir, source)
parts.append(common)
if solids:
notes.append(f"{solids} 块纯色平面未写进预设(运行时没有贴图可画)")
background: str | None = None
if cover:
source = image_file(scene_dir, cover)
if source is None:
notes.append(f"指定的背景图 {cover} 不在 scene/ 里,改为自动挑选")
else:
background = _rel(scene_dir, source)
if background is None:
background = _pick_background(scene_dir, parts)
scene_config: dict[str, Any] = {"ui": scene.get("ui"), "parts": parts}
camera_node = scene.get("camera") or {}
camera = camera_node.get("camera") or {}
if camera.get("type") is not None:
# 相机必须带进预设:透视场景(type 1)忽略 fov 与 z 就会画成一块糊满屏的贴图。
scene_config["camera"] = {
"type": camera.get("type"),
"fov": camera.get("fov"),
"position": camera_node.get("position"),
}
return {"backgroundImage": background or "", "sceneConfig": scene_config}, notes
def _pick_background(scene_dir: Path, parts: list[dict[str, Any]]) -> str | None:
"""自动挑背景图:贴图平面里**面积最大**的那块(画布比例的来源),没有就退到骨架贴图页。
为什么按面积:`backgroundImage` 在运行时决定 `document.body` 的底图与取景用的宽高比
(`src/runtime/index.ts` 的 `measureImageAspect`)。挑到一块小按钮会让整幅画的比例全错,
而背景/天空/远景恰好总是场景里最大的那块平面(实测:nico-tea 的 `main_sky_jpg` 2500×1064)。
"""
images = [p for p in parts if p.get("image")]
if images:
best = max(
images,
key=lambda p: (float(p.get("width") or 0) * float(p.get("height") or 0), -int(p.get("order") or 0)),
)
return str(best["image"])
for part in parts:
if not part.get("jsonUrl"):
continue
source = spine_page_file(scene_dir, str(part["id"]))
if source is not None:
return _rel(scene_dir, source)
return None
def preset_paths(preset: dict[str, Any]) -> list[str]:
"""预设里所有必须真实存在的 `./…` 路径。"""
out: list[str] = []
background = preset.get("backgroundImage")
if isinstance(background, str) and background:
out.append(background)
scene_config = preset.get("sceneConfig") or {}
for part in scene_config.get("parts") or []:
for key in ("image", "jsonUrl", "atlasUrl"):
value = part.get(key)
if isinstance(value, str) and value:
out.append(value)
return out
def missing_preset_paths(preset: dict[str, Any], scene_dir: Path) -> list[str]:
"""预设里指向磁盘上不存在的文件的路径(写出后立刻自检,不留到构建期才炸)。"""
missing: list[str] = []
for rel in preset_paths(preset):
# 映射键一律用**正斜杠**:Windows 的 str(Path) 是反斜杠,写进预设就搬到别的机器上读不到。
if not (scene_dir / rel.replace("\\", "/").removeprefix("./")).is_file():
missing.append(rel)
return missing
def build_meta(
*,
wallpaper_id: str,
name: str,
title: str,
description: str,
game: str,
page: str,
scene: str,
source: str,
fetched_at: str | None = None,
) -> dict[str, Any]:
"""场景目录的 `meta.json`:既有字段语义(id/name/title/description/audio)原样保留,
另加溯源字段(game/page/scene/source),方便搬进 `wallpapers/` 后回查来源。"""
meta: dict[str, Any] = {
"id": wallpaper_id,
"name": name,
"title": title,
"description": description,
"game": game,
"page": page,
"scene": scene,
"source": source,
# 音源清单为空:页面的 BGM 还没抓(构建会据此省掉 bgm 属性)。
"audio": {"choices": []},
}
if fetched_at:
meta["fetchedAt"] = fetched_at
return meta
def write_json(path: Path, data: Any, *, indent: int = 2) -> int:
"""写一份 JSON(UTF-8、不转义中文),返回字节数。"""
path.parent.mkdir(parents=True, exist_ok=True)
text = json.dumps(data, ensure_ascii=False, indent=indent) + "\n"
path.write_text(text, encoding="utf-8")
return len(text.encode("utf-8"))
+230
View File
@@ -0,0 +1,230 @@
"""把一个 staging 的**场景目录** promote 成 `wallpapers/<游戏>/<壁纸id>/`。
`fetch` 现在产出的就是终态形状(一个场景 = 一个可直接搬走的壁纸目录),所以这一步退化成
**拷贝 + 校验 + 写 meta**:把 `spines/` / `scene/` / `audios/` 三个目录与 `preset.template.json`
搬过去,把 `meta.json` 的 id / 文案换成发布值,再把预设里每条 `./…` 路径在目标目录上核一遍。
刻意不做的事:不抓音源(页面的 BGM 是另一条链)、不生成预览图、不动已存在的目录(除非 --force)。
"""
from __future__ import annotations
import json
import shutil
from dataclasses import dataclass
from pathlib import Path
from typing import Any
from . import layout as layout_mod
__all__ = ["promote_page", "resolve_scene_dir", "PromoteResult"]
@dataclass
class PromoteResult:
target: Path
scene_dir: Path
parts: int
spines: int
images: int
solids_skipped: int
bytes: int
def _read_json(path: Path) -> dict[str, Any]:
return json.loads(path.read_text(encoding="utf-8"))
def _is_scene_dir(path: Path) -> bool:
"""场景目录的判据:`preset.template.json` + `scene.json` 都在(fetch 的产物形状)。"""
return (path / "preset.template.json").is_file() and (path / "scene.json").is_file()
def scene_dirs(staged: Path) -> list[Path]:
"""页面目录下的全部场景目录(按名字排序,保证可复现)。"""
if not staged.is_dir():
return []
return sorted(p for p in staged.iterdir() if p.is_dir() and _is_scene_dir(p))
def resolve_scene_dir(staged: Path, scene: str | None = None) -> Path:
"""把「场景目录」或「页面目录 + 场景 id」统一解析成**一个**场景目录。
认三种输入(与 CLI 的 `--page` 对应):
1. 场景目录本身(`_out/<游戏>/<页面>/<场景id>`)
2. `.scratch` 式的「页面/场景」路径——由调用方拼好后传进来
3. 页面目录(`_out/<游戏>/<页面>`):用 `--scene` 指定;没指定就取 `page.json` 的
`chosenScene`;再没有就要求页面下只有一个场景(多个时报错,不猜)。
"""
if _is_scene_dir(staged):
return staged
if not staged.is_dir():
raise FileNotFoundError(f"staging 里没有这个东西:{staged}")
candidates = scene_dirs(staged)
if not candidates:
raise FileNotFoundError(f"{staged} 下没有场景目录(先跑:python -m tools.downloader fetch)")
def matches(name: str) -> Path | None:
for candidate in candidates:
if candidate.name == name:
return candidate
try:
meta = _read_json(candidate / "meta.json")
except (OSError, json.JSONDecodeError):
continue
if meta.get("scene") == name or meta.get("id") == name:
return candidate
return None
if scene:
hit = matches(scene)
if hit is None:
names = "、".join(p.name for p in candidates)
raise FileNotFoundError(f"{staged} 下没有场景 {scene}(有:{names})")
return hit
chosen = None
page_file = staged / "page.json"
if page_file.is_file():
try:
chosen = _read_json(page_file).get("chosenDir") or _read_json(page_file).get("chosenScene")
except (OSError, json.JSONDecodeError):
chosen = None
if chosen:
hit = matches(str(chosen))
if hit is not None:
return hit
if len(candidates) == 1:
return candidates[0]
names = "、".join(p.name for p in candidates)
raise FileNotFoundError(
f"{staged} 下有 {len(candidates)} 个场景目录,要用 --scene 指定一个(有:{names})"
)
def promote_page(
staged: Path,
wallpapers_root: Path,
*,
game: str,
wallpaper_id: str,
scene: str | None = None,
name: str | None = None,
title: str | None = None,
description: str | None = None,
cover: str | None = None,
force: bool = False,
) -> PromoteResult:
"""执行一次 promote;返回统计。"""
scene_dir = resolve_scene_dir(staged, scene)
if not (scene_dir / "scene.json").is_file():
raise FileNotFoundError(f"场景目录里没有 scene.json:{scene_dir}")
scene_payload = _read_json(scene_dir / "scene.json")
src_meta: dict[str, Any] = {}
meta_file = scene_dir / "meta.json"
if meta_file.is_file():
try:
src_meta = _read_json(meta_file)
except json.JSONDecodeError as exc:
raise FileNotFoundError(f"{meta_file} 不是合法 JSON:{exc}") from exc
target = wallpapers_root / game / wallpaper_id
if target.exists() and not force:
raise FileExistsError(f"目标已存在(要覆盖就加 --force):{target}")
if target.exists():
shutil.rmtree(target)
# 布局 A:一具骨架一组(json + atlas + 贴图页同居),场景图单独一层。
# 页必须与 atlas 同目录——spine-ts 按 `<atlas 目录>/<页名>` 解析(见 README)。
written = 0
for segment in (layout_mod.SPINE_DIR, layout_mod.SCENE_DIR, layout_mod.AUDIO_DIR):
source = scene_dir / segment
if source.is_dir():
written += _copy_tree(source, target / segment)
# 预设优先沿用场景目录里已生成的那份(fetch 已产出终态);缺失就按 scene.json 现算。
preset_file = scene_dir / "preset.template.json"
if preset_file.is_file():
preset = _read_json(preset_file)
else:
preset, _ = layout_mod.build_preset(scene_payload, scene_dir)
if cover:
source = layout_mod.image_file(scene_dir, cover)
if source is None:
raise FileNotFoundError(f"封面图不在场景目录里:scene/{cover}.*")
dest = target / layout_mod.SCENE_DIR / source.name
if not dest.exists():
written += _copy(source, dest)
preset["backgroundImage"] = f"./{layout_mod.SCENE_DIR}/{source.name}"
if not ((preset.get("sceneConfig") or {}).get("parts") or []):
if target.exists():
shutil.rmtree(target, ignore_errors=True)
raise FileNotFoundError(
f"{scene_dir} 里一件可搬走的 part 都没有(平面全是运行时渲染目标),换一个场景目录"
)
missing = layout_mod.missing_preset_paths(preset, target)
if missing:
# 别把半个目标目录留在 wallpapers/ 里——它看起来像一档发布壁纸,实际缺件。
shutil.rmtree(target, ignore_errors=True)
raise FileNotFoundError(
"promote 后预设里有路径找不到对应文件(已拷贝的资产不自洽):\n - " + "\n - ".join(missing)
)
written += layout_mod.write_json(target / "preset.template.json", preset)
final_meta = dict(src_meta)
# 这两个是 staging 专用字段("这个目录能不能直接搬走"),发布形状里没有它们的位置。
final_meta.pop("usable", None)
final_meta.pop("reason", None)
final_meta["id"] = wallpaper_id
final_meta.setdefault("name", wallpaper_id)
final_meta.setdefault("title", str(final_meta["name"]))
final_meta.setdefault("description", str(final_meta["name"]))
if name:
final_meta["name"] = name
if title:
final_meta["title"] = title
if description:
final_meta["description"] = description
# 音源清单为空:这一页的 BGM 还没抓(构建会据此省掉 bgm 属性)。
final_meta.setdefault("audio", {"choices": []})
final_meta.setdefault("game", game)
final_meta.setdefault("page", scene_dir.parent.name)
written += layout_mod.write_json(target / "meta.json", final_meta)
game_dir = wallpapers_root / game
if not (game_dir / "meta.json").exists():
(game_dir / "meta.json").write_text(
json.dumps({"id": game, "name": game, "audios": []}, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
parts = (preset.get("sceneConfig") or {}).get("parts") or []
solids = sum(1 for p in scene_payload.get("parts") or [] if p.get("kind") == "solid")
return PromoteResult(
target=target,
scene_dir=scene_dir,
parts=len(parts),
spines=sum(1 for p in parts if p.get("kind") == "spine"),
images=sum(1 for p in parts if p.get("kind") == "image"),
solids_skipped=solids,
bytes=written,
)
def _copy(src: Path, dest: Path) -> int:
dest.parent.mkdir(parents=True, exist_ok=True)
shutil.copy2(src, dest)
return dest.stat().st_size
def _copy_tree(src: Path, dest: Path) -> int:
total = 0
for item in sorted(src.rglob("*")):
if item.is_file():
total += _copy(item, dest / item.relative_to(src))
return total
+1
View File
@@ -0,0 +1 @@
PyYAML>=6
+372
View File
@@ -0,0 +1,372 @@
"""页面场景树(``sceneList``)的解析与"part"提取。
米哈游活动页用自研引擎(three.js 系)描述场景:顶层是一个 ``sceneList`` 数组,每个元素是一个
**场景**(``scene_main`` / ``scene_loading`` / ``scene_game`` …),场景里是一棵树。节点有两类
我们关心的负载:
* ``spine:{id:"main_nike"}`` —— 一具骨架(一具骨架 = 一个视觉元件:人物、云、箱子、光效…)
* ``geometry:{type:2,config:{...}}`` + ``material:[{uniforms:[{key:"diffuse",value:"main_sky_jpg"}]}]``
—— 一块**几何平面**,贴的是场景图。用户口中的 "geometric" 就是这个(Spine 官方没有这个概念)。
这里不用正则去"抠字段"——那正是旧实现踩过的坑(父节点的 position/scale 会从子孙子树里被误读)。
我们把整段字面量解析成对象再走树,字段天然只属于自己那一层。
"""
from __future__ import annotations
from dataclasses import dataclass, field
from typing import Any
from . import jslit
__all__ = ["Scene", "Part", "find_scene_list", "parse_scene", "pick_default_scene"]
_SCENE_KEY = "sceneList"
@dataclass
class Part:
"""场景里的一个可绘制件。"""
kind: str # "spine" | "image"(有贴图的平面) | "solid"(纯色平面,不需要资源)
id: str # 骨架逻辑名,或几何平面贴的图名
node_path: str
order: int # 树序遍历序 = 绘制层级
render_order: int
position: tuple[float, float, float] # 世界位移(父 scale 已乘进来)
scale: tuple[float, float, float] # 世界缩放(逐层连乘)
local_position: tuple[float, float, float] | None
local_scale: tuple[float, float, float] | None
rotation: tuple[float, float, float] | None
geometry_type: int | None
auto_matrix: bool
geometry_size: tuple[float, float] | None = None
geometry_center: tuple[float, float] | None = None
modifiers: tuple[str, ...] = ()
# 页面在这个节点上指定的动画(spine.defaultAnimation);缺省 = 骨架第一个动画。
animation: str | None = None
skin: str | None = None
time_scale: float | None = None
def as_dict(self) -> dict[str, Any]:
out: dict[str, Any] = {
"kind": self.kind,
"id": self.id,
"path": self.node_path,
"order": self.order,
"renderOrder": self.render_order,
"position": list(self.position),
"scale": list(self.scale),
}
if self.local_position is not None:
out["localPosition"] = list(self.local_position)
if self.local_scale is not None:
out["localScale"] = list(self.local_scale)
if self.rotation is not None:
out["rotation"] = list(self.rotation)
if self.geometry_type is not None:
out["geometryType"] = self.geometry_type
if self.geometry_size is not None:
out["geometrySize"] = list(self.geometry_size)
if self.geometry_center is not None:
out["geometryCenter"] = list(self.geometry_center)
if self.auto_matrix:
out["autoMatrix"] = True
if self.modifiers:
out["modifiers"] = list(self.modifiers)
if self.animation:
out["animation"] = self.animation
if self.skin:
out["skin"] = self.skin
if self.time_scale is not None:
out["timeScale"] = self.time_scale
return out
@dataclass
class Scene:
"""一个场景及其全部可绘制件。"""
id: str
ui: tuple[int, int] | None
camera: dict[str, Any] | None
parts: list[Part] = field(default_factory=list)
node_count: int = 0
max_depth: int = 0
rotated_nodes: int = 0 # 带 rotation 的节点数(当前不参与合成,见下)
@property
def spine_ids(self) -> list[str]:
seen: list[str] = []
for p in self.parts:
if p.kind == "spine" and p.id not in seen:
seen.append(p.id)
return seen
def as_dict(self, *, page: str | None = None, game: str | None = None) -> dict[str, Any]:
out: dict[str, Any] = {
"version": 1,
"id": self.id,
"ui": list(self.ui) if self.ui else None,
"camera": self.camera,
"parts": [p.as_dict() for p in self.parts],
"stats": {
"nodes": self.node_count,
"maxDepth": self.max_depth,
"spines": len(self.spine_ids),
"images": sum(1 for p in self.parts if p.kind == "image"),
"solids": sum(1 for p in self.parts if p.kind == "solid"),
"rotatedNodes": self.rotated_nodes,
},
}
if game:
out["game"] = game
if page:
out["page"] = page
return out
def find_scene_list(text: str) -> list[dict[str, Any]]:
"""定位并解析 ``sceneList`` 数组;找不到返回空列表。
两种书写风格:压缩后的 ``sceneList:[`` 与已序列化的 ``"sceneList": [``。
"""
for pattern in (f'"{_SCENE_KEY}"', f"{_SCENE_KEY}:"):
idx = text.find(pattern)
if idx < 0:
continue
start = text.find("[", idx)
if start < 0:
continue
end = jslit.match_literal(text, start)
if end < 0:
continue
data = jslit.loads(text[start : end + 1])
if isinstance(data, list) and data:
return [s for s in data if isinstance(s, dict)]
return []
def _vec3(value: Any) -> tuple[float, float, float] | None:
if isinstance(value, (list, tuple)) and len(value) >= 3:
try:
return (float(value[0]), float(value[1]), float(value[2]))
except (TypeError, ValueError):
return None
if isinstance(value, (int, float)):
return (float(value),) * 3
return None
def _material_info(node: dict[str, Any]) -> tuple[str | None, bool]:
"""返回(diffuse 名, 是否真的用贴图)。
``defines.USE_TEXTURE == 0`` 的平面是**纯色块**(diffuse 通常写 "DEFAULT"),
页面上并不存在对应图片;把它当资源去下载只会抓到一个不存在的 URL。
"""
diffuse: str | None = None
textured = True
for material in node.get("material") or []:
if not isinstance(material, dict):
continue
defines = material.get("defines")
if isinstance(defines, dict) and defines.get("USE_TEXTURE") == 0:
textured = False
for uniform in material.get("uniforms") or []:
if isinstance(uniform, dict) and uniform.get("key") == "diffuse":
value = uniform.get("value")
if isinstance(value, str):
diffuse = value
return diffuse, textured
def _modifiers(node: dict[str, Any]) -> tuple[str, ...]:
"""节点的 modifier id 列表(``cacheContainer`` = 渲染到贴图缓冲、``CSS3DObject`` = DOM 叠层)。"""
ids: list[str] = []
for modifier in node.get("modifier") or []:
if isinstance(modifier, dict) and isinstance(modifier.get("id"), str):
ids.append(modifier["id"])
return tuple(ids)
def mesh_bbox(mesh: Any) -> tuple[tuple[float, float], tuple[float, float]] | None:
"""glTF 网格的顶点包围盒 → (尺寸, 相对节点原点的中心)。
`position.array` 是扁平的 xyz 三元组序列(实测这些网格都是四边形)。中心不为零时,
平面并不以节点原点为中心,画的时候要按它偏移。
"""
if not isinstance(mesh, dict):
return None
position = ((mesh.get("attributes") or {}).get("position") or {})
array = position.get("array")
if not isinstance(array, list) or len(array) < 9:
return None
xs: list[float] = []
ys: list[float] = []
for index, value in enumerate(array):
if not isinstance(value, (int, float)):
continue
if index % 3 == 0:
xs.append(float(value))
elif index % 3 == 1:
ys.append(float(value))
if not xs or not ys:
return None
return (
(max(xs) - min(xs), max(ys) - min(ys)),
((min(xs) + max(xs)) / 2, (min(ys) + max(ys)) / 2),
)
def parse_scene(
scene: dict[str, Any],
geometries: dict[str, Any] | None = None,
timeline: dict[str, Any] | None = None,
) -> Scene:
"""把一棵场景树摊平成 part 列表,并合成世界变换。
合成语义照抄引擎(three.js 系)的节点语义,与旧项目 v2 侧车一致:
world_position = parent_position + parent_scale ⊙ local_position
world_scale = parent_scale ⊙ local_scale
旋转暂不参与合成(旧项目同样如此)——只在 stats 里报出数量,别假装算了。
"""
scene_id = str(scene.get("id") or scene.get("name") or "scene")
config = scene.get("sceneConfig") or {}
ui = None
if isinstance(config, dict) and config.get("uiWidth") and config.get("uiHeight"):
ui = (int(config["uiWidth"]), int(config["uiHeight"]))
camera = None
cameras = scene.get("camera")
if isinstance(cameras, list) and cameras:
first = cameras[0]
if isinstance(first, dict):
camera = {
"id": first.get("id"),
"position": first.get("position"),
"rotation": first.get("rotation"),
"camera": first.get("camera"),
}
result = Scene(id=scene_id, ui=ui, camera=camera)
order = 0
def walk(node: Any, path: str, depth: int, parent_pos: tuple[float, float, float], parent_scale: tuple[float, float, float]) -> None:
nonlocal order
if not isinstance(node, dict):
return
name = str(node.get("name") or node.get("id") or "?")
node_path = f"{path}/{name}"
result.node_count += 1
result.max_depth = max(result.max_depth, depth)
# 时间线**按节点名覆盖** position/scale(源码:parsePath → getObjectByName → target[prop] = 值)。
# 静态值是入场前的状态,稳定态要用轨道末帧值。
overridden_pos = None
overridden_scale = None
if timeline:
tracks = timeline.get(scene_id) or {}
overridden_pos = (tracks.get("positions") or {}).get(name)
overridden_scale = (tracks.get("scales") or {}).get(name)
local_pos = _vec3(overridden_pos) if overridden_pos else _vec3(node.get("position"))
local_scale = _vec3(overridden_scale) if overridden_scale else _vec3(node.get("scale"))
rotation = _vec3(node.get("rotation"))
if rotation and any(abs(v) > 1e-9 for v in rotation):
result.rotated_nodes += 1
world_pos = parent_pos
if local_pos:
world_pos = tuple(parent_pos[i] + parent_scale[i] * local_pos[i] for i in range(3))
world_scale = parent_scale
if local_scale:
world_scale = tuple(parent_scale[i] * local_scale[i] for i in range(3))
spine = node.get("spine")
spine_id = spine.get("id") if isinstance(spine, dict) else None
spine_animation = None
spine_skin = None
spine_time_scale = None
if isinstance(spine, dict):
# 页面在这个节点上指定的播放方式——不读它就会去播骨架的第一个动画(常常是入场动画 in)。
if isinstance(spine.get("defaultAnimation"), str):
spine_animation = spine["defaultAnimation"]
if isinstance(spine.get("skin"), str):
spine_skin = spine["skin"]
if isinstance(spine.get("timeScale"), (int, float)):
spine_time_scale = float(spine["timeScale"])
diffuse, textured = _material_info(node)
modifiers = _modifiers(node)
geometry = node.get("geometry")
geometry_type = geometry.get("type") if isinstance(geometry, dict) else None
geometry_size = None
geometry_center = None
if isinstance(geometry, dict):
config = geometry.get("config")
if isinstance(config, dict) and config.get("width") and config.get("height"):
geometry_size = (float(config["width"]), float(config["height"]))
elif geometry.get("type") == 1 and isinstance(geometry.get("id"), str) and geometries:
box = mesh_bbox(geometries.get(geometry["id"]))
if box is not None:
geometry_size, geometry_center = box
if isinstance(spine_id, str) and spine_id:
result.parts.append(
Part(
kind="spine",
id=spine_id,
node_path=node_path,
order=order,
render_order=int(node.get("renderOrder") or 0),
position=world_pos,
scale=world_scale,
local_position=local_pos,
local_scale=local_scale,
rotation=rotation,
geometry_type=None,
auto_matrix=bool(node.get("autoMatrix")),
modifiers=modifiers,
animation=spine_animation,
skin=spine_skin,
time_scale=spine_time_scale,
)
)
order += 1
elif isinstance(diffuse, str) and diffuse and geometry is not None:
result.parts.append(
Part(
kind="image" if textured and diffuse != "DEFAULT" else "solid",
id=diffuse,
node_path=node_path,
order=order,
render_order=int(node.get("renderOrder") or 0),
position=world_pos,
scale=world_scale,
local_position=local_pos,
local_scale=local_scale,
rotation=rotation,
geometry_type=int(geometry_type) if isinstance(geometry_type, int) else None,
auto_matrix=bool(node.get("autoMatrix")),
geometry_size=geometry_size,
geometry_center=geometry_center,
modifiers=modifiers,
)
)
order += 1
for child in node.get("children") or []:
walk(child, node_path, depth + 1, world_pos, world_scale)
for child in scene.get("children") or []:
walk(child, scene_id, 1, (0.0, 0.0, 0.0), (1.0, 1.0, 1.0))
return result
def pick_default_scene(scenes: list[Scene]) -> Scene | None:
"""默认场景 = 骨架最多的那个;并列时取先出现的。
这与旧项目的规则一致("取 spine 最多的一个场景")。它只是**默认值**:
一旦 `sources.yml` 或 `selection.yml` 指定了场景,就以人/记录为准。
"""
if not scenes:
return None
return max(scenes, key=lambda s: (len(s.spine_ids), -scenes.index(s)))
+165
View File
@@ -0,0 +1,165 @@
"""场景的选择:人写的 `sources.yml` 优先,其次机器写的 `selection.yml`,最后才问人。
为什么要第二份文件:`sources.yml` 是人手写的(注释、顺序、措辞都是人的),用 PyYAML 回写会把
这些全吃掉。所以机器只写自己那一份 `selection.yml`,人写的文件永远不被机器改写。
优先级(明确写死,避免两份文件打架):
1. `sources.yml` 里这一页写了 `scene:` / `spines:` → 用它(人的意志最高)
2. `selection.yml` 里有这一页的记录 → 用它(上次交互的结果)
3. 都没有 → 默认取骨架最多的场景 + 该场景全部骨架;
开了 `--interactive` 才问人
**这份选择不再裁剪产物**:`fetch` 会把页面里每个有资源的场景都落成一个壁纸目录(见
`tools/downloader/README.md`)。选择只决定"哪一档是推荐的"——写进 `page.json` 的
`chosenScene` / `chosenDir`,当 `promote` 没给 `--scene` 时的默认值。记录里的 `spines` /
`images` 字段保留只为兼容旧文件,不再有任何裁剪语义。
"""
from __future__ import annotations
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Iterable, Sequence
import yaml
__all__ = ["Choice", "load_selection", "save_selection", "resolve", "ask"]
_HEADER = (
"# 场景/骨架选择记录(由 `python -m tools.downloader select` 写入,机器所有)。\n"
"# 人写的来源清单是 wallpapers/sources.yml;这里只记「哪一页推荐哪个场景」。\n"
"# 入库是为了换台机器重跑能复现同一份推荐值。\n"
"# 注意:fetch 会落盘**全部**有资源的场景,这份记录不再裁剪产物。\n"
"# 优先级:sources.yml 的显式声明 > 本文件 > 默认值(骨架最多的场景)。\n"
)
@dataclass
class Choice:
"""一页的选择结果。"""
page: str
scene: str
spines: list[str] = field(default_factory=list)
source: str = "default" # default | sources | selection | interactive
images: list[str] = field(default_factory=list)
def as_dict(self) -> dict[str, Any]:
return {"scene": self.scene, "spines": list(self.spines), "images": list(self.images)}
def load_selection(path: Path) -> dict[str, Any]:
if not path.exists():
return {}
data = yaml.safe_load(path.read_text(encoding="utf-8")) or {}
return data if isinstance(data, dict) else {}
def save_selection(path: Path, selection: dict[str, Any]) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
body = yaml.safe_dump(selection, allow_unicode=True, sort_keys=True, default_flow_style=False)
path.write_text(_HEADER + body, encoding="utf-8")
def resolve(
page: str,
scenes: Sequence[Any],
*,
sources_entry: dict[str, Any] | None = None,
selection: dict[str, Any] | None = None,
default_scene: str | None = None,
default_spines: Iterable[str] = (),
interactive: bool = False,
input_fn: Any = input,
output_fn: Any = print,
) -> Choice:
"""按优先级算出一页要抓的场景与骨架。"""
selection = selection or {}
record = selection.get(page) or {}
scene = None
spines: list[str] = []
source = "default"
if sources_entry and sources_entry.get("scene"):
scene = str(sources_entry["scene"])
source = "sources"
if sources_entry and sources_entry.get("spines"):
spines = [str(s) for s in sources_entry["spines"]]
source = "sources"
if scene is None and record.get("scene"):
scene = str(record["scene"])
source = "selection"
if not spines and record.get("spines"):
spines = [str(s) for s in record["spines"]]
if source == "default":
source = "selection"
if scene is None:
scene = default_scene
if not spines:
spines = list(default_spines)
# 记录里的场景若一个骨架都没有,它就是个**选不出壁纸**的记录(交互误选、或页面改版了)。
# 这种情况退回默认并标明来源,别让人拿到一个必然报红的产物。
if scenes and scene and scene != default_scene:
chosen_scene = next((s for s in scenes if s.id == scene), None)
if chosen_scene is not None and not chosen_scene.spine_ids and default_scene:
source = f"{source}(无效→默认)"
scene = default_scene
spines = list(default_spines)
if interactive and scenes:
chosen = ask(page, scenes, scene=scene, spines=spines, input_fn=input_fn, output_fn=output_fn)
if chosen is not None:
scene, spines = chosen
source = "interactive"
return Choice(page=page, scene=scene or "", spines=spines, source=source)
def ask(
page: str,
scenes: Sequence[Any],
*,
scene: str | None,
spines: Sequence[str],
input_fn: Any = input,
output_fn: Any = print,
) -> tuple[str, list[str]] | None:
"""交互式选择:只选**推荐场景**。
为什么不再多选骨架:产物现在是"一个场景一个自包含的壁纸目录",砍掉几具骨架会让这个场景的
`sceneConfig` 缺件——那是坏壁纸,不是精简。要裁就在预设里裁(`preset.template.json` 是纯数据)。
"""
output_fn(f"\n[{page}] 检测到 {len(scenes)} 个场景(**全部**会落盘,这里选的是推荐值):")
for i, s in enumerate(scenes, 1):
mark = " ← 推荐" if s.id == scene else ""
has_assets = any(p.kind in ("spine", "image") for p in s.parts)
note = "" if has_assets else "(无资源,不会落盘)"
output_fn(
f" {i}) {s.id} 骨架 {len(s.spine_ids)} 个,图片 "
f"{sum(1 for p in s.parts if p.kind == 'image')} 个{note}{mark}"
)
raw = input_fn("推荐场景编号(回车取推荐值):").strip()
if raw:
try:
index = int(raw) - 1
if not 0 <= index < len(scenes):
output_fn("编号超出范围,取推荐值。")
else:
scene = scenes[index].id
except ValueError:
output_fn("不是数字,取推荐值。")
target = next((s for s in scenes if s.id == scene), None)
if target is None:
return None
if not any(p.kind in ("spine", "image") for p in target.parts):
# 一件资源都没有的场景落不成壁纸(`fetch` 会跳过它)。与其让人推荐一个必然没有产物的
# 场景,不如当场退回默认并说明原因。
output_fn(f"[{page}] {scene} 里没有任何需要资源的 part,退回默认场景。")
return None
return scene, list(spines)
+125
View File
@@ -0,0 +1,125 @@
# 场景/骨架选择记录(由 `python -m tools.downloader select` 写入,机器所有)。
# 人写的来源清单是 wallpapers/sources.yml;这里只记「哪一页推荐哪个场景」。
# 入库是为了换台机器重跑能复现同一份推荐值。
# 注意:fetch 会落盘**全部**有资源的场景,这份记录不再裁剪产物。
# 优先级:sources.yml 的显式声明 > 本文件 > 默认值(骨架最多的场景)。
back-moon:
images: []
scene: scene_main
spines:
- kv_dahua_la
- kv_dahua_lb
- kv_dahua_b
- kv_dahua_c
- kv_dahua_d
- kv_dahua_e
- kv_dahua_f
- kv_dahua_g
- kv_dahua_h
- kv_dahua_i
- kv_dahua_k
- kv_dahua_l
- kv_qj_xhua_a
- kv_qj_xhua_b
- kv_fc_xhua_b
- kv_fc_xhua_c
- kv_fc_xhua_d
- kv_fc_xhua_e
- kv_fc_xhua_f
- kv_fc_xhua_g
- kv_fc_xhua_h
- kv_fc_xhua_i
- kv_xhua_c
- kv_yezhi_a
- kv_yc_hua_a
- kv_zc_hua_b
- kv_zc_hua_c
- kv_zc_hua_d
- kv_zc_hua_e
- kv_yc_hua_b
- kv_xf_hua_a
- kv_shaonv
get-memory:
images: []
scene: P1
spines:
- loading_pass
- p1_icon
- p2_huiyi_d
- p2_huiyi_c
- p2_huiyi_b
- p2_huiyi_a
- p1_traveler
- p1_starline
- p1_huiyi_line
- p2_huiyi_line
- p1_floor
- p1_leaf
- p1_moon_stone
- p1_girl
- p1_moon
- p1_star
- p1_montain
kv45:
images: []
scene: scene_main
spines:
- 01_beijing
- 02_shajin
- 03_zhigengniao
- 04_qianjing
kv46:
images: []
scene: scene_main
spines:
- jh01_46kv_bg
- jh02_46kv_npc
- jh03_46kv_boss
- jh04_46kv_an
- jh05_46kv_zhengzhu
nico-tea:
images: []
scene: scene_main
spines:
- main_beizi_e
- main_beizi_f
- main_beizi_g
- main_beizi_d
- main_d_book_xfl
- main_d_books
- main_xinlin
- main_d_hua
- main_d_dianxin
- main_nike
- main_tree_a
- main_tree_b
- main_cao_g
- main_cao_h
- main_cao_f
start-ndkl:
images: []
scene: scene_main
spines:
- 7FEI_BG
- 6AI_BG7
- 6AI_BG6
- 6AI_BG5
- 6AI_BG4
- 6AI_BG3
- 6AI_BG2
- 6AI_BG1
- 5GUODU
- 4FEI_REN
- 3LA_BG
- 2AI_REN
- 1LA_REN
- 0FG
zhidong-wonder:
images: []
scene: scene_draw
spines:
- d_box_light
- box_click
- box
- d_tips_bg
- d_snow_icon
+23
View File
@@ -0,0 +1,23 @@
"""站点适配层。
一个站点适配器只做三件事:**发现入口 bundle**、**从 bundle 里把资源清单抠出来**、**把逻辑名解析成
可下载的 URL**。通用管线(选择场景、落盘、自检)不认站点,只认这里吐出的数据结构——将来接别的
网站,新增一个同形状的模块即可。
"""
from __future__ import annotations
from . import mihoyo
__all__ = ["mihoyo", "for_url", "SITE_MODULES"]
SITE_MODULES = [mihoyo]
def for_url(url: str) -> mihoyo.Site | None:
"""按 URL 选站点适配器;没有匹配的返回 None。"""
for module in SITE_MODULES:
site = module.match(url)
if site is not None:
return site
return None
+730
View File
@@ -0,0 +1,730 @@
"""米哈游活动页适配器。
事实基线(2026-09-20 对 6 个页面的实爬与逆向,卷宗见 .scratch/mhy-recon/REPORT.md):
* **骨架数据一律内联在入口 bundle 里**——6 个页面里 `.atlas` / `.skel` 的网络请求数是 0。
只有 hsr 的 kv45 例外:它的 skeleton json 走网络(页面根目录 `<hash>.json`),atlas 仍是内联。
* 内联有三种家族:
- **A**(nico-tea / back-moon / zhidong-wonder):
``Object.values(Object.assign({"…/spine/<N>.json":{…}}))[0]``,值是 **JS 对象字面量**(键不带引号)。
- **B**(get-memory / start-ndkl):atlas 与骨架都是匿名 webpack 模块,靠配对表
``"<N>":{atlas:<fn>(<id>),json:<fn>(<id>)}`` 关联。
- **C**(kv45):``spineSetting`` 表里 ``{atlas:"<文本>", json:<模块变量>}``。
* **贴图页走网络**,URL 由 bundle 里的资源表给出,形如
``images/<逻辑名>.<hash>..png``(ys)与 ``assets/images/<逻辑名>.<hash>.<hash>.png``(hsr)。
逻辑名就是 atlas 第一行的页名——**去 hash 就是把它换回逻辑名**。
* 同一逻辑名可能有多个候选(引擎有桌面/移动两套预载表)。判据:**数组式描述表是基准集,
字典式表是移动端覆盖**(引擎里是 `desktop() || base.forEach(e => override[e.id] && …)`),
桌面取基准集。
"""
from __future__ import annotations
import json
import re
import urllib.error
import urllib.request
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
from .. import atlas as atlas_mod
from .. import jslit
from .. import scene as scene_mod
__all__ = ["Site", "PageData", "SpineAsset", "match", "fetch_page", "HttpError"]
_DATA_URL = re.compile(r"^data:(?P<mime>[^;,]+)(?P<b64>;base64)?,(?P<body>.*)$", re.DOTALL)
_MIME_EXT = {"image/png": ".png", "image/jpeg": ".jpg", "image/webp": ".webp", "image/gif": ".gif"}
def decode_data_url(url: str) -> bytes:
"""解内联 data: URL(webpack 把小于阈值的图直接塞进 bundle)。"""
import base64
m = _DATA_URL.match(url)
if m is None:
raise HttpError(f"不是合法的 data: URL:{url[:40]}…")
body = m.group("body")
if m.group("b64"):
return base64.b64decode(body)
from urllib.parse import unquote_to_bytes
return unquote_to_bytes(body)
def image_ext(rel: str) -> str:
"""资源路径 → 落盘扩展名(data: URL 走 MIME 判定)。"""
if rel.startswith("data:"):
mime = rel[5:].split(";", 1)[0].split(",", 1)[0]
return _MIME_EXT.get(mime, ".png")
ext = Path(rel.split("?")[0]).suffix.lower()
return ext if ext in (".png", ".webp", ".jpg", ".jpeg", ".gif") else ".png"
_UA = (
"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) "
"Chrome/146.0.0.0 Safari/537.36"
)
_TIMEOUT = 30
class HttpError(RuntimeError):
"""抓取失败(离线缺缓存、HTTP 非 200、网络异常)。"""
@dataclass(frozen=True)
class Site:
"""站点标识。"""
id: str
game: str
host: str
path_prefix: str
@dataclass
class SpineAsset:
"""一具骨架的抓取素材。"""
name: str
atlas_text: str
pages: list[str]
version: str | None = None
animations: list[str] = field(default_factory=list)
skins: list[str] = field(default_factory=list)
json_data: dict[str, Any] | None = None # 内联家族
json_rel: str | None = None # 走网络时相对页面目录的路径
json_text: str | None = None # 已下载/待下载的原文
def as_meta(self, source: str) -> dict[str, Any]:
return {
"name": self.name,
"spine": self.version,
"animations": self.animations,
"skins": self.skins,
"pages": self.pages,
"source": source,
}
@dataclass
class PageData:
"""一个页面的全部抓取素材。"""
id: str
game: str
url: str
site: Site
entry_url: str | None = None
bundles: dict[str, str] = field(default_factory=dict)
spines: dict[str, SpineAsset] = field(default_factory=dict)
scenes: list[scene_mod.Scene] = field(default_factory=list)
images: dict[str, str] = field(default_factory=dict) # 逻辑名 → 相对页面目录的路径
geometries: dict[str, Any] = field(default_factory=dict) # glTF 网格哈希 → 网格数据
timeline: dict[str, dict[str, dict[str, list[float]]]] = field(default_factory=dict) # 场景 → 轨道末帧
warnings: list[str] = field(default_factory=list)
@property
def base_url(self) -> str:
return self.url.rsplit("/", 1)[0] + "/"
def asset_url(self, rel: str) -> str:
"""相对路径 → 绝对 URL;内联 data: URL 原样返回。"""
if rel.startswith("data:"):
return rel
return self.base_url + rel.lstrip("/")
# --------------------------------------------------------------------------------------
# HTTP + 缓存
# --------------------------------------------------------------------------------------
def _cache_path(cache_dir: Path, page_id: str, rel: str) -> Path:
safe = rel.replace("/", "__").split("?")[0]
return cache_dir / page_id / safe
def http_get(url: str, cache_dir: Path | None, *, page_id: str = "_", rel: str | None = None,
offline: bool = False, binary: bool = False) -> bytes:
"""取一个 URL,带磁盘缓存。``offline=True`` 时只用缓存,缺了就报错。"""
if url.startswith("data:"):
return decode_data_url(url)
path = _cache_path(cache_dir, page_id, rel or url) if cache_dir else None
if path is not None and path.exists():
return path.read_bytes()
if offline:
raise HttpError(f"离线模式但缓存缺失:{url}")
req = urllib.request.Request(url, headers={"User-Agent": _UA, "Referer": url})
try:
with urllib.request.urlopen(req, timeout=_TIMEOUT) as resp:
if resp.status != 200:
raise HttpError(f"HTTP {resp.status}:{url}")
data = resp.read()
except urllib.error.HTTPError as exc:
raise HttpError(f"HTTP {exc.code}:{url}") from exc
except urllib.error.URLError as exc:
raise HttpError(f"网络错误:{url}({exc.reason})") from exc
if path is not None:
path.parent.mkdir(parents=True, exist_ok=True)
path.write_bytes(data)
return data
def http_get_text(url: str, cache_dir: Path | None, **kw: Any) -> str:
return http_get(url, cache_dir, **kw).decode("utf-8", errors="replace")
# --------------------------------------------------------------------------------------
# 通用解析辅助
# --------------------------------------------------------------------------------------
_STRING = re.compile(r'"(?:[^"\\]|\\.)*"')
_PUBLIC_PATH = r"[A-Za-z_$][\w$]*\.p"
_ASSET_LITERAL = re.compile(r'([A-Za-z_$][\w$]*)\s*=\s*' + _PUBLIC_PATH + r'\s*\+\s*"([^"]+)"')
_MODULE_LITERAL = re.compile(
r'(\d+):function\([^)]*\)\s*\{[^}]*?e\.exports\s*=\s*' + _PUBLIC_PATH + r'\s*\+\s*"([^"]+)"'
)
# 模块直接导出**字符串**:小图被 webpack 内联成 `e.exports="data:image/png;base64,\u2026"`,也有
# `e.exports="images/x.png"` 这种不带公共路径前缀的写法。描述表里写 `src:a(38458)` 时只能靠
# 这张表还原——不认它就会把"内联的图"当成"根本没引用"(实测 get-memory 的 loading_moutain_a 与
# a01_lizi、start-ndkl 的 loading_start_1 都栽在这里,它们是真资源,不是运行时残留)。
_MODULE_STRING = re.compile(r'(\d+)\s*:\s*function\s*\([^)]*\)\s*\{[^}]*?e\.exports\s*=\s*"((?:[^"\\]|\\.)*)"')
# `src:a(38458)`:直接调 require 函数、参数是模块号。
_REQUIRE_CALL = re.compile(r"\s*[A-Za-z_$][\w$]*\s*\(\s*(\d+)\s*\)\s*")
# `$w="data:image/png;base64,\u2026"`:变量直接存字符串字面量。描述表里写 `{src:$w,…}` 时靠它还原
# (kv45 的 loading_dt1 就是这样,不认它就会被当成"缺资源"直接判红)。
_STRING_ASSIGN = re.compile(r'([A-Za-z_$][\w$]*)\s*=\s*"((?:[^"\\]|\\.)*)"')
_REQUIRE_ID = re.compile(r'([A-Za-z_$][\w$]*)\s*=\s*n\((\d+)\)')
_REQUIRE_URL_ID = re.compile(r'([A-Za-z_$][\w$]*)\s*=\s*new URL\(n\((\d+)\)')
_REQUIRE_NS = re.compile(r'([A-Za-z_$][\w$]*)\s*=\s*n\.n\(([A-Za-z_$][\w$]*)\)')
# 描述表有两种 src 形状,分开匹配——一条 `.+?` 通吃会让它跨过整段 bundle 去够后面那条
# `,id:"…",type:"image"`:kv45 的 `{src:e.activeIcon,alt:""}` 就是这么把
# `{src:$w,id:"loading_dt1",type:"image"}` 整个吞掉的(真表项永远收集不到)。
# inline:`Object.values(Object.assign({"<路径>":"data:…"}))[0]`(webpack 内联的小图)
# simple:`a(38458)` / `$w` / `a.p+"images/x.png"` 这类不含 `,` `{` `}` 的表达式
_ASSIGN_EXPR = r'Object\.values\(Object\.assign\(\{[^{}]*\}\)\)\[0\]'
_DESCRIPTOR_INLINE = re.compile(r'\{src:(' + _ASSIGN_EXPR + r'),id:"([^"]+)",type:"image"\}')
_DESCRIPTOR_SIMPLE = re.compile(r'\{src:([^,{}]+),id:"([^"]+)",type:"image"\}')
# webpack 会把小图内联:`Object.values(Object.assign({"<源路径>":"data:image/png;base64,…"}))[0]`
_DATA_IN_ASSIGN = re.compile(r'Object\.values\(Object\.assign\(\{"[^"]*":("(?:[^"\\]|\\.)*")\}\)\)')
_OVERRIDE_TABLE = re.compile(r'([A-Za-z_$][\w$]*)\s*=\s*\{\s*"([^"]+)"\s*:\s*([A-Za-z_$][\w$]*)\(\)')
def _unescape_js_string(raw: str) -> str:
"""把 ``"…"``(含两端引号)还原成文本。"""
return json.loads(raw)
def _literal_of(expr: str) -> str | None:
"""从 ``Object.values(Object.assign({"<源路径>":<VAR>}))[0]`` 这类表达式里取出变量名。"""
m = re.search(r'Object\.values\(Object\.assign\(\{[^:]*:([A-Za-z_$][\w$]*)\}\)\)', expr)
if m:
return m.group(1)
m = re.fullmatch(r"([A-Za-z_$][\w$]*)\(\)", expr.strip())
if m:
return m.group(1)
m = re.fullmatch(r"([A-Za-z_$][\w$]*)", expr.strip())
if m:
return m.group(1)
return None
class _AssetResolver:
"""把 bundle 里的变量/模块引用解析成真正的资源路径。
压缩后的变量名在**模块作用域**内唯一,所以回退查找要限定在同一个 webpack 模块里;
这里用"从引用处往前找最近的模块头"作为边界,避免跨模块误配。
"""
def __init__(self, text: str) -> None:
self.text = text
self.var_to_literal = {m.group(1): m.group(2) for m in _ASSET_LITERAL.finditer(text)}
self.module_to_literal = {m.group(1): m.group(2) for m in _MODULE_LITERAL.finditer(text)}
self.module_to_string: dict[str, str] = {}
for m in _MODULE_STRING.finditer(text):
try:
# 捕获组里没有引号,而 _unescape_js_string 要的是含引号的原文。
value = _unescape_js_string('"' + m.group(2) + '"')
except json.JSONDecodeError:
continue
# 只认"看起来是资源"的字符串导出:atlas 文本与内联 JSON 也是字符串导出,
# 把它们当路径只会造出一条永远下载失败的假引用。
if value.startswith("data:") or value.lower().endswith((".png", ".webp", ".jpg", ".jpeg", ".gif")):
self.module_to_string.setdefault(m.group(1), value)
self.var_to_string: dict[str, str] = {}
for m in _STRING_ASSIGN.finditer(text):
raw = m.group(2)
if not (raw.startswith("data:") or raw.lower().endswith((".png", ".webp", ".jpg", ".jpeg", ".gif"))):
continue
try:
self.var_to_string.setdefault(m.group(1), _unescape_js_string('"' + raw + '"'))
except json.JSONDecodeError:
continue
self.var_to_module = {m.group(1): m.group(2) for m in _REQUIRE_ID.finditer(text)}
self.var_to_module.update({m.group(1): m.group(2) for m in _REQUIRE_URL_ID.finditer(text)})
self.ns_to_var = {m.group(1): m.group(2) for m in _REQUIRE_NS.finditer(text)}
def resolve_var(self, var: str) -> str | None:
if var in self.var_to_literal:
return self.var_to_literal[var]
if var in self.var_to_string:
return self.var_to_string[var]
target = self.ns_to_var.get(var, var)
if target in self.var_to_literal:
return self.var_to_literal[target]
module_id = self.var_to_module.get(target)
if module_id is not None:
return self.module_to_literal.get(module_id)
return None
def resolve_expr(self, expr: str) -> str | None:
inline = _DATA_IN_ASSIGN.search(expr)
if inline is not None:
try:
return _unescape_js_string(inline.group(1))
except json.JSONDecodeError:
return None
# `src:a(38458)` 这种直接调 require 的写法:模块导出可能是 `p+"\u2026"`,也可能是内联 data:。
call = _REQUIRE_CALL.fullmatch(expr)
if call is not None:
module_id = call.group(1)
return self.module_to_literal.get(module_id) or self.module_to_string.get(module_id)
var = _literal_of(expr)
if var is None:
return None
return self.resolve_var(var)
def _collect_geometries(text: str) -> dict[str, Any]:
"""收集内联的 glTF 网格表(`geometries:{<哈希>:{type,attributes:{position:{array}}}}`)。
场景节点里 `geometry:{type:1,id:"<哈希>"}` 的尺寸不写在场景树上,只能来这里算顶点包围盒。
键有时带引号(JSON.parse 的模块)有时不带(JS 字面量),所以两种都要认。
"""
out: dict[str, Any] = {}
for m in re.finditer(r"geometries", text):
brace = text.find("{", m.end())
if brace < 0 or brace - m.end() > 4:
continue
end = jslit.match_literal(text, brace)
if end < 0:
continue
try:
data = jslit.loads(text[brace : end + 1])
except jslit.JsLitError:
continue
if isinstance(data, dict):
for key, value in data.items():
if isinstance(value, dict) and key not in out:
out[key] = value
return out
_TRACK = re.compile(r'name:"([\w.]+)\.(position|scale)"')
_CLIP_START = "{type:"
# 场景自己在 modifier 里点名播哪条轨道:playTimeline{sceneName, trackName1..6, frame1..6}
_PLAY_TIMELINE = re.compile(r'playTimeline",data:\{sceneName:"([\w-]+)",([^}]*)\}')
_TRACK_ENTRY = re.compile(r'trackName(\d+):"([\w-]*)"')
_FRAME_ENTRY = re.compile(r"frame(\d+):(\d+)")
def _block_tracks(
text: str, block: str, frame: int | None = None
) -> dict[str, dict[str, list[float]]]:
"""取某个**块**(如 `pv`)里 position/scale 轨道在**指定帧**的值。
`frame=None` 时取末帧(块的结束状态);给了帧就取该帧(场景声明的 `frame1`)。
为什么按块取:同名节点(layout/bg/img)在 `loading` 与 `pv` 里各有自己的关键帧,
全局扫名字会把**入场块**的值套到主场景上(实测差得很远:layout.z 539 vs 439,x/y 也不同)。
源码依据:`parsePath` 把 `a.b` 解析成 `getObjectByName("a")` 再写 `.b`——轨道名第一段是节点名,值是**覆盖**。
"""
positions: dict[str, list[float]] = {}
scales: dict[str, list[float]] = {}
anchor = text.find(f"{{{block}:{{type:1,")
if anchor < 0:
return {"positions": positions, "scales": scales}
end = jslit.match_literal(text, anchor)
if end < 0:
return {"positions": positions, "scales": scales}
for m in _TRACK.finditer(text, anchor, end):
name, prop = m.group(1), m.group(2)
back = text.rfind(_CLIP_START, anchor, m.start())
if back < 0 or m.start() - back > 4000:
continue
clip_end = jslit.match_literal(text, back)
if clip_end < 0 or clip_end > end:
continue
try:
clip = jslit.loads(text[back : clip_end + 1])
except jslit.JsLitError:
continue
if not isinstance(clip, dict) or clip.get("name") != f"{name}.{prop}":
continue
data = clip.get("data") or []
if not isinstance(data, list) or not data:
continue
frames = (data[-1] or {}).get("frames") or []
if not isinstance(frames, list) or not frames:
continue
# 帧号语义:块的 clip 从 start 起算,`frames` 是整个区间的关键帧序列。
# 场景声明的帧是「从这条轨道第几帧开始」,直接按该下标取(越界则回退到首/末)。
value = frames[frame] if frame is not None and 0 <= frame < len(frames) else frames[-1]
if not isinstance(value, list):
continue
target = positions if prop == "position" else scales
target.setdefault(name, [float(v) for v in value if isinstance(v, (int, float))])
return {"positions": positions, "scales": scales}
def _collect_timeline(text: str) -> dict[str, dict[str, dict[str, list[float]]]]:
"""按场景归并:`{场景id: {positions: {...}, scales: {...}}}`。
场景播哪条轨道由它自己的 `playTimeline` modifier 声明;只取那条轨道的末帧(= 稳定态)。
"""
per_scene: dict[str, dict[str, dict[str, list[float]]]] = {}
cache: dict[str, dict[str, dict[str, list[float]]]] = {}
for m in _PLAY_TIMELINE.finditer(text):
scene, body = m.group(1), m.group(2)
bucket = per_scene.setdefault(scene, {"positions": {}, "scales": {}})
frames_by_slot = {entry.group(1): int(entry.group(2)) for entry in _FRAME_ENTRY.finditer(body)}
for entry in _TRACK_ENTRY.finditer(body):
track = entry.group(2)
if not track:
continue
frame = frames_by_slot.get(entry.group(1))
key = f"{track}@{frame}"
if key not in cache:
cache[key] = _block_tracks(text, track, frame)
for kind in ("positions", "scales"):
for node, vec in cache[key][kind].items():
bucket[kind].setdefault(node, vec)
return per_scene
def _collect_images(text: str) -> dict[str, str]:
"""从 bundle 的预载描述表里取「逻辑名 → 资源路径」。
数组式描述表(``[{src:…,id:"x",type:"image"}]``)是基准集,先收;字典式
(``{"x":fn()}``)是引擎在非桌面端才套用的覆盖集,只在没有基准候选时兜底。
"""
resolver = _AssetResolver(text)
images: dict[str, str] = {}
for pattern in (_DESCRIPTOR_INLINE, _DESCRIPTOR_SIMPLE):
for m in pattern.finditer(text):
expr, name = m.group(1), m.group(2)
rel = resolver.resolve_expr(expr)
if rel and name not in images:
images[name] = rel
for m in _OVERRIDE_TABLE.finditer(text):
var, name = m.group(3), m.group(2)
rel = resolver.resolve_var(var)
if rel and name not in images:
images[name] = rel
# 兜底:描述表没覆盖到的名字,直接用 `n.p+"images/<逻辑名>.<hash>..png"` 字面量。
# ys 页 99 个字面量与 99 个逻辑名一一对应(无重名),这条兜底因此是确定的;
# hsr 有桌面/移动两套候选,兜底只取先出现的那个。
for m in _ASSET_LITERAL.finditer(text):
rel = m.group(2)
if not re.match(r"^(assets/)?images/", rel):
continue
stem = rel.rsplit("/", 1)[-1].split(".")[0]
images.setdefault(stem, rel)
return images
# --------------------------------------------------------------------------------------
# 骨架提取:三个内联家族
# --------------------------------------------------------------------------------------
_A_JSON = re.compile(r'Object\.values\(Object\.assign\(\{"((?:\.\./\.\.|/)[^"]*?/spine/[\w.-]+\.json)":')
_A_ATLAS = re.compile(r'Object\.values\(Object\.assign\(\{"((?:\.\./\.\.|/)[^"]*?/spine/[\w.-]+\.atlas)":')
_B_ATLAS_MODULE = re.compile(r'(\d+):function\(([\w,\s]*)\)\{"?(?:use strict"?;)?e\.exports=("(?:[^"\\]|\\.)*")\}')
_B_JSON_MODULE = re.compile(r'(\d+):function\(([\w,\s]*)\)\{"?(?:use strict"?;)?e\.exports=JSON\.parse\(')
_B_PAIR = re.compile(r'"?([\w\u4e00-\u9fa5@-]+)"?:\{atlas:([A-Za-z_$][\w$]*)\((\d+)\),json:\2\((\d+)\)\}')
_C_SPINE = re.compile(r'"?([\w@-]+)"?:\{atlas:"((?:[^"\\]|\\.)*)",json:([A-Za-z_$][\w$]*)\}')
def _summary(data: dict[str, Any]) -> tuple[str | None, list[str], list[str]]:
skeleton = data.get("skeleton") or {}
version = skeleton.get("spine") if isinstance(skeleton, dict) else None
animations = sorted((data.get("animations") or {}).keys()) if isinstance(data.get("animations"), dict) else []
raw_skins = data.get("skins")
if isinstance(raw_skins, list):
skins = [s.get("name") if isinstance(s, dict) else s for s in raw_skins]
elif isinstance(raw_skins, dict):
skins = list(raw_skins.keys())
else:
skins = []
return version, animations, [s for s in skins if isinstance(s, str)]
def _extract_family_a(text: str) -> dict[str, SpineAsset]:
assets: dict[str, SpineAsset] = {}
for m in _A_JSON.finditer(text):
name = m.group(1).rsplit("/", 1)[-1][: -len(".json")]
brace = text.find("{", m.end())
if brace < 0:
continue
end = jslit.match_literal(text, brace)
if end < 0:
continue
data = jslit.loads(text[brace : end + 1])
if not isinstance(data, dict):
continue
version, animations, skins = _summary(data)
assets[name] = SpineAsset(
name=name,
atlas_text="",
pages=[],
version=version,
animations=animations,
skins=skins,
json_data=data,
)
for m in _A_ATLAS.finditer(text):
name = m.group(1).rsplit("/", 1)[-1][: -len(".atlas")]
quote = m.end() # 正则正好停在值字符串的左引号上
if quote >= len(text) or text[quote] != '"':
continue
end = quote + 1
while end < len(text):
if text[end] == "\\":
end += 2
elif text[end] == '"':
break
else:
end += 1
atlas_text = _unescape_js_string(text[quote : end + 1])
info = atlas_mod.parse(atlas_text)
asset = assets.get(name)
if asset is None:
continue
asset.atlas_text = atlas_text
asset.pages = info.pages
return {k: v for k, v in assets.items() if v.atlas_text}
def _extract_family_b(text: str) -> dict[str, SpineAsset]:
atlas_modules: dict[str, str] = {}
for m in _B_ATLAS_MODULE.finditer(text):
try:
value = _unescape_js_string(m.group(3))
except json.JSONDecodeError:
continue
first = (value.splitlines() or [""])[0]
if re.search(r"\.(png|webp|jpg|jpeg)\s*$", first, re.IGNORECASE) and re.search(
r"^\s*(size|filter|format)\s*:", value, re.IGNORECASE | re.MULTILINE
):
atlas_modules[m.group(1)] = value
json_modules: dict[str, dict[str, Any]] = {}
for m in _B_JSON_MODULE.finditer(text):
paren = m.end() - 1
end = jslit.match_literal(text, paren)
if end < 0:
continue
raw = text[paren + 1 : end].strip()
if len(raw) >= 2 and raw[0] == "'" and raw[-1] == "'":
raw = jslit.unescape(raw[1:-1])
try:
data = json.loads(raw)
except json.JSONDecodeError:
continue
if isinstance(data, dict) and (data.get("skeleton") or {}).get("spine"):
json_modules[m.group(1)] = data
assets: dict[str, SpineAsset] = {}
for m in _B_PAIR.finditer(text):
name, atlas_id, json_id = m.group(1), m.group(3), m.group(4)
if name in assets:
continue
data = json_modules.get(json_id)
atlas_text = atlas_modules.get(atlas_id, "")
if data is None or not atlas_text:
continue
version, animations, skins = _summary(data)
assets[name] = SpineAsset(
name=name,
atlas_text=atlas_text,
pages=atlas_mod.parse(atlas_text).pages,
version=version,
animations=animations,
skins=skins,
json_data=data,
)
return assets
def _extract_family_c(text: str) -> dict[str, SpineAsset]:
"""kv45:``{atlas:"<文本>", json:<模块变量>}``,json 走网络。"""
resolver = _AssetResolver(text)
assets: dict[str, SpineAsset] = {}
for m in _C_SPINE.finditer(text):
name, raw_atlas, var = m.group(1), m.group(2), m.group(3)
if name in assets:
continue
try:
atlas_text = _unescape_js_string(f'"{raw_atlas}"')
except json.JSONDecodeError:
continue
rel = resolver.resolve_var(var)
if not rel:
continue
assets[name] = SpineAsset(
name=name,
atlas_text=atlas_text,
pages=atlas_mod.parse(atlas_text).pages,
json_rel=rel,
)
return assets
# --------------------------------------------------------------------------------------
# 页面抓取
# --------------------------------------------------------------------------------------
_SCRIPT_SRC = re.compile(r'<script[^>]+src="([^"]+)"', re.IGNORECASE)
_YS_PATH = re.compile(r"/(ys|bh3|hkrpg)/event/")
_HSR_PATH = re.compile(r"/puzzle/hkrpg/")
def match(url: str) -> Site | None:
"""认页面:ys 活动页与 hsr 的 puzzle 页各一个站点标识。"""
if _HSR_PATH.search(url):
return Site(id="mihoyo-hsr-puzzle", game="hsr", host="act.mihoyo.com", path_prefix="/puzzle/hkrpg/")
if _YS_PATH.search(url):
return Site(id="mihoyo-ys-event", game="ys", host="act.mihoyo.com", path_prefix="/ys/event/")
if "act.mihoyo.com" in url:
return Site(id="mihoyo-act", game="mihoyo", host="act.mihoyo.com", path_prefix="/")
return None
def _absolute(base: str, rel: str) -> str:
if rel.startswith("http://") or rel.startswith("https://"):
return rel
if rel.startswith("//"):
return "https:" + rel
if rel.startswith("/"):
return "https://act.mihoyo.com" + rel
return base + rel.lstrip("./")
def _discover_scripts(html: str, base: str) -> list[str]:
return [_absolute(base, src) for src in _SCRIPT_SRC.findall(html)]
_CHUNK_FN = re.compile(r"\.u\s*=\s*function\s*\([^)]*\)\s*\{")
_CHUNK_PAIR = re.compile(r'(\d+)\s*:\s*"([^"]*)"')
def discover_chunk_urls(text: str, base: str) -> list[str]:
"""从 webpack 运行时里解出异步 chunk 的 URL。
形如 ``s.u=function(e){return({416:"lib.pc",833:"lib.m"}[e]||e)+"."+{258:"ccb0954b",…}[e]+".js"}``:
取函数体里所有 ``id:"值"``,按值的样子分成**名字表**(`lib.pc` 这类含点的)与**哈希表**(纯十六进制),
再拼成 ``<名字|id>.<哈希>.js``。
"""
urls: list[str] = []
for m in _CHUNK_FN.finditer(text):
end = jslit.match_literal(text, m.end() - 1)
if end < 0:
continue
body = text[m.end() : end]
names: dict[str, str] = {}
hashes: dict[str, str] = {}
for pair in _CHUNK_PAIR.finditer(body):
chunk_id, value = pair.group(1), pair.group(2)
if re.fullmatch(r"[0-9a-f]{6,}", value):
hashes[chunk_id] = value
elif value:
names[chunk_id] = value
for chunk_id, digest in hashes.items():
urls.append(_absolute(base, f"{names.get(chunk_id, chunk_id)}.{digest}.js"))
return urls
def fetch_page(page_id: str, game: str, url: str, *, cache_dir: Path | None = None,
offline: bool = False, log: Any = print) -> PageData:
"""抓一个页面并解析出骨架、场景、图片表。"""
site = match(url)
if site is None:
raise HttpError(f"没有匹配的站点适配器:{url}")
data = PageData(id=page_id, game=game, url=url, site=site)
base = url.rsplit("/", 1)[0] + "/"
html = http_get_text(url, cache_dir, page_id=page_id, rel="index.html", offline=offline)
scripts = _discover_scripts(html, base)
if not scripts:
raise HttpError(f"页面里没有找到任何 <script src>:{url}")
# hsr 的骨架表在 258.*.js 这类 chunk 里,ys 在入口 index_*.js 里;两者都扫一遍更稳。
ordered = sorted(scripts, key=lambda s: (0 if re.search(r"/(index|258)\.", s) else 1, s))
# 队列而不是 for:hsr 的骨架表在 webpack 异步 chunk 里,入口 HTML 根本没列它,
# 只能边扫边把 `.u=` 映射出来的 chunk 追加进队列。
queue = list(ordered)
seen: set[str] = set()
while queue:
script_url = queue.pop(0)
if script_url in seen:
continue
seen.add(script_url)
name = script_url.rsplit("/", 1)[-1]
try:
text = http_get_text(script_url, cache_dir, page_id=page_id, rel=name, offline=offline)
except HttpError as exc:
data.warnings.append(f"跳过 {name}:{exc}")
continue
data.bundles[name] = text
# 网格表要先收:场景树里的 type-1 平面靠它算尺寸与中心。
for key, mesh in _collect_geometries(text).items():
data.geometries.setdefault(key, mesh)
for scene_id, tracks in _collect_timeline(text).items():
bucket = data.timeline.setdefault(scene_id, {"positions": {}, "scales": {}})
for kind in ("positions", "scales"):
for node, vec in tracks[kind].items():
bucket[kind].setdefault(node, vec)
if not data.scenes:
data.scenes = [
scene_mod.parse_scene(s, data.geometries, data.timeline) for s in scene_mod.find_scene_list(text)
]
for key, rel in _collect_images(text).items():
data.images.setdefault(key, rel)
if not data.spines:
spines = _extract_family_a(text) or _extract_family_b(text) or _extract_family_c(text)
if spines:
data.spines = spines
data.entry_url = script_url
if data.spines and data.scenes and data.images:
break
for chunk_url in discover_chunk_urls(text, base):
if chunk_url not in seen:
queue.append(chunk_url)
if not data.spines:
data.warnings.append("没有从任何 bundle 里提取到骨架(页面结构可能变了)")
if not data.scenes:
data.warnings.append("没有找到 sceneList(场景树)")
return data
def load_spine_json(data: PageData, asset: SpineAsset, *, cache_dir: Path | None = None,
offline: bool = False) -> dict[str, Any]:
"""取骨架数据:内联家族直接用内存里的对象,kv45 去下载 ``<hash>.json``。"""
if asset.json_data is not None:
return asset.json_data
if asset.json_rel is None:
raise HttpError(f"骨架 {asset.name} 既没有内联数据也没有 json 路径")
if asset.json_text is None:
asset.json_text = http_get_text(
data.asset_url(asset.json_rel), cache_dir, page_id=data.id, rel=asset.json_rel, offline=offline
)
parsed = json.loads(asset.json_text)
version, animations, skins = _summary(parsed)
asset.version = version
asset.animations = animations
asset.skins = skins
return parsed
+233
View File
@@ -0,0 +1,233 @@
"""`_out/` 的自检:产物是不是真的自洽。
抓取最怕的不是"下不动",是"下了一半还长得像成功":侧车引用的骨架漏了一个、atlas 声明的贴图页
没落地、页名与文件名对不上(运行时就会去请求一个不存在的页)、预设里的 `./…` 路径找不到文件。
所以每次 `fetch` 结束都会跑一遍,`verify` 子命令也能单独跑;**退出码非 0 = 有硬伤**。
形状(一个场景 = 一个可直接搬走的壁纸目录,见 `layout.py`)::
<root>/<游戏>/<页面>/page.json
<root>/<游戏>/<页面>/<场景id>/meta.json | preset.template.json | scene.json
<root>/<游戏>/<页面>/<场景id>/spines/<名>/… scene/… audios/
约定:这里只做"引用闭包 + 文件齐全 + 页名一致 + 版本可读 + 预设路径存在"五件事,不评判画面对不对。
"""
from __future__ import annotations
import json
import re
from pathlib import Path
from typing import Any
from . import atlas as atlas_mod
from . import layout as layout_mod
__all__ = ["verify_page", "verify_all", "verify_scene_dir", "SUPPORTED_RUNTIME"]
# 本仓库 vendored 的播放器是 spine-ts 4.2 线:4.0/4.1/4.2 的数据都能读,
# 4.3 起把约束并进 root.constraints,4.2 解析器会静默丢掉全部约束——所以要显式报出来。
SUPPORTED_RUNTIME = (4, 2)
_EXT = (".png", ".webp", ".jpg", ".jpeg")
_ID_OK = re.compile(r"^[a-z0-9][a-z0-9_-]*$", re.IGNORECASE)
def _read_json(path: Path) -> Any:
return json.loads(path.read_text(encoding="utf-8"))
def _runtime_of(version: str | None) -> tuple[int, int] | None:
if not version:
return None
m = re.match(r"(\d+)\.(\d+)", version)
if not m:
return None
return (int(m.group(1)), int(m.group(2)))
def _is_scene_dir(path: Path) -> bool:
"""场景目录 = 有 `preset.template.json`。判据不能只看目录名——目录名是转义过的场景 id。"""
return (path / "preset.template.json").is_file()
def verify_scene_dir(scene_dir: Path) -> tuple[list[str], dict[str, Any]]:
"""检查一个**场景目录**(= 一份可直接搬走的壁纸目录);返回(问题列表, 统计)。"""
problems: list[str] = []
stats: dict[str, Any] = {"spines": 0, "images": 0, "versions": {}, "bytes": 0}
for name in ("meta.json", "preset.template.json", "scene.json"):
if not (scene_dir / name).is_file():
problems.append(f"缺少 {name}")
try:
scene = _read_json(scene_dir / "scene.json")
except (OSError, json.JSONDecodeError) as exc:
return [*problems, f"scene.json 读不出来:{exc}"], stats
parts = scene.get("parts") or []
if not parts:
problems.append("scene.json 里一个 part 都没有(场景选择可能失败)")
# meta.json:这些字段是构建期 fail-fast 校验的同一批(tools/lib/vault.ts),这里提前红。
meta_file = scene_dir / "meta.json"
if meta_file.is_file():
try:
meta = _read_json(meta_file)
except json.JSONDecodeError as exc:
problems.append(f"meta.json 不是合法 JSON:{exc}")
meta = None
if isinstance(meta, dict):
if meta.get("id") != scene_dir.name:
problems.append(f"meta.json 的 id {meta.get('id')!r} 与目录名 {scene_dir.name!r} 不一致")
if not _ID_OK.match(str(meta.get("id") or "")):
problems.append(f"meta.json 的 id {meta.get('id')!r} 含非法字符(只能是字母数字与 _-)")
for key in ("name", "title", "description"):
if not meta.get(key):
problems.append(f"meta.json 缺 {key}")
audio = meta.get("audio")
if not isinstance(audio, dict) or not isinstance(audio.get("choices"), list):
problems.append("meta.json 的 audio.choices 不是列表")
else:
for index, choice in enumerate(audio["choices"]):
rel = (choice or {}).get("file") if isinstance(choice, dict) else None
if not rel:
problems.append(f"meta.json 的 audio.choices[{index}] 缺 file")
elif not (scene_dir / str(rel).replace("\\", "/").removeprefix("./")).is_file():
problems.append(f"meta.json 声明了音源 {rel},但磁盘上没有这个文件")
preset_file = scene_dir / "preset.template.json"
if preset_file.is_file():
try:
preset = _read_json(preset_file)
except json.JSONDecodeError as exc:
problems.append(f"preset.template.json 不是合法 JSON:{exc}")
preset = None
if isinstance(preset, dict):
scene_config = preset.get("sceneConfig")
spine_config = preset.get("spineConfig")
if bool(scene_config) == bool(spine_config):
problems.append("preset.template.json 必须二选一地写 sceneConfig 或 spineConfig")
# 全是运行时纹理的场景(各页的 scene_ui:drawScene 渲染目标)本来就凑不出 part,
# 那种目录只是"留档不可搬走";但 scene.json 里有该落盘的 part 而预设漏了,是硬伤。
drawable = [
p for p in parts
if p.get("kind") == "spine" or (p.get("kind") == "image" and not p.get("runtime"))
]
if isinstance(scene_config, dict) and not (scene_config.get("parts") or []) and drawable:
problems.append("preset.template.json 的 sceneConfig.parts 是空的,但 scene.json 里有可落盘的 part")
for rel in layout_mod.missing_preset_paths(preset, scene_dir):
problems.append(f"preset.template.json 引用的 {rel} 不在磁盘上")
for part in parts:
kind, pid = part.get("kind"), part.get("id")
if kind == "spine":
stats["spines"] += 1
spine_dir = scene_dir / layout_mod.SPINE_DIR / str(pid)
json_file = spine_dir / f"{pid}.json"
atlas_file = spine_dir / f"{pid}.atlas"
spine_meta = spine_dir / "meta.json"
for path in (json_file, atlas_file, spine_meta):
if not path.exists():
problems.append(f"骨架 {pid} 缺少 {path.name}")
if not atlas_file.exists():
continue
atlas_text = atlas_file.read_text(encoding="utf-8")
info = atlas_mod.parse(atlas_text)
if not info.pages:
problems.append(f"骨架 {pid} 的 atlas 里没有贴图页行")
# 页名**读 atlas 自己声明的**,与落盘文件名逐字比对——按 `<stem>_N` 猜会漏掉多页。
for page_name in info.pages:
if not (spine_dir / page_name).exists():
problems.append(f"骨架 {pid} 缺少贴图页 {page_name}")
if spine_meta.exists():
try:
meta = _read_json(spine_meta)
except json.JSONDecodeError as exc:
problems.append(f"骨架 {pid} 的 meta.json 不是合法 JSON:{exc}")
continue
version = meta.get("spine")
runtime = _runtime_of(version)
if version:
stats["versions"][str(version)] = stats["versions"].get(str(version), 0) + 1
if runtime and runtime > SUPPORTED_RUNTIME:
problems.append(
f"骨架 {pid} 是 {version} 格式,超出 vendored 播放器(4.2 线)能读的范围"
)
elif kind == "image":
if part.get("runtime"):
continue # 由骨架渲染进贴图缓冲的平面没有独立文件,不是缺失
stats["images"] += 1
if layout_mod.image_file(scene_dir, str(pid)) is None:
problems.append(f"几何平面的贴图缺失:scene/{pid}.*")
for path in scene_dir.rglob("*"):
if path.is_file():
stats["bytes"] += path.stat().st_size
return problems, stats
def verify_page(page_dir: Path) -> tuple[list[str], dict[str, Any]]:
"""检查一个页面目录(它下面每个场景目录各是一份壁纸目录)。"""
problems: list[str] = []
stats: dict[str, Any] = {"sceneDirs": 0, "spines": 0, "images": 0, "versions": {}, "bytes": 0}
page_file = page_dir / "page.json"
if not page_file.exists():
return [f"缺少 page.json:{page_dir}"], stats
try:
_read_json(page_file)
except json.JSONDecodeError as exc:
return [f"page.json 不是合法 JSON:{exc}"], stats
# 旧形状(一页一个场景)的残留:页面根的 scene.json / spine/ / scene/ 现在都不该存在。
for legacy in ("scene.json", "spine", "scene"):
if (page_dir / legacy).exists():
problems.append(f"页面根上还留着旧形状的 {legacy}(新形状里它属于每个场景目录)")
scenes = sorted(p for p in page_dir.iterdir() if p.is_dir() and _is_scene_dir(p))
strays = sorted(
p.name for p in page_dir.iterdir() if p.is_dir() and not _is_scene_dir(p) and p.name not in ("spine", "scene")
)
if not scenes:
problems.append("页面下没有任何场景目录(先跑 fetch)")
for name in strays:
problems.append(f"页面下的 {name}/ 不是场景目录(缺 preset.template.json)")
for scene_dir in scenes:
scene_problems, scene_stats = verify_scene_dir(scene_dir)
problems.extend(f"[{scene_dir.name}] {p}" for p in scene_problems)
stats["sceneDirs"] += 1
stats["spines"] += scene_stats["spines"]
stats["images"] += scene_stats["images"]
stats["bytes"] += scene_stats["bytes"]
for version, count in scene_stats["versions"].items():
stats["versions"][version] = stats["versions"].get(version, 0) + count
return problems, stats
def verify_all(root: Path, pages: list[str] | None = None) -> tuple[list[str], dict[str, Any]]:
"""检查 `_out/` 下的全部(或指定)页面。"""
problems: list[str] = []
total: dict[str, Any] = {"pages": 0, "sceneDirs": 0, "spines": 0, "images": 0, "bytes": 0, "versions": {}}
if not root.exists():
return [f"产物目录不存在:{root}"], total
candidates = []
for game_dir in sorted(p for p in root.iterdir() if p.is_dir()):
for page_dir in sorted(p for p in game_dir.iterdir() if p.is_dir()):
if pages and page_dir.name not in pages:
continue
candidates.append(page_dir)
if not candidates:
problems.append(f"没有任何页面产物:{root}")
for page_dir in candidates:
page_problems, stats = verify_page(page_dir)
problems.extend(f"[{page_dir.parent.name}/{page_dir.name}] {p}" for p in page_problems)
total["pages"] += 1
total["sceneDirs"] += stats["sceneDirs"]
total["spines"] += stats["spines"]
total["images"] += stats["images"]
total["bytes"] += stats["bytes"]
for version, count in stats["versions"].items():
total["versions"][version] = total["versions"].get(version, 0) + count
return problems, total