Turn scattered PDFs into a living Chinese research corpus: source files, extraction evidence, readable Obsidian notes, topic archives, and overview / academic / engineering surveys stay connected.
Paper Reading Workflow ZH packages a local paper-reading practice into reusable Codex Skills and Python helpers. It is designed for readers who do not want a disposable summary: every processed paper keeps a backend evidence workspace, a clean Chinese Obsidian note, and a path into a topic-level archive.
The workflow is intentionally split into two levels:
- Single paper: download or register a PDF, extract text and visuals, keep evidence artifacts, write a 14-section Chinese note, mark it as read.
- Topic corpus: fold same-type papers into a topic archive, maintain an inventory, and refresh overview, academic, and engineering surveys.
Private paths and local service addresses are kept out of the repository. Real locations live in environment variables or paper-notes.config.json, which is ignored by git.
| Layer | Output | Purpose |
|---|---|---|
| Paper inbox | PDFs and read markers | Keep source files and reading status explicit |
| Evidence workspace | MinerU JSON/Markdown, figures, manifests, reports | Preserve extraction and claim-level grounding |
| Formal note | 14-section Chinese Obsidian note | Make the paper readable and reviewable |
| Topic archive | Notes, PDFs, _evidence_reports, index, inventory |
Treat same-type papers as a corpus |
| Topic surveys | 全景综述, 学术综述, 工程综述 |
Turn a corpus into reusable synthesis |
paper title / URL / PDF
-> OpenCLI-assisted discovery
-> PDF inbox
-> MinerU extraction
-> evidence workspace
-> visual manifest
-> clean Chinese Obsidian note
-> read marker + reading index
-> topic archive
-> overview / academic / engineering surveys
-> recurring maintenance
All Skills live in skills/. The single-paper Skill is skills/evidence-grounded-paper-notes-zh/; the topic-family Skills sit next to it.
.
├── skills/
│ ├── evidence-grounded-paper-notes-zh/
│ ├── paper-knowledge-family-zh/
│ ├── paper-topic-archive-zh/
│ ├── paper-topic-survey-zh/
│ └── paper-topic-maintenance-zh/
├── paper_notes/
│ ├── config.py
│ ├── download.py
│ ├── mineru_extract.py
│ ├── figure_manifest.py
│ └── validate_notes.py
├── scripts/ # thin wrappers around paper_notes modules
├── tests/
├── examples/
│ ├── paper-notes.config.example.json
│ ├── read-marker.example.md
│ └── sample-vault/ # runnable demo vault for the validator
├── images/
├── NOTICE.md
└── README.zh-CN.md
- Python 3.10+
curl- OpenCLI, optional but recommended for paper discovery
- A local MinerU API for PDF extraction
- Poppler
pdftoppmandPillowonly when cropping visual evidence:
python -m pip install ".[visual]"Clone and install locally:
git clone https://github.com/LeoLin990405/paper-reading-workflow-zh.git
cd paper-reading-workflow-zh
python -m pip install -e .Installing exposes paper-notes-download, paper-notes-extract, paper-notes-figures, and paper-notes-validate as commands. The python scripts/*.py forms below also work straight from a clone without installing.
Try the validator on the bundled sample vault (no MinerU or OpenCLI needed):
paper-notes-validate --notes-dir "examples/sample-vault/注意力机制"Create a private config:
cp examples/paper-notes.config.example.json paper-notes.config.jsonDownload or register a paper:
python scripts/download_paper_opencli.py \
--query "Attention Is All You Need" \
--config paper-notes.config.jsonExtract PDFs with MinerU:
python scripts/mineru_local_batch_extract.py --config paper-notes.config.jsonBuild visual manifests:
python scripts/build_figure_manifest.py --config paper-notes.config.jsonValidate formal notes:
python scripts/validate_clean_notes.py --config paper-notes.config.jsonUse paper-notes.config.json, or set environment variables:
export PAPER_NOTES_PDF_INBOX="$HOME/Papers/inbox"
export PAPER_NOTES_READING_INBOX="$HOME/Documents/Obsidian/ResearchVault/论文阅读"
export PAPER_NOTES_ARCHIVE_ROOT="$HOME/Documents/Obsidian/ResearchVault/论文笔记"
export PAPER_NOTES_STAGING_ROOT="$HOME/Archives/paper-reading-workflow"
export MINERU_LOCAL_URL="http://127.0.0.1:8010"examples/paper-notes.config.example.json documents the same fields. Keep the real config local.
| Skill | Role |
|---|---|
evidence-grounded-paper-notes-zh |
Single-paper download, extraction, note writing, read markers |
paper-knowledge-family-zh |
Orchestrates the whole paper knowledge family |
paper-topic-archive-zh |
Folds same-type notes into topic archives |
paper-topic-survey-zh |
Maintains 全景综述, 学术综述, and 工程综述 |
paper-topic-maintenance-zh |
Keeps living corpora current after new papers arrive |
The core rule is:
single-paper note -> topic archive -> survey refresh -> recurring maintenance
Formal Obsidian notes use a fixed 14-section structure:
PDF 原文
已读信息
1. 一句话总览
2. 论文基本信息
3. 研究问题
4. 核心方法
5. 关键图表解读
6. 实验与结果
7. 公式与技术细节
8. 核心概念解释
9. 创新点与深度理解
10. 局限与批判
11. 和已读论文的关系
12. 对我的启发
Backend terms such as Evidence, Inference, MinerU, full_text, paper_index, and evidence_cards should stay out of the polished note body.
python -m pip install -e ".[dev]"
ruff check paper_notes scripts tests
pytestCI runs lint, unit tests, and a validation pass over examples/sample-vault/ on Python 3.10 and 3.12.
This project was iterated after studying MoonKirito/evidence-grounded-paper-deep-read, a paper-reading Skill released under the MIT License.
Thanks also to OpenCLI for adapter-style paper discovery and MinerU for document-to-Markdown/JSON extraction.
See NOTICE.md for attribution details.
MIT. See LICENSE.