PRODUCTS & SYSTEMS · 2026 · 12 VOLUMES / 5,817-PAGE SCOPE · CORPUS UNVERIFIED
5,817 was the planned input, not a completion count. I designed a workflow for checking OCR against the original pages of 12 Korean-medicine volumes and aligning Korean readings beside the classical Chinese text.
A start package, manifest, checkpoint and QA rules, and some partial outputs survive. I found no evidence that all 12 volumes were completed as a single corrected text-and-reading corpus. There is no external validation.
The workflow was designed to stop safely
A single final export would make it hard to locate when an error entered a large OCR run. At fixed page intervals, the plan preserved the page image reference, OCR state and reading alignment together. The manifest recorded the covered range so another pass could resume from a known point.
OCR and reading annotation fail differently. The checks separated a wrongly recognised character, a correct character with a wrong reading, and a shifted line that broke the relationship between the two. Volume order and integrity values were meant to be checked before final assembly.
A start package was not the corpus
The available record does not establish a completed-page count, sample accuracy, error rate or final checksum. That is why the 5,817-page scope is not reported as processed output. A completion audit would have to recover the volume-level artifacts and page-comparison records.
Redistribution rights for the books, PDFs and OCR text have not been resolved item by item. This note therefore omits titles, authors, page text and reading examples, and describes only the workflow.
Project record
Period: 2026
Role: page-comparison OCR, reading alignment, checkpoints, QA and integrity workflow
Internal evidence: start package, manifest, partial outputs and working rules
External validation: none
Final outcome: completion of the full 12-volume, 5,817-page corpus is unverified
