A searchable, re-typeset edition of Fiona Ross, "The Evolution of the Printed Bengali Character from 1778 to 1978" (Ph.D., SOAS, 1988), transcribed from the 431-leaf ProQuest scan. All 431 pages done; 178 plates and 410 inline type specimens cut from the scan; 51 errata. Tracked: the transcription (src/pages), the preamble and its typographic decisions, the cut images (plates/ — not reliably regenerable, the crop specs for the inline cuts were never scripted), tools, and the four working documents. Not tracked: the built PDF, which `make` remakes from src/ and plates/; the ProQuest scan under source/, which is third-party and needed only by `make prep` and `make plate`; scans/ and work/, both regenerable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
150 lines
14 KiB
Markdown
150 lines
14 KiB
Markdown
# Ross 1988 re-typeset — working instructions for Claude Code
|
||
|
||
Goal: convert `source/10731406.pdf` (scan of Fiona Ross, *The Evolution of the Printed
|
||
Bengali Character from 1778 to 1978*, Ph.D., SOAS 1988; 431 PDF pages) into a searchable,
|
||
re-typeset PDF. Design decisions are final and recorded in `PLAN.md`; current state and the
|
||
next batch are in `CONTINUATION.md`; open readings for the owner (Sammay) in `QUESTIONS.md`.
|
||
Read those three files first, every session.
|
||
|
||
## Hard rules
|
||
1. **The scan's text layer is not a source.** Transcribe from the page image
|
||
(`scans/hi-NNN.jpg`, 130 dpi; re-rasterize at 300 dpi for anything small). The
|
||
tesseract text in `work/ocr/` is only a typing scaffold — proofread every word against
|
||
the image before it goes in a `.tex` file. Never `pdftotext` the source for content.
|
||
2. **Never guess.** A word, diacritic, number or glyph you cannot read with confidence is
|
||
written as `\unsure{best reading}` and added to `QUESTIONS.md` with PDF page, original
|
||
page and a description. Do not silently normalise or "correct" the author's text.
|
||
3. **Original spelling and typography stand**, including the author's inconsistencies.
|
||
Handwritten corrections in the scanned copy are adopted and logged in `QUESTIONS.md`.
|
||
4. **One file per original page**: `src/pages/pNNNN.tex`, NNNN = zero-padded **PDF** page.
|
||
Row in `manifest.tsv` updated in the same step (printed page, kind, status=done, src_file, notes).
|
||
5. **Bengali: modern script is text, historical typefaces are images.** Bengali that shows
|
||
the modern script — the Scheme of Transliteration on orig. p.13 — is typed as Unicode and
|
||
sets itself in Tiro Bangla. Bengali that shows a **historical typeface** (every specimen the
|
||
original pastes into its own sentences, the List-of-Plates glyphs, plate captions) is never
|
||
typed: cut it from the scan with `tools/crop_inline.py` and set it with `\ig`. The argument in
|
||
the text is about those shapes — retyping them destroys the evidence. (Sammay, 2026-09-13.)
|
||
6. **Plates keep native resolution**: crop with `make plate N=…`; never resample — the sole
|
||
exception is the automatic deskew, which rotates (and so resamples) a page that is measurably
|
||
out of square. The tool also sweeps scanner bars and the grey gutter shadow from the edges.
|
||
Tune per page
|
||
with `ARGS`: `--top/--bottom/--edge` (strip page number, caption, scanner edge), `--rotate 270`
|
||
for plates the original prints sideways (lossless), `--trim T,R,B,L` after rotation for skewed
|
||
scanner edges the heuristics cannot see; `--box L,T,R,B` gives the crop explicitly and skips the
|
||
bar/page-number heuristics altogether — needed for dark plates (photographed manuscripts,
|
||
photographic plates), where “dark means scanner bar” is false; `--raw` (with `--box`) takes the
|
||
box **exactly**, skipping the shadow sweep, deskew, bar sweep and autocrop — needed for a plate
|
||
that reproduces a whole **framed** page, whose own border rules the edge sweeps mistake for
|
||
scanner bars and strip along with content. `--box` values are fractions of the raw page, never
|
||
pixels. Record the recipe in the manifest `notes` column.
|
||
7. After each batch: `make check` must pass, then `make build` must succeed, and
|
||
`python3 tools/reprocheck.py` must print `0 tokens short` — it compares every token of
|
||
`src/pages/*.tex` against the built PDF and is the only check that catches content the
|
||
*typesetter* loses (see CONTINUATION.md for the `\hangafter` case it was written for). Do not leave
|
||
the tree in a state where either fails. `make verify` runs all three.
|
||
8. Do not edit `src/main.tex` (generated by `build.sh`). **Git is the version history**
|
||
(initialised 2026-09-14, Sammay): the build no longer archives the previous PDF, the PDF
|
||
is not tracked — `make` remakes it — and the 68 archived copies in `older-versions/` were
|
||
deleted. Superseded design drafts stay in `older-versions/` as text and are tracked.
|
||
9. Ask before changing anything in `PLAN.md`'s design table.
|
||
|
||
## Macro reference (`src/preamble.tex`)
|
||
| Macro | Use |
|
||
|---|---|
|
||
| `\chapstart{Title}` / `\chaphead{Title}` | Starts a chapter/section: new page, centred bold head, bookmark, note group "Notes to Title". **Must precede `\origpage` in the file.** Resets note numbering: next `\fn` is `\fn{1}`. |
|
||
| `\origpage{N}` | Original page N starts here (mid-sentence is normal). Exactly one per text page file, at the point of the break. |
|
||
| `\fn{n}{text}` | Footnote n of the original → endnote. n must be the original's number; `make check` enforces 1,2,3… per chapter. |
|
||
| `\plateop{N}{no}{plates/pNNNN.png}{caption}` | A plate that begins original page N (the usual case: one plate per page). |
|
||
| `\plate{no}{img}{caption}` | Plate not at a page start. |
|
||
| `\subhead{text}` | Italic run-in subheading (as the original's *Introduction:* etc.). |
|
||
| `\chapnum{Chapter 1}{Charles Wilkins}` | Chapter head as the original sets it: number over title, both centred bold; notes group as “Chapter 1: Charles Wilkins”. Precedes `\origpage`. |
|
||
| `\begin{extract}…\end{extract}` | A displayed quotation as the original sets them — indented both sides, set solid, no quotation marks. An environment, so a quotation broken by an original page break opens in one page file and closes in the next (`\origpage` then sits inside it). |
|
||
| `\ig{img}` | An inline type specimen cut from the scan by `tools/crop_inline.py`, set at native size (1 px = 1/300 in). Use wherever the original pastes typeforms into its own text — never retype those as Tiro Bangla, the argument is about their shapes. |
|
||
| `\pg{N}` | Hyperlink to original page N (used in Contents/List of Plates). |
|
||
| `\begin{biblist}…\end{biblist}`, `\bibgroup{…}`, `\bibhead{…}` | The Bibliography (orig. pp.419–429): the original sets it solid — a bold group heading (`\bibgroup`), then one entry per paragraph with the continuation lines indented. `biblist` gives that block (no paragraph skip, 1.6em hanging indent); `\bibhead` is a centred bold heading inside it. |
|
||
| `\fig{img}` | An **unnumbered** illustration the original sets in the run of the text (no plate number, no caption — the preceding sentence carries the reference). Centred, native size, shrunk only if wider than the measure. Crop it by explicit pixel box with PIL to `plates/fig-pNNNN.png` and record the box in the manifest `notes` column. |
|
||
| `\pl{no}{caption}{page}`, `\tocl{indent}{label}{title}{page}` | List-of-Plates / Contents lines. |
|
||
| `\unsure{text}` | Doubtful reading (logged to `unsure.log`). |
|
||
| `\qslip{as printed}` | A suspected slip **inside quoted matter**. Never corrected — the slip may be the quoted source's, not Ross's. Sets the text unchanged, logs to `qslips.log`; list it in `QUESTIONS.md` with the source so it can be checked. |
|
||
| `\erratum{corrected}{as printed}` | A slip in the **original** that this edition corrects: sets the corrected reading, logs to `errata.log`. Arguments must be plain text (put `\emph` outside). Every use needs a matching `\erratumline{orig page}{as printed}{corrected}` in `src/errata.tex` — `make check` enforces it — and the pair prints in the Errata section. Handwritten corrections in the scan are *not* errata: adopt them silently and log in `QUESTIONS.md`. |
|
||
| `\partstart{Part I}{Title}` | Part-title page: label and title centred about two fifths down, bookmark, new note group. Precedes `\origpage` like `\chapstart`. |
|
||
| `\sectionstart{Section A}{Title}` | Section-title page — flush left, as the original sets them. |
|
||
| `\emph{}` for italics, `` ` '' `` for quotes, `\ldots\ ` for …, `\&`, `\%`. | Typewriter conventions of the original (no ligatures etc.) need no special handling. |
|
||
|
||
## Typography (set once in `src/preamble.tex` — do not re-litigate per page)
|
||
12pt **TeX Gyre ScholaX** (Century Schoolbook; loaded by path from the TeX tree, which fontconfig does
|
||
not index) on a 5.95in measure; Liberation Sans for margin marks and running
|
||
heads; Tiro Bangla for Bengali. **Leading 1.59 — a 22.9pt line, the original's own, measured off the scan
|
||
at 300 dpi; the inline specimens set it, not the text** (see CONTINUATION.md). Inline specimens at native
|
||
size (`\igscale` = 72.27/300), which is the 74% of the line the original gives them. First line indented
|
||
1.6em with `\parskip` 0.15em, `\frenchspacing`
|
||
(the original is a typescript with uniform spacing, and TeX's sentence spacing after `p.` is wrong here).
|
||
**Ragged right and unhyphenated** (`\RaggedRight`; `\lefthyphenmin=62` plus `\hyphenpenalty=10000` — the original typescript breaks only at spaces and at hyphens it actually typed, so `\exhyphenpenalty` stays low and `\tolerance`/`\emergencystretch` are raised to pay for the lost breaks). (Sammay, 2026-09-14.) **Hanging punctuation** via microtype protrusion,
|
||
dashes hanging far less. **Widows and orphans forbidden** with `\raggedbottom`. Running head: current
|
||
chapter (a `chap` mark class, set by every heading macro) on the left, original page range on the right.
|
||
Endnotes set a step smaller and tighter via `\notesection`; plate captions are `\footnotesize` with their own leading, a clear step below the text as in the original. Contents and List-of-Plates entries are
|
||
whole-line links to `page.N`.
|
||
|
||
## `tools/polish.py` — the build-time typographic pass
|
||
`build.sh` runs it first: it reads `src/pages/*.tex` and writes `work/pages/*.tex`, and the build inputs
|
||
the latter. **The transcription in `src/pages` always stays exactly as typed** — never hand-edit the
|
||
polished copies. The pass applies what the typewriter could not (Sammay, 2026-09-13):
|
||
* ties, so a reference or unit never breaks across a line (`p.~63`, `Part~II`, `240~lbs`)
|
||
* en-dashes for numeric ranges (1778--1978)
|
||
* small caps for institutional acronyms — an explicit whitelist in the script (IOL, BMS, OUP, SOAS,
|
||
MS, BL …). Fount codes (CW1, SB4, BM V, VF1) and Roman numerals are deliberately excluded
|
||
* `\mbox` round every inline specimen so it never breaks from adjacent punctuation
|
||
* **cross-reference links**: `pl. N` and `chapter N` always link (they can only be this thesis's own);
|
||
`p. N`/`pp. N` link **only where the thesis cites itself**. A page reference inside a citation
|
||
belongs to somebody else's book: the script tests the preceding clause for citation markers (a
|
||
title in `\emph`, `ibid.`, a shelfmark, publication data, a surname followed by a comma) before
|
||
accepting self-reference words (`above`, `below`, `see`, `mentioned`). Neither signal → no link; a
|
||
missing link is cheaper than one that lands on the wrong page. **Re-audit after changing this rule**:
|
||
print every link with its context and read them.
|
||
Protected from all of the above: comment lines, file paths, and the arguments of `\erratum`, `\qslip`,
|
||
`\ig`, `\origpage`, `\fn`, and the heading macros (their text becomes bookmarks and marks, and a link
|
||
inside them breaks the build).
|
||
|
||
## Inline type specimens — the workflow
|
||
For any page where the original pastes typeforms into its own text:
|
||
1. `docker compose run --rm -T tex python3 work/lines.py N` — lists the page's text lines with their
|
||
y-fractions. **Its line N is usually the line *after* the one you counted by eye; if a sheet comes
|
||
back with the wrong line, try the band one earlier.**
|
||
2. `… work/glyphs.py N y0,y1 [y0,y1 …]` — numbered contact sheet of the ink runs on those bands,
|
||
written to `work/sheetN.png`. Read it and pick the runs that are specimens.
|
||
3. `… tools/crop_inline.py N name=SPEC …` — cuts them to `plates/inline/pNNNN-name.png`.
|
||
SPEC is a run index plus, optionally, `:left|right|first|last|mid|trimright|trimleft` to split a
|
||
run that merged with its neighbour, `+PX` to start PX into the run, `~PX` to keep only its first
|
||
PX. Use `+PX`/`~PX` when the typewriter's space is too noisy to register as a gap.
|
||
4. Where a run defeats all of that (a word set solid inside a footnote), crop by explicit pixel box
|
||
with PIL and **record the box in the manifest `notes` column** so the cut is reproducible.
|
||
|
||
## Per-batch procedure (≈12–15 pages per batch)
|
||
1. `make prep A=<pdf> B=<pdf>` → `scans/hi-*.jpg` + `work/ocr/p*.txt`.
|
||
2. For each page, **look at the image**. Decide kind: `text`, `plate`, `mixed`, `blank`.
|
||
3. Text page: `python3 tools/ocr2tex.py <pdf> <printed> [--head "Chapter title"]` gives a
|
||
draft; then proofread against the image: italics (titles, transliterated words), `\fn{}`
|
||
marks placed exactly where the superscript sits, footnote text moved into `\fn{}`, all
|
||
diacritics (ā ī ū ṛ ṃ ṅ ñ ṭ ḍ ṇ ś ṣ ḥ), hyphenation joined, en-dashes kept as `-` as the
|
||
original types them. **The typescript never hyphenates a word at a line end** — verified across
|
||
all 431 leaves: every one of its 59 line-ending hyphens is either a compound the author typed
|
||
(`offset-lithography`, `Hof-und`, `forty-eight`) or OCR noise off a reproduced plate. So a hyphen
|
||
in a page file is always intrinsic to the word; never invent one, and never keep one that splits a
|
||
word. Output hyphenation is off to match (Sammay, 2026-09-14). Where the author compounds
|
||
inconsistently (`type-casting`/`typecasting`, `punch-cutter`/`punchcutter`) rule 3 applies: each
|
||
instance stands as printed. Page break inside a paragraph: the previous file ends mid-sentence
|
||
and this file begins `\origpage{N}` followed by the continuation — no blank line between
|
||
files (the build concatenates them), so do **not** start the file with `\par`.
|
||
4. Plate page: `make plate N=<pdf>` (inspect `plates/pNNNN.png`; if the caption or page
|
||
number survived or content was clipped, rerun with `ARGS="--top 0.05 --bottom 0.11"`
|
||
etc.), then a one-line file with `\plateop{...}` and the caption transcribed exactly.
|
||
5. Update `manifest.tsv` rows (`python3 tools/setrow.py <pdf> <printed> <kind> "notes"`); `make check`; `make build`; spot-check 2–3 output pages
|
||
(`pdftoppm -jpeg -r 60 -f X -l Y ross-1988-retypeset.pdf work/o`).
|
||
6. Update `CONTINUATION.md` (state, next batch) and `QUESTIONS.md`. Stop at a chapter or
|
||
batch boundary, never mid-file.
|
||
|
||
## Environment
|
||
Tools run in Docker: `make image` once, then the `make` targets above (`NODOCKER=1 make …`
|
||
if xelatex, poppler-utils, tesseract, python3-pil/numpy/pdfplumber are installed locally).
|
||
Fonts: Liberation (system), FreeSerif (system), Tiro Bangla (`fonts/`). Do not add fonts.
|