Files
tepbc-ross-retyped/CLAUDE.md
T
bdeshiandClaude Opus 5 82849eb94d Initial commit: Ross 1988 re-typeset edition
A searchable, re-typeset edition of Fiona Ross, "The Evolution of the
Printed Bengali Character from 1778 to 1978" (Ph.D., SOAS, 1988),
transcribed from the 431-leaf ProQuest scan. All 431 pages done; 178
plates and 410 inline type specimens cut from the scan; 51 errata.

Tracked: the transcription (src/pages), the preamble and its typographic
decisions, the cut images (plates/ — not reliably regenerable, the crop
specs for the inline cuts were never scripted), tools, and the four
working documents.

Not tracked: the built PDF, which `make` remakes from src/ and plates/;
the ProQuest scan under source/, which is third-party and needed only by
`make prep` and `make plate`; scans/ and work/, both regenerable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 19:03:25 +06:00

150 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Ross 1988 re-typeset — working instructions for Claude Code
Goal: convert `source/10731406.pdf` (scan of Fiona Ross, *The Evolution of the Printed
Bengali Character from 1778 to 1978*, Ph.D., SOAS 1988; 431 PDF pages) into a searchable,
re-typeset PDF. Design decisions are final and recorded in `PLAN.md`; current state and the
next batch are in `CONTINUATION.md`; open readings for the owner (Sammay) in `QUESTIONS.md`.
Read those three files first, every session.
## Hard rules
1. **The scan's text layer is not a source.** Transcribe from the page image
(`scans/hi-NNN.jpg`, 130 dpi; re-rasterize at 300 dpi for anything small). The
tesseract text in `work/ocr/` is only a typing scaffold — proofread every word against
the image before it goes in a `.tex` file. Never `pdftotext` the source for content.
2. **Never guess.** A word, diacritic, number or glyph you cannot read with confidence is
written as `\unsure{best reading}` and added to `QUESTIONS.md` with PDF page, original
page and a description. Do not silently normalise or "correct" the author's text.
3. **Original spelling and typography stand**, including the author's inconsistencies.
Handwritten corrections in the scanned copy are adopted and logged in `QUESTIONS.md`.
4. **One file per original page**: `src/pages/pNNNN.tex`, NNNN = zero-padded **PDF** page.
Row in `manifest.tsv` updated in the same step (printed page, kind, status=done, src_file, notes).
5. **Bengali: modern script is text, historical typefaces are images.** Bengali that shows
the modern script — the Scheme of Transliteration on orig. p.13 — is typed as Unicode and
sets itself in Tiro Bangla. Bengali that shows a **historical typeface** (every specimen the
original pastes into its own sentences, the List-of-Plates glyphs, plate captions) is never
typed: cut it from the scan with `tools/crop_inline.py` and set it with `\ig`. The argument in
the text is about those shapes — retyping them destroys the evidence. (Sammay, 2026-09-13.)
6. **Plates keep native resolution**: crop with `make plate N=…`; never resample — the sole
exception is the automatic deskew, which rotates (and so resamples) a page that is measurably
out of square. The tool also sweeps scanner bars and the grey gutter shadow from the edges.
Tune per page
with `ARGS`: `--top/--bottom/--edge` (strip page number, caption, scanner edge), `--rotate 270`
for plates the original prints sideways (lossless), `--trim T,R,B,L` after rotation for skewed
scanner edges the heuristics cannot see; `--box L,T,R,B` gives the crop explicitly and skips the
bar/page-number heuristics altogether — needed for dark plates (photographed manuscripts,
photographic plates), where “dark means scanner bar” is false; `--raw` (with `--box`) takes the
box **exactly**, skipping the shadow sweep, deskew, bar sweep and autocrop — needed for a plate
that reproduces a whole **framed** page, whose own border rules the edge sweeps mistake for
scanner bars and strip along with content. `--box` values are fractions of the raw page, never
pixels. Record the recipe in the manifest `notes` column.
7. After each batch: `make check` must pass, then `make build` must succeed, and
`python3 tools/reprocheck.py` must print `0 tokens short` — it compares every token of
`src/pages/*.tex` against the built PDF and is the only check that catches content the
*typesetter* loses (see CONTINUATION.md for the `\hangafter` case it was written for). Do not leave
the tree in a state where either fails. `make verify` runs all three.
8. Do not edit `src/main.tex` (generated by `build.sh`). **Git is the version history**
(initialised 2026-09-14, Sammay): the build no longer archives the previous PDF, the PDF
is not tracked — `make` remakes it — and the 68 archived copies in `older-versions/` were
deleted. Superseded design drafts stay in `older-versions/` as text and are tracked.
9. Ask before changing anything in `PLAN.md`'s design table.
## Macro reference (`src/preamble.tex`)
| Macro | Use |
|---|---|
| `\chapstart{Title}` / `\chaphead{Title}` | Starts a chapter/section: new page, centred bold head, bookmark, note group "Notes to Title". **Must precede `\origpage` in the file.** Resets note numbering: next `\fn` is `\fn{1}`. |
| `\origpage{N}` | Original page N starts here (mid-sentence is normal). Exactly one per text page file, at the point of the break. |
| `\fn{n}{text}` | Footnote n of the original → endnote. n must be the original's number; `make check` enforces 1,2,3… per chapter. |
| `\plateop{N}{no}{plates/pNNNN.png}{caption}` | A plate that begins original page N (the usual case: one plate per page). |
| `\plate{no}{img}{caption}` | Plate not at a page start. |
| `\subhead{text}` | Italic run-in subheading (as the original's *Introduction:* etc.). |
| `\chapnum{Chapter 1}{Charles Wilkins}` | Chapter head as the original sets it: number over title, both centred bold; notes group as “Chapter 1: Charles Wilkins”. Precedes `\origpage`. |
| `\begin{extract}…\end{extract}` | A displayed quotation as the original sets them — indented both sides, set solid, no quotation marks. An environment, so a quotation broken by an original page break opens in one page file and closes in the next (`\origpage` then sits inside it). |
| `\ig{img}` | An inline type specimen cut from the scan by `tools/crop_inline.py`, set at native size (1 px = 1/300 in). Use wherever the original pastes typeforms into its own text — never retype those as Tiro Bangla, the argument is about their shapes. |
| `\pg{N}` | Hyperlink to original page N (used in Contents/List of Plates). |
| `\begin{biblist}…\end{biblist}`, `\bibgroup{…}`, `\bibhead{…}` | The Bibliography (orig. pp.419–429): the original sets it solid — a bold group heading (`\bibgroup`), then one entry per paragraph with the continuation lines indented. `biblist` gives that block (no paragraph skip, 1.6em hanging indent); `\bibhead` is a centred bold heading inside it. |
| `\fig{img}` | An **unnumbered** illustration the original sets in the run of the text (no plate number, no caption — the preceding sentence carries the reference). Centred, native size, shrunk only if wider than the measure. Crop it by explicit pixel box with PIL to `plates/fig-pNNNN.png` and record the box in the manifest `notes` column. |
| `\pl{no}{caption}{page}`, `\tocl{indent}{label}{title}{page}` | List-of-Plates / Contents lines. |
| `\unsure{text}` | Doubtful reading (logged to `unsure.log`). |
| `\qslip{as printed}` | A suspected slip **inside quoted matter**. Never corrected — the slip may be the quoted source's, not Ross's. Sets the text unchanged, logs to `qslips.log`; list it in `QUESTIONS.md` with the source so it can be checked. |
| `\erratum{corrected}{as printed}` | A slip in the **original** that this edition corrects: sets the corrected reading, logs to `errata.log`. Arguments must be plain text (put `\emph` outside). Every use needs a matching `\erratumline{orig page}{as printed}{corrected}` in `src/errata.tex` — `make check` enforces it — and the pair prints in the Errata section. Handwritten corrections in the scan are *not* errata: adopt them silently and log in `QUESTIONS.md`. |
| `\partstart{Part I}{Title}` | Part-title page: label and title centred about two fifths down, bookmark, new note group. Precedes `\origpage` like `\chapstart`. |
| `\sectionstart{Section A}{Title}` | Section-title page — flush left, as the original sets them. |
| `\emph{}` for italics, `` ` '' `` for quotes, `\ldots\ ` for …, `\&`, `\%`. | Typewriter conventions of the original (no ligatures etc.) need no special handling. |
## Typography (set once in `src/preamble.tex` — do not re-litigate per page)
12pt **TeX Gyre ScholaX** (Century Schoolbook; loaded by path from the TeX tree, which fontconfig does
not index) on a 5.95in measure; Liberation Sans for margin marks and running
heads; Tiro Bangla for Bengali. **Leading 1.59 — a 22.9pt line, the original's own, measured off the scan
at 300 dpi; the inline specimens set it, not the text** (see CONTINUATION.md). Inline specimens at native
size (`\igscale` = 72.27/300), which is the 74% of the line the original gives them. First line indented
1.6em with `\parskip` 0.15em, `\frenchspacing`
(the original is a typescript with uniform spacing, and TeX's sentence spacing after `p.` is wrong here).
**Ragged right and unhyphenated** (`\RaggedRight`; `\lefthyphenmin=62` plus `\hyphenpenalty=10000` — the original typescript breaks only at spaces and at hyphens it actually typed, so `\exhyphenpenalty` stays low and `\tolerance`/`\emergencystretch` are raised to pay for the lost breaks). (Sammay, 2026-09-14.) **Hanging punctuation** via microtype protrusion,
dashes hanging far less. **Widows and orphans forbidden** with `\raggedbottom`. Running head: current
chapter (a `chap` mark class, set by every heading macro) on the left, original page range on the right.
Endnotes set a step smaller and tighter via `\notesection`; plate captions are `\footnotesize` with their own leading, a clear step below the text as in the original. Contents and List-of-Plates entries are
whole-line links to `page.N`.
## `tools/polish.py` — the build-time typographic pass
`build.sh` runs it first: it reads `src/pages/*.tex` and writes `work/pages/*.tex`, and the build inputs
the latter. **The transcription in `src/pages` always stays exactly as typed** — never hand-edit the
polished copies. The pass applies what the typewriter could not (Sammay, 2026-09-13):
* ties, so a reference or unit never breaks across a line (`p.~63`, `Part~II`, `240~lbs`)
* en-dashes for numeric ranges (1778--1978)
* small caps for institutional acronyms — an explicit whitelist in the script (IOL, BMS, OUP, SOAS,
MS, BL …). Fount codes (CW1, SB4, BM V, VF1) and Roman numerals are deliberately excluded
* `\mbox` round every inline specimen so it never breaks from adjacent punctuation
* **cross-reference links**: `pl. N` and `chapter N` always link (they can only be this thesis's own);
`p. N`/`pp. N` link **only where the thesis cites itself**. A page reference inside a citation
belongs to somebody else's book: the script tests the preceding clause for citation markers (a
title in `\emph`, `ibid.`, a shelfmark, publication data, a surname followed by a comma) before
accepting self-reference words (`above`, `below`, `see`, `mentioned`). Neither signal → no link; a
missing link is cheaper than one that lands on the wrong page. **Re-audit after changing this rule**:
print every link with its context and read them.
Protected from all of the above: comment lines, file paths, and the arguments of `\erratum`, `\qslip`,
`\ig`, `\origpage`, `\fn`, and the heading macros (their text becomes bookmarks and marks, and a link
inside them breaks the build).
## Inline type specimens — the workflow
For any page where the original pastes typeforms into its own text:
1. `docker compose run --rm -T tex python3 work/lines.py N` — lists the page's text lines with their
y-fractions. **Its line N is usually the line *after* the one you counted by eye; if a sheet comes
back with the wrong line, try the band one earlier.**
2. `… work/glyphs.py N y0,y1 [y0,y1 …]` — numbered contact sheet of the ink runs on those bands,
written to `work/sheetN.png`. Read it and pick the runs that are specimens.
3. `… tools/crop_inline.py N name=SPEC …` — cuts them to `plates/inline/pNNNN-name.png`.
SPEC is a run index plus, optionally, `:left|right|first|last|mid|trimright|trimleft` to split a
run that merged with its neighbour, `+PX` to start PX into the run, `~PX` to keep only its first
PX. Use `+PX`/`~PX` when the typewriter's space is too noisy to register as a gap.
4. Where a run defeats all of that (a word set solid inside a footnote), crop by explicit pixel box
with PIL and **record the box in the manifest `notes` column** so the cut is reproducible.
## Per-batch procedure (≈12–15 pages per batch)
1. `make prep A=<pdf> B=<pdf>` → `scans/hi-*.jpg` + `work/ocr/p*.txt`.
2. For each page, **look at the image**. Decide kind: `text`, `plate`, `mixed`, `blank`.
3. Text page: `python3 tools/ocr2tex.py <pdf> <printed> [--head "Chapter title"]` gives a
draft; then proofread against the image: italics (titles, transliterated words), `\fn{}`
marks placed exactly where the superscript sits, footnote text moved into `\fn{}`, all
diacritics (ā ī ū ṛ ṃ ṅ ñ ṭ ḍ ṇ ś ṣ ḥ), hyphenation joined, en-dashes kept as `-` as the
original types them. **The typescript never hyphenates a word at a line end** — verified across
all 431 leaves: every one of its 59 line-ending hyphens is either a compound the author typed
(`offset-lithography`, `Hof-und`, `forty-eight`) or OCR noise off a reproduced plate. So a hyphen
in a page file is always intrinsic to the word; never invent one, and never keep one that splits a
word. Output hyphenation is off to match (Sammay, 2026-09-14). Where the author compounds
inconsistently (`type-casting`/`typecasting`, `punch-cutter`/`punchcutter`) rule 3 applies: each
instance stands as printed. Page break inside a paragraph: the previous file ends mid-sentence
and this file begins `\origpage{N}` followed by the continuation — no blank line between
files (the build concatenates them), so do **not** start the file with `\par`.
4. Plate page: `make plate N=<pdf>` (inspect `plates/pNNNN.png`; if the caption or page
number survived or content was clipped, rerun with `ARGS="--top 0.05 --bottom 0.11"`
etc.), then a one-line file with `\plateop{...}` and the caption transcribed exactly.
5. Update `manifest.tsv` rows (`python3 tools/setrow.py <pdf> <printed> <kind> "notes"`); `make check`; `make build`; spot-check 2–3 output pages
(`pdftoppm -jpeg -r 60 -f X -l Y ross-1988-retypeset.pdf work/o`).
6. Update `CONTINUATION.md` (state, next batch) and `QUESTIONS.md`. Stop at a chapter or
batch boundary, never mid-file.
## Environment
Tools run in Docker: `make image` once, then the `make` targets above (`NODOCKER=1 make …`
if xelatex, poppler-utils, tesseract, python3-pil/numpy/pdfplumber are installed locally).
Fonts: Liberation (system), FreeSerif (system), Tiro Bangla (`fonts/`). Do not add fonts.