Files
tepbc-ross-retyped/CLAUDE.md
T
bdeshiandClaude Opus 5 82849eb94d Initial commit: Ross 1988 re-typeset edition
A searchable, re-typeset edition of Fiona Ross, "The Evolution of the
Printed Bengali Character from 1778 to 1978" (Ph.D., SOAS, 1988),
transcribed from the 431-leaf ProQuest scan. All 431 pages done; 178
plates and 410 inline type specimens cut from the scan; 51 errata.

Tracked: the transcription (src/pages), the preamble and its typographic
decisions, the cut images (plates/ — not reliably regenerable, the crop
specs for the inline cuts were never scripted), tools, and the four
working documents.

Not tracked: the built PDF, which `make` remakes from src/ and plates/;
the ProQuest scan under source/, which is third-party and needed only by
`make prep` and `make plate`; scans/ and work/, both regenerable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 19:03:25 +06:00

14 KiB
Raw Blame History

Ross 1988 re-typeset — working instructions for Claude Code

Goal: convert source/10731406.pdf (scan of Fiona Ross, The Evolution of the Printed Bengali Character from 1778 to 1978, Ph.D., SOAS 1988; 431 PDF pages) into a searchable, re-typeset PDF. Design decisions are final and recorded in PLAN.md; current state and the next batch are in CONTINUATION.md; open readings for the owner (Sammay) in QUESTIONS.md. Read those three files first, every session.

Hard rules

  1. The scan's text layer is not a source. Transcribe from the page image (scans/hi-NNN.jpg, 130 dpi; re-rasterize at 300 dpi for anything small). The tesseract text in work/ocr/ is only a typing scaffold — proofread every word against the image before it goes in a .tex file. Never pdftotext the source for content.
  2. Never guess. A word, diacritic, number or glyph you cannot read with confidence is written as \unsure{best reading} and added to QUESTIONS.md with PDF page, original page and a description. Do not silently normalise or "correct" the author's text.
  3. Original spelling and typography stand, including the author's inconsistencies. Handwritten corrections in the scanned copy are adopted and logged in QUESTIONS.md.
  4. One file per original page: src/pages/pNNNN.tex, NNNN = zero-padded PDF page. Row in manifest.tsv updated in the same step (printed page, kind, status=done, src_file, notes).
  5. Bengali: modern script is text, historical typefaces are images. Bengali that shows the modern script — the Scheme of Transliteration on orig. p.13 — is typed as Unicode and sets itself in Tiro Bangla. Bengali that shows a historical typeface (every specimen the original pastes into its own sentences, the List-of-Plates glyphs, plate captions) is never typed: cut it from the scan with tools/crop_inline.py and set it with \ig. The argument in the text is about those shapes — retyping them destroys the evidence. (Sammay, 2026-09-13.)
  6. Plates keep native resolution: crop with make plate N=…; never resample — the sole exception is the automatic deskew, which rotates (and so resamples) a page that is measurably out of square. The tool also sweeps scanner bars and the grey gutter shadow from the edges. Tune per page with ARGS: --top/--bottom/--edge (strip page number, caption, scanner edge), --rotate 270 for plates the original prints sideways (lossless), --trim T,R,B,L after rotation for skewed scanner edges the heuristics cannot see; --box L,T,R,B gives the crop explicitly and skips the bar/page-number heuristics altogether — needed for dark plates (photographed manuscripts, photographic plates), where “dark means scanner bar” is false; --raw (with --box) takes the box exactly, skipping the shadow sweep, deskew, bar sweep and autocrop — needed for a plate that reproduces a whole framed page, whose own border rules the edge sweeps mistake for scanner bars and strip along with content. --box values are fractions of the raw page, never pixels. Record the recipe in the manifest notes column.
  7. After each batch: make check must pass, then make build must succeed, and python3 tools/reprocheck.py must print 0 tokens short — it compares every token of src/pages/*.tex against the built PDF and is the only check that catches content the typesetter loses (see CONTINUATION.md for the \hangafter case it was written for). Do not leave the tree in a state where either fails. make verify runs all three.
  8. Do not edit src/main.tex (generated by build.sh). Git is the version history (initialised 2026-09-14, Sammay): the build no longer archives the previous PDF, the PDF is not tracked — make remakes it — and the 68 archived copies in older-versions/ were deleted. Superseded design drafts stay in older-versions/ as text and are tracked.
  9. Ask before changing anything in PLAN.md's design table.

Macro reference (src/preamble.tex)

Macro Use
\chapstart{Title} / \chaphead{Title} Starts a chapter/section: new page, centred bold head, bookmark, note group "Notes to Title". Must precede \origpage in the file. Resets note numbering: next \fn is \fn{1}.
\origpage{N} Original page N starts here (mid-sentence is normal). Exactly one per text page file, at the point of the break.
\fn{n}{text} Footnote n of the original → endnote. n must be the original's number; make check enforces 1,2,3… per chapter.
\plateop{N}{no}{plates/pNNNN.png}{caption} A plate that begins original page N (the usual case: one plate per page).
\plate{no}{img}{caption} Plate not at a page start.
\subhead{text} Italic run-in subheading (as the original's Introduction: etc.).
\chapnum{Chapter 1}{Charles Wilkins} Chapter head as the original sets it: number over title, both centred bold; notes group as “Chapter 1: Charles Wilkins”. Precedes \origpage.
\begin{extract}…\end{extract} A displayed quotation as the original sets them — indented both sides, set solid, no quotation marks. An environment, so a quotation broken by an original page break opens in one page file and closes in the next (\origpage then sits inside it).
\ig{img} An inline type specimen cut from the scan by tools/crop_inline.py, set at native size (1 px = 1/300 in). Use wherever the original pastes typeforms into its own text — never retype those as Tiro Bangla, the argument is about their shapes.
\pg{N} Hyperlink to original page N (used in Contents/List of Plates).
\begin{biblist}…\end{biblist}, \bibgroup{…}, \bibhead{…} The Bibliography (orig. pp.419–429): the original sets it solid — a bold group heading (\bibgroup), then one entry per paragraph with the continuation lines indented. biblist gives that block (no paragraph skip, 1.6em hanging indent); \bibhead is a centred bold heading inside it.
\fig{img} An unnumbered illustration the original sets in the run of the text (no plate number, no caption — the preceding sentence carries the reference). Centred, native size, shrunk only if wider than the measure. Crop it by explicit pixel box with PIL to plates/fig-pNNNN.png and record the box in the manifest notes column.
\pl{no}{caption}{page}, \tocl{indent}{label}{title}{page} List-of-Plates / Contents lines.
\unsure{text} Doubtful reading (logged to unsure.log).
\qslip{as printed} A suspected slip inside quoted matter. Never corrected — the slip may be the quoted source's, not Ross's. Sets the text unchanged, logs to qslips.log; list it in QUESTIONS.md with the source so it can be checked.
\erratum{corrected}{as printed} A slip in the original that this edition corrects: sets the corrected reading, logs to errata.log. Arguments must be plain text (put \emph outside). Every use needs a matching \erratumline{orig page}{as printed}{corrected} in src/errata.tex — make check enforces it — and the pair prints in the Errata section. Handwritten corrections in the scan are not errata: adopt them silently and log in QUESTIONS.md.
\partstart{Part I}{Title} Part-title page: label and title centred about two fifths down, bookmark, new note group. Precedes \origpage like \chapstart.
\sectionstart{Section A}{Title} Section-title page — flush left, as the original sets them.
\emph{} for italics, ` '' for quotes, \ldots\ for …, \&, \%. Typewriter conventions of the original (no ligatures etc.) need no special handling.

Typography (set once in src/preamble.tex — do not re-litigate per page)

12pt TeX Gyre ScholaX (Century Schoolbook; loaded by path from the TeX tree, which fontconfig does not index) on a 5.95in measure; Liberation Sans for margin marks and running heads; Tiro Bangla for Bengali. Leading 1.59 — a 22.9pt line, the original's own, measured off the scan at 300 dpi; the inline specimens set it, not the text (see CONTINUATION.md). Inline specimens at native size (\igscale = 72.27/300), which is the 74% of the line the original gives them. First line indented 1.6em with \parskip 0.15em, \frenchspacing (the original is a typescript with uniform spacing, and TeX's sentence spacing after p. is wrong here). Ragged right and unhyphenated (\RaggedRight; \lefthyphenmin=62 plus \hyphenpenalty=10000 — the original typescript breaks only at spaces and at hyphens it actually typed, so \exhyphenpenalty stays low and \tolerance/\emergencystretch are raised to pay for the lost breaks). (Sammay, 2026-09-14.) Hanging punctuation via microtype protrusion, dashes hanging far less. Widows and orphans forbidden with \raggedbottom. Running head: current chapter (a chap mark class, set by every heading macro) on the left, original page range on the right. Endnotes set a step smaller and tighter via \notesection; plate captions are \footnotesize with their own leading, a clear step below the text as in the original. Contents and List-of-Plates entries are whole-line links to page.N.

tools/polish.py — the build-time typographic pass

build.sh runs it first: it reads src/pages/*.tex and writes work/pages/*.tex, and the build inputs the latter. The transcription in src/pages always stays exactly as typed — never hand-edit the polished copies. The pass applies what the typewriter could not (Sammay, 2026-09-13):

  • ties, so a reference or unit never breaks across a line (p.~63, Part~II, 240~lbs)
  • en-dashes for numeric ranges (1778--1978)
  • small caps for institutional acronyms — an explicit whitelist in the script (IOL, BMS, OUP, SOAS, MS, BL …). Fount codes (CW1, SB4, BM V, VF1) and Roman numerals are deliberately excluded
  • \mbox round every inline specimen so it never breaks from adjacent punctuation
  • cross-reference links: pl. N and chapter N always link (they can only be this thesis's own); p. N/pp. N link only where the thesis cites itself. A page reference inside a citation belongs to somebody else's book: the script tests the preceding clause for citation markers (a title in \emph, ibid., a shelfmark, publication data, a surname followed by a comma) before accepting self-reference words (above, below, see, mentioned). Neither signal → no link; a missing link is cheaper than one that lands on the wrong page. Re-audit after changing this rule: print every link with its context and read them. Protected from all of the above: comment lines, file paths, and the arguments of \erratum, \qslip, \ig, \origpage, \fn, and the heading macros (their text becomes bookmarks and marks, and a link inside them breaks the build).

Inline type specimens — the workflow

For any page where the original pastes typeforms into its own text:

  1. docker compose run --rm -T tex python3 work/lines.py N — lists the page's text lines with their y-fractions. Its line N is usually the line after the one you counted by eye; if a sheet comes back with the wrong line, try the band one earlier.
  2. … work/glyphs.py N y0,y1 [y0,y1 …] — numbered contact sheet of the ink runs on those bands, written to work/sheetN.png. Read it and pick the runs that are specimens.
  3. … tools/crop_inline.py N name=SPEC … — cuts them to plates/inline/pNNNN-name.png. SPEC is a run index plus, optionally, :left|right|first|last|mid|trimright|trimleft to split a run that merged with its neighbour, +PX to start PX into the run, ~PX to keep only its first PX. Use +PX/~PX when the typewriter's space is too noisy to register as a gap.
  4. Where a run defeats all of that (a word set solid inside a footnote), crop by explicit pixel box with PIL and record the box in the manifest notes column so the cut is reproducible.

Per-batch procedure (≈12–15 pages per batch)

  1. make prep A=<pdf> B=<pdf> → scans/hi-*.jpg + work/ocr/p*.txt.
  2. For each page, look at the image. Decide kind: text, plate, mixed, blank.
  3. Text page: python3 tools/ocr2tex.py <pdf> <printed> [--head "Chapter title"] gives a draft; then proofread against the image: italics (titles, transliterated words), \fn{} marks placed exactly where the superscript sits, footnote text moved into \fn{}, all diacritics (ā ī ū ṛ ṃ ṅ ñ ṭ ḍ ṇ ś ṣ ḥ), hyphenation joined, en-dashes kept as - as the original types them. The typescript never hyphenates a word at a line end — verified across all 431 leaves: every one of its 59 line-ending hyphens is either a compound the author typed (offset-lithography, Hof-und, forty-eight) or OCR noise off a reproduced plate. So a hyphen in a page file is always intrinsic to the word; never invent one, and never keep one that splits a word. Output hyphenation is off to match (Sammay, 2026-09-14). Where the author compounds inconsistently (type-casting/typecasting, punch-cutter/punchcutter) rule 3 applies: each instance stands as printed. Page break inside a paragraph: the previous file ends mid-sentence and this file begins \origpage{N} followed by the continuation — no blank line between files (the build concatenates them), so do not start the file with \par.
  4. Plate page: make plate N=<pdf> (inspect plates/pNNNN.png; if the caption or page number survived or content was clipped, rerun with ARGS="--top 0.05 --bottom 0.11" etc.), then a one-line file with \plateop{...} and the caption transcribed exactly.
  5. Update manifest.tsv rows (python3 tools/setrow.py <pdf> <printed> <kind> "notes"); make check; make build; spot-check 2–3 output pages (pdftoppm -jpeg -r 60 -f X -l Y ross-1988-retypeset.pdf work/o).
  6. Update CONTINUATION.md (state, next batch) and QUESTIONS.md. Stop at a chapter or batch boundary, never mid-file.

Environment

Tools run in Docker: make image once, then the make targets above (NODOCKER=1 make … if xelatex, poppler-utils, tesseract, python3-pil/numpy/pdfplumber are installed locally). Fonts: Liberation (system), FreeSerif (system), Tiro Bangla (fonts/). Do not add fonts.