The line was 78 characters at 12pt, above the 60-75 band — not the "~73"
I had recorded. That figure came from a font-metric estimate (measure
divided by the unweighted mean lowercase advance, counting no spaces) and
was simply wrong. Measured directly off the built PDF — 79 full lines,
glyphs plus inter-word gaps — 12pt ran 78 and 13pt runs 73.
So this is not restoring a lost measure, it is fixing a line that was too
long, and raising the size is the better of the two fixes: narrowing the
measure to 5.7in would have shortened the line while leaving the type
small. (Sammay proposed the size bump.)
Set through fontspec Scale=1.0833 over the 12pt class, deliberately: it
scales every size in the family, including the plate captions, but leaves
\baselineskip class-derived, so \setstretch{1.59} still yields the 22.9pt
line measured off the original. That leading is anchored to the inline
specimens and must not drift with the type size.
Verified after the build:
characters per line 78 -> 73 (target band 60-75)
pages 455 -> 479
underfull boxes 84 -> 71
Bengali/Latin ratio 2.08 -> 2.11 Scale=MatchLowercase is applied
after the main font's Scale, so
Tiro Bangla needed no adjustment
specimens still clear of the lines above and below
overfull 0, font warnings 0, reprocheck 0 tokens short
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
15 KiB
Ross 1988 re-typeset — working instructions for Claude Code
Goal: convert source/10731406.pdf (scan of Fiona Ross, The Evolution of the Printed
Bengali Character from 1778 to 1978, Ph.D., SOAS 1988; 431 PDF pages) into a searchable,
re-typeset PDF. Design decisions are final and recorded in PLAN.md; current state and the
next batch are in CONTINUATION.md; open readings for the owner (Sammay) in QUESTIONS.md.
Read those three files first, every session.
Hard rules
- The scan's text layer is not a source. Transcribe from the page image
(
scans/hi-NNN.jpg, 130 dpi; re-rasterize at 300 dpi for anything small). The tesseract text inwork/ocr/is only a typing scaffold — proofread every word against the image before it goes in a.texfile. Neverpdftotextthe source for content. - Never guess. A word, diacritic, number or glyph you cannot read with confidence is
written as
\unsure{best reading}and added toQUESTIONS.mdwith PDF page, original page and a description. Do not silently normalise or "correct" the author's text. - Original spelling and typography stand, including the author's inconsistencies.
Handwritten corrections in the scanned copy are adopted and logged in
QUESTIONS.md. - One file per original page:
src/pages/pNNNN.tex, NNNN = zero-padded PDF page. Row inmanifest.tsvupdated in the same step (printed page, kind, status=done, src_file, notes). - Bengali: modern script is text, historical typefaces are images. Bengali that shows
the modern script — the Scheme of Transliteration on orig. p.13 — is typed as Unicode and
sets itself in Tiro Bangla. Bengali that shows a historical typeface (every specimen the
original pastes into its own sentences, the List-of-Plates glyphs, plate captions) is never
typed: cut it from the scan with
tools/crop_inline.pyand set it with\ig. The argument in the text is about those shapes — retyping them destroys the evidence. (Sammay, 2026-09-13.) - Plates keep native resolution: crop with
make plate N=…; never resample — the sole exception is the automatic deskew, which rotates (and so resamples) a page that is measurably out of square. The tool also sweeps scanner bars and the grey gutter shadow from the edges. Tune per page withARGS:--top/--bottom/--edge(strip page number, caption, scanner edge),--rotate 270for plates the original prints sideways (lossless),--trim T,R,B,Lafter rotation for skewed scanner edges the heuristics cannot see;--box L,T,R,Bgives the crop explicitly and skips the bar/page-number heuristics altogether — needed for dark plates (photographed manuscripts, photographic plates), where “dark means scanner bar” is false;--raw(with--box) takes the box exactly, skipping the shadow sweep, deskew, bar sweep and autocrop — needed for a plate that reproduces a whole framed page, whose own border rules the edge sweeps mistake for scanner bars and strip along with content.--boxvalues are fractions of the raw page, never pixels. Record the recipe in the manifestnotescolumn. - After each batch:
make checkmust pass, thenmake buildmust succeed, andpython3 tools/reprocheck.pymust print0 tokens short— it compares every token ofsrc/pages/*.texagainst the built PDF and is the only check that catches content the typesetter loses (see CONTINUATION.md for the\hangaftercase it was written for). Do not leave the tree in a state where either fails.make verifyruns all three. - Do not edit
src/main.tex(generated bybuild.sh). Git is the version history (initialised 2026-09-14, Sammay): the build no longer archives the previous PDF, the PDF is not tracked —makeremakes it — and the 68 archived copies inolder-versions/were deleted. Superseded design drafts stay inolder-versions/as text and are tracked. - Ask before changing anything in
PLAN.md's design table.
Macro reference (src/preamble.tex)
| Macro | Use |
|---|---|
\chapstart[Short]{Title} / \chaphead[Short]{Title} |
Starts a chapter/section: new page, centred bold head, bookmark, note group "Notes to Title". Must precede \origpage in the file. Resets note numbering: next \fn is \fn{1}. The optional [Short] is the form used for the PDF outline, the running head and the note group, for a head the original sets on two lines — the Abstract repeats the thesis title above its own, and without it the bookmark reads as one glued string. |
\origpage{N} |
Original page N starts here (mid-sentence is normal). Exactly one per text page file, at the point of the break. |
\fn{n}{text} |
Footnote n of the original → endnote. n must be the original's number; make check enforces 1,2,3… per chapter. |
\plateop{N}{no}{plates/pNNNN.png}{caption} |
A plate that begins original page N (the usual case: one plate per page). |
\plate{no}{img}{caption} |
Plate not at a page start. |
\subhead{text} |
Italic run-in subheading (as the original's Introduction: etc.). |
\chapnum{Chapter 1}{Charles Wilkins} |
Chapter head as the original sets it: number over title, both centred bold; notes group as “Chapter 1: Charles Wilkins”. Precedes \origpage. |
\begin{extract}…\end{extract} |
A displayed quotation as the original sets them — indented both sides, set solid, no quotation marks. An environment, so a quotation broken by an original page break opens in one page file and closes in the next (\origpage then sits inside it). |
\ig{img} |
An inline type specimen cut from the scan by tools/crop_inline.py, set at native size (1 px = 1/300 in). Use wherever the original pastes typeforms into its own text — never retype those as Tiro Bangla, the argument is about their shapes. |
\pg{N} |
Hyperlink to original page N (used in Contents/List of Plates). |
\begin{biblist}…\end{biblist}, \bibgroup{…}, \bibhead{…} |
The Bibliography (orig. pp.419–429): the original sets it solid — a bold group heading (\bibgroup), then one entry per paragraph with the continuation lines indented. biblist gives that block (no paragraph skip, 1.6em hanging indent); \bibhead is a centred bold heading inside it. |
\fig{img} |
An unnumbered illustration the original sets in the run of the text (no plate number, no caption — the preceding sentence carries the reference). Centred, native size, shrunk only if wider than the measure. Crop it by explicit pixel box with PIL to plates/fig-pNNNN.png and record the box in the manifest notes column. |
\pl{no}{caption}{page}, \tocl{indent}{label}{title}{page}, \toclnp{indent}{label}{title} |
List-of-Plates / Contents lines. \tocl carries a page number and links the whole line to page.N; \toclnp is the no-page form for a front-matter entry whose number is set separately with \pg. |
\unsure{text} |
Doubtful reading (logged to unsure.log). |
\qslip{as printed} |
A suspected slip inside quoted matter. Never corrected — the slip may be the quoted source's, not Ross's. Sets the text unchanged, logs to qslips.log; list it in QUESTIONS.md with the source so it can be checked. |
\erratum{corrected}{as printed} |
A slip in the original that this edition corrects: sets the corrected reading, logs to errata.log. Arguments must be plain text (put \emph outside). Every use needs a matching \erratumline{orig page}{as printed}{corrected} in src/errata.tex — make check enforces it — and the pair prints in the Errata section. Handwritten corrections in the scan are not errata: adopt them silently and log in QUESTIONS.md. |
\partstart{Part I}{Title} |
Part-title page: label and title centred about two fifths down, bookmark, new note group. Precedes \origpage like \chapstart. |
\sectionstart{Section A}{Title} |
Section-title page — flush left, as the original sets them. |
\emph{} for italics, ` '' for quotes, \ldots\ for …, \&, \%. |
Typewriter conventions of the original (no ligatures etc.) need no special handling. |
Typography (set once in src/preamble.tex — do not re-litigate per page)
13pt XCharter (Carter's Charter, extended; loaded by path from the TeX tree, which fontconfig does
not index — note the upright face is *-Roman, not *-Regular; set via fontspec Scale=1.0833 over a
12pt class, which scales every size in the family but leaves \baselineskip class-derived so the
leading stays put) on a 5.95in measure — 73 characters a line, measured off the built PDF; Liberation Sans for margin marks and running
heads; Tiro Bangla for Bengali. Leading 1.59 — a 22.9pt line, the original's own, measured off the scan
at 300 dpi; the inline specimens set it, not the text (see CONTINUATION.md). Inline specimens at native
size (\igscale = 72.27/300), which is the 74% of the line the original gives them. First line indented
1.6em with \parskip 0.15em, \frenchspacing
(the original is a typescript with uniform spacing, and TeX's sentence spacing after p. is wrong here).
Ragged right and unhyphenated (\RaggedRight; \lefthyphenmin=62 plus \hyphenpenalty=10000 — the original typescript breaks only at spaces and at hyphens it actually typed, so \exhyphenpenalty stays low and \tolerance/\emergencystretch are raised to pay for the lost breaks). (Sammay, 2026-09-14.) Hanging punctuation via microtype protrusion,
dashes hanging far less. Widows and orphans forbidden with \raggedbottom. Running head: current
chapter (a chap mark class, set by every heading macro) on the left, original page range on the right.
Endnotes set a step smaller and tighter via \notesection; plate captions are \footnotesize with their own leading, a clear step below the text as in the original. Contents and List-of-Plates entries are
whole-line links to page.N.
tools/polish.py — the build-time typographic pass
build.sh runs it first: it reads src/pages/*.tex and writes work/pages/*.tex, and the build inputs
the latter. The transcription in src/pages always stays exactly as typed — never hand-edit the
polished copies. The pass applies what the typewriter could not (Sammay, 2026-09-13):
- ties, so a reference or unit never breaks across a line (
p.~63,Part~II,240~lbs) - en-dashes for numeric ranges (1778--1978)
- no small caps. The pass small-capped a whitelist of institutional acronyms until 2026-09-15
(Sammay). Abbreviations are now plain uppercase throughout: the whitelist could only ever be
partial, and on the Abbreviations page BFBS, LMS and MLCo stood unconverted beside small-capped
BL, BMS, EIC, IOL, OUP, SOAS, SPG — it also split
MS EUR 30across two styles. Completing the list is the wrong fix: the original is a typescript, a typewriter cannot set small caps, so the source defines none and rule 3 leaves every acronym as the plain uppercase it prints \mboxround every inline specimen so it never breaks from adjacent punctuation- cross-reference links:
pl. Nandchapter Nalways link (they can only be this thesis's own);p. N/pp. Nlink only where the thesis cites itself. A page reference inside a citation belongs to somebody else's book: the script tests the preceding clause for citation markers (a title in\emph,ibid., a shelfmark, publication data, a surname followed by a comma) before accepting self-reference words (above,below,see,mentioned). Neither signal → no link; a missing link is cheaper than one that lands on the wrong page. Re-audit after changing this rule: print every link with its context and read them. Protected from all of the above: comment lines, file paths, and the arguments of\erratum,\qslip,\ig,\origpage,\fn, and the heading macros (their text becomes bookmarks and marks, and a link inside them breaks the build).
Inline type specimens — the workflow
For any page where the original pastes typeforms into its own text:
docker compose run --rm -T tex python3 work/lines.py N— lists the page's text lines with their y-fractions. Its line N is usually the line after the one you counted by eye; if a sheet comes back with the wrong line, try the band one earlier.… work/glyphs.py N y0,y1 [y0,y1 …]— numbered contact sheet of the ink runs on those bands, written towork/sheetN.png. Read it and pick the runs that are specimens.… tools/crop_inline.py N name=SPEC …— cuts them toplates/inline/pNNNN-name.png. SPEC is a run index plus, optionally,:left|right|first|last|mid|trimright|trimleftto split a run that merged with its neighbour,+PXto start PX into the run,~PXto keep only its first PX. Use+PX/~PXwhen the typewriter's space is too noisy to register as a gap.- Where a run defeats all of that (a word set solid inside a footnote), crop by explicit pixel box
with PIL and record the box in the manifest
notescolumn so the cut is reproducible.
Per-batch procedure (≈12–15 pages per batch)
make prep A=<pdf> B=<pdf>→scans/hi-*.jpg+work/ocr/p*.txt.- For each page, look at the image. Decide kind:
text,plate,mixed,blank. - Text page:
python3 tools/ocr2tex.py <pdf> <printed> [--head "Chapter title"]gives a draft; then proofread against the image: italics (titles, transliterated words),\fn{}marks placed exactly where the superscript sits, footnote text moved into\fn{}, all diacritics (ā ī ū ṛ ṃ ṅ ñ ṭ ḍ ṇ ś ṣ ḥ), hyphenation joined, en-dashes kept as-as the original types them. The typescript never hyphenates a word at a line end — verified across all 431 leaves: every one of its 59 line-ending hyphens is either a compound the author typed (offset-lithography,Hof-und,forty-eight) or OCR noise off a reproduced plate. So a hyphen in a page file is always intrinsic to the word; never invent one, and never keep one that splits a word. Output hyphenation is off to match (Sammay, 2026-09-14). Where the author compounds inconsistently (type-casting/typecasting,punch-cutter/punchcutter) rule 3 applies: each instance stands as printed. Page break inside a paragraph: the previous file ends mid-sentence and this file begins\origpage{N}followed by the continuation — no blank line between files (the build concatenates them), so do not start the file with\par. - Plate page:
make plate N=<pdf>(inspectplates/pNNNN.png; if the caption or page number survived or content was clipped, rerun withARGS="--top 0.05 --bottom 0.11"etc.), then a one-line file with\plateop{...}and the caption transcribed exactly. - Update
manifest.tsvrows (python3 tools/setrow.py <pdf> <printed> <kind> "notes");make check;make build; spot-check 2–3 output pages (pdftoppm -jpeg -r 60 -f X -l Y ross-1988-retypeset.pdf work/o). - Update
CONTINUATION.md(state, next batch) andQUESTIONS.md. Stop at a chapter or batch boundary, never mid-file.
Environment
Tools run in Docker: make image once, then the make targets above (NODOCKER=1 make …
if xelatex, poppler-utils, tesseract, python3-pil/numpy/pdfplumber are installed locally).
Fonts: Liberation (system), FreeSerif (system), Tiro Bangla (fonts/). Do not add fonts.