Files
tepbc-ross-retyped/PLAN.md
T
bdeshiandClaude Opus 5 82849eb94d Initial commit: Ross 1988 re-typeset edition
A searchable, re-typeset edition of Fiona Ross, "The Evolution of the
Printed Bengali Character from 1778 to 1978" (Ph.D., SOAS, 1988),
transcribed from the 431-leaf ProQuest scan. All 431 pages done; 178
plates and 410 inline type specimens cut from the scan; 51 errata.

Tracked: the transcription (src/pages), the preamble and its typographic
decisions, the cut images (plates/ — not reliably regenerable, the crop
specs for the inline cuts were never scripted), tools, and the four
working documents.

Not tracked: the built PDF, which `make` remakes from src/ and plates/;
the ProQuest scan under source/, which is third-party and needed only by
`make prep` and `make plate`; scans/ and work/, both regenerable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-14 19:03:25 +06:00

4.0 KiB
Raw Blame History

Ross 1988 thesis — re-typeset edition: plan (v2, confirmed 2026-09-13)

Source: 10731406.pdf — Fiona G. E. Ross, The Evolution of the Printed Bengali Character from 1778 to 1978, Ph.D., SOAS, 1988. ProQuest scan, 431 PDF pages, 300 dpi bitonal (a few plates in colour), ABBYY text layer (NOT used as source).

Confirmed design

Aspect Decision
Layout Reflowed text. \origpage{N} marks where original page N begins: bold sans number in the outer margin + thin bar in the line + page.N hyper-target. Running head shows the original page range on each edition page; edition page number in the footer
Notes Footnotes become endnotes (enotez) in a final Notes section, grouped "Notes to ", keeping original numbers via \fn{n}{text}; marks link to notes, notes link back
Type 11pt TeX Gyre ScholaX (Century Schoolbook) on a 5.2in measure (≈70 characters a line); Liberation Sans for marks and heads; leading 1.22; ragged right (ragged2e \RaggedRight, hyphenation on); hanging punctuation (microtype protrusion, dashes hang far less); widows and orphans forbidden with \raggedbottom. Tiro Bangla for Bengali that shows the modern script (the Scheme of Transliteration). Bengali that shows a historical typeface is never typed: it is cut from the scan by tools/crop_inline.py and set with \ig at native size
Structure \chapstart{Title} = new page + centred bold head + PDF bookmark + notes group. One .tex per original page in src/pages/pNNNN.tex (NNNN = PDF page); \chapstart must come before \origpage in a file
Plates Cropped from the 300 dpi scan by tools/crop_plate.py (scanner bars, page number, caption stripped; gutter shadow swept; deskewed when the page is out of square — the one step that resamples), embedded lossless at native resolution. \plateop{N}{no}{img}{caption} for a plate that begins page N; \plate{no}{img}{caption} otherwise. Per-page crop recipe recorded in the manifest notes column
Colophon src/colophon.tex = leaf 2: description of the conversion + ProQuest notice text. Blank stamped leaf (PDF p.3) dropped and noted there
Editorial Original spelling and typography stand. Handwritten corrections in the scan are adopted silently and logged in QUESTIONS.md. The author's own slips are corrected via \erratum{corrected}{as printed} and listed in the Errata section. Slips inside quoted matter are never corrected — mark \qslip{as printed} and table them in QUESTIONS.md so they can be checked against the source. A grammatical slip stands, with an editorial [sic]. All confirmed by Sammay, 2026-09-13. The three typographic intentions the typewriter could not realise — en-dashes for numeric ranges, true quotation marks, small caps for institutional acronyms — are applied at set time by tools/polish.py; the transcription itself stays as typed (Sammay, 2026-09-13)
Uncertain \unsure{...} → unsure.log → QUESTIONS.md; never silently guessed
Contents Contents and List-of-Plates entries are whole-line hyperlinks to the original page target (page.N)
Versioning Git (from 2026-09-14, Sammay). The built PDF is not tracked and not archived — make regenerates it from src/ and plates/. The ProQuest scan (source/) is untracked too, being third-party; only make prep and make plate need it. Superseded design drafts stay in older-versions/ as text

Workflow per batch

  1. tools/prep.sh A B → scans/hi-NNN.jpg (130 dpi) + work/ocr/pNNNN.txt (tesseract scaffold).
  2. Read every page image. For clean prose pages tools/ocr2tex.py N printed [--head T] gives a draft; proofread word-by-word against the image, add \emph, \fn, diacritics.
  3. Plates: python3 tools/crop_plate.py NNNN then a one-line \plateop file.
  4. Update manifest.tsv; ./build.sh; spot-check.

Remaining

Introduction pp.17–30 (PDF 19–32) → Part I from p.31 (PDF 33) … Bibliography p.419 (PDF 421) → end (PDF 431). ~245 text + ~170 plate pages: ~10–12 sessions.