Für die Übersetzung unserer mittelalterlichen Rezepte gleichen wir unklare Wörter gegen historische Wörterbücher ab. Fünf davon waren nur in schlechter Qualität digital verfügbar, auf drei verschiedene Weisen kaputt: bei den beiden Dialekt-Idiotika hatte eine generische OCR das historische lange „ſ" durchgängig als „f" gelesen, bei Diefenbachs zwei Glossar-Bänden waren die Quellenangaben verstümmelt und das „ß" verschwunden. Die Idiotika haben wir mit einer auf Fraktur spezialisierten OCR-Engine (Kraken) neu digitalisiert, bei Diefenbach den vorhandenen Volltext gezielt repariert. Bei seinem dreisprachigen Glossar von 1846 half beides nicht: dort erfand die vorhandene OCR Umlaute, die im Druck nicht stehen - den Band haben wir deshalb komplett neu gelesen. Da die Arbeit ohnehin gemacht ist und alle Werke gemeinfrei sind, stellen wir die Ergebnisse hier zum Download bereit - falls sie mal jemand anders brauchen kann.
Lizenz: Alle Originalwerke sind gemeinfrei (die Autoren starben 1868, 1883 und 1900, die Schutzfrist ist seit über hundert Jahren abgelaufen). Die Textdateien sind eine rein maschinelle OCR-Transkription ohne eigene kreative Leistung - auch sie sind damit gemeinfrei nutzbar. Wir bitten um einen Hinweis auf die Quelle bei Weiterverwendung, verpflichtend ist das aber nicht.
August Friedrich Christian Vilmar, Marburg/Leipzig 1868. Regionales Dialekt-Wörterbuch für den Raum Marburg/Kassel/Fulda.
Digitalisiert aus einem durchschossenen Exemplar der Bibliothek des Geographischen Instituts Marburg (983 Buchseiten, davon 492 mit Wörterbuch-Druckseiten - die übrigen eingeschossenen Blätter enthalten handschriftliche Notizen früherer Nutzer und wurden nicht transkribiert, da unser OCR-Modell auf Drucktext trainiert ist).
Georg Autenrieth, Zweibrücken 1899. Regionales Dialekt-Wörterbuch für die Pfalz (linguistisch Teil des rheinfränkischen Dialektraums).
Digitalisiert aus dem vollständigen Scan (Google Books/archive.org), 204 Buchseiten, davon 195 mit Wörterbuch-Inhalt.
Lorenz Diefenbach, Frankfurt am Main 1857. Mittellateinisch-deutsches Glossar, zusammengetragen aus fast 200 Handschriften und Frühdrucken des 14. bis 16. Jahrhunderts. Es übersetzt lateinische Begriffe nicht ins Neuhochdeutsche, sondern direkt in zeitgenössische deutsche Vokabeln und Dialektformen - und nennt zu jeder Form die Handschrift, die sie belegt.
Anders als bei den beiden Idiotika oben haben wir hier nicht neu digitalisiert, sondern den vorhandenen Volltext repariert. Der Band ist in Antiqua gesetzt, nicht in Fraktur, und dreispaltig in sehr kleinem Schriftgrad. Unser Fraktur-Modell liest die deutschen Glossen zwar besser, verstümmelt dabei aber die kleinen Zahlen in Klammern - und genau die sind Diefenbachs Quellenapparat. Wir haben es ausprobiert und den Weg wieder verworfen.
Behoben sind zwei Schadensklassen der alten OCR. Erstens war das ß durchgängig als B gelesen: im ganzen Band stand vorher kein einziges ß, sondern „groB", „fuB", „auB" statt „groß", „fuß", „auß" - 7361 Stellen. Zweitens waren die Quellensiglen verstümmelt, etwa „Aeifr." statt „Aelfr." für das Glossarium Aelfredi - 124 Stellen. Dazu 55 Zeilen Scan-Müll vom marmorierten Vorsatzpapier am Dateianfang.
Lorenz Diefenbach, Frankfurt am Main 1867. Eigenständige Fortsetzung des Glossariums mit zusätzlichen Einträgen - kein alphabetischer Fortlauf des ersten Bandes, sondern ein zweiter Durchgang durch das Alphabet. Wer im Hauptwerk sucht, sollte hier ebenfalls nachsehen.
Hier waren vor allem die Quellensiglen betroffen, 487 Stellen. Die häufigste Abkürzung des Bandes, „Symb." für die Symbolae, war sogar in der Minderheit: 152 korrekte Schreibungen gegen 340 verstümmelte („Symh.", „Sijmh.", „Sgmh." und weitere Varianten). Das ß ist in diesem Band intakt und wurde nicht angetastet.
Lorenz Diefenbach (Hrsg.), Frankfurt am Main 1846. Das Glossar einer Handschrift von 1470, dreisprachig: lateinisches Stichwort, deutsche Glosse, und - wo die Handschrift sie hatte - eine böhmische. Diefenbach hat es zuerst herausgegeben und mit Anmerkungen versehen.
Diesen Band haben wir vollständig neu gelesen, mit spaltenweiser Erkennung über 160 Seiten. Der Grund war nicht schlechte Schrifterkennung im Üblichen, sondern ein tückischerer Fehler: die vorhandene Fassung erfand Umlaute. Am Faksimile nachgeprüft steht gedruckt „Archus i. princeps furst" und „frucht der baum" - die alte OCR machte daraus „fürst" und „bäum". Solche Punkte sieht man im Text nicht als Fehler, sie lesen sich einfach wie ein Wort. Die neue Fassung hat 4767 statt 4681 Stichwörter und gibt die Diakritika so wieder, wie sie auf dem Papier stehen.
⚠️ Eine Besonderheit beim Nachschlagen: das Böhmische steht in der vorhussitischen
Schreibung von 1470, also in Digraphen statt mit Häkchen - cz für č,
ss für š, w für v. „šíje" heißt dort ssyge, „barva" ist
barwa, „zeť" ist zeth. Wer moderne tschechische Schreibung sucht, findet
nichts. Das ist quellentreu und kein Fehler der Erkennung.
Die Kraken-Transkription ist deutlich besser als die vorher verfügbaren generischen OCR-Fassungen, aber nicht fehlerfrei - vereinzelt vertauscht die Engine ähnlich aussehende Buchstaben (z.B. b/d). Wie bei jeder OCR-Digitalisierung eines historischen Drucks gilt: als Lesehilfe und Recherche-Ausgangspunkt nutzen, bei wichtigen Stellen gegen ein Original oder eine gedruckte Ausgabe gegenprüfen.
Bei den beiden Diefenbach-Bänden ist nur das Beschriebene behoben, nicht alles. Die alte OCR verwechselt an anderen Stellen weiter „e" mit „c" (etwa „dcr" statt „der") und löst „m" in „ni" oder „iii" auf; im Vorspann mit seinem großen Schriftgrad ist das dichter als im Wörterbuchteil. Für diese Fehler gibt es keine geschlossene Regel, sie stehen noch drin.
To translate our medieval recipes, we cross-check unclear words against historical dictionaries. Five of these were only available online in poor quality, broken in three different ways: in the two dialect dictionaries a generic OCR had read the historical long „ſ" as „f" throughout, while in Diefenbach's two glossary volumes the source references were mangled and the „ß" had vanished entirely. We re-digitized the dialect dictionaries with an OCR engine specialized for Fraktur/blackletter print (Kraken); for Diefenbach we repaired the existing full text instead. For his trilingual glossary of 1846 neither approach helped: there the available OCR was inventing umlauts that aren't in the print, so we re-read that volume from scratch. Since the work is already done and all the works are in the public domain, we're sharing the results here - in case anyone else can use them.
License: All the original works are in the public domain (their authors died in 1868, 1883 and 1900; the copyright term expired over a century ago). The text files are a purely mechanical OCR transcription with no creative authorship of their own - they are likewise freely usable. We'd appreciate a source credit if you reuse them, but it's not required.
August Friedrich Christian Vilmar, Marburg/Leipzig, 1868. Regional dialect dictionary for the Marburg/Kassel/Fulda area of central Germany.
Digitized from an interleaved copy held by the library of the Geographical Institute in Marburg (983 book pages, 492 of which are actual dictionary print pages - the interleaved blank sheets carry handwritten notes from earlier readers and were not transcribed, since our OCR model is trained on print, not handwriting).
Georg Autenrieth, Zweibrücken, 1899. Regional dialect dictionary for the Palatinate (linguistically part of the Rhenish Franconian dialect area).
Digitized from the full scan (Google Books/archive.org), 204 book pages, 195 of which carry dictionary content.
Lorenz Diefenbach, Frankfurt am Main, 1857. A Medieval-Latin-to-German glossary compiled from nearly 200 manuscripts and early printed books of the 14th to 16th centuries. It renders Latin terms not into modern German but directly into contemporary German words and dialect forms - and names, for each form, the manuscript that attests it.
Unlike the two dialect dictionaries above, we did not re-digitize this one; we repaired the existing text. The volume is set in roman type rather than Fraktur, in three columns at a very small size. Our Fraktur model does read the German glosses better, but it mangles the small parenthesized numbers - and those are precisely Diefenbach's apparatus of sources. We tried it and abandoned the approach.
Two classes of damage are fixed. First, the ß had been read as B throughout: the entire volume previously contained not a single ß, giving „groB", „fuB", „auB" for „groß", „fuß", „auß" - 7,361 instances. Second, the source sigla were corrupted, for example „Aeifr." for „Aelfr." (the Glossarium Aelfredi) - 124 instances. Plus 55 lines of scanner noise from the marbled endpapers at the start of the file.
Lorenz Diefenbach, Frankfurt am Main, 1867. A free-standing continuation of the Glossarium with additional entries - not an alphabetical continuation of the first volume but a second pass through the alphabet. If you search the main work, search this one too.
Here the source sigla were the main problem: 487 instances corrected. The volume's most frequent abbreviation, „Symb." for the Symbolae, was actually in the minority: 152 correct spellings against 340 corrupted ones („Symh.", „Sijmh.", „Sgmh." and further variants). The ß in this volume is intact and was left alone.
Lorenz Diefenbach (ed.), Frankfurt am Main, 1846. The glossary of a manuscript of 1470, in three languages: a Latin headword, a German gloss, and - where the manuscript had one - a Bohemian (Czech) gloss. Diefenbach was its first editor and added commentary throughout.
We re-read this volume in full, column by column across 160 pages. The reason was not ordinary misrecognition but something more insidious: the existing text invented umlauts. Checked against the facsimile, the print reads „Archus i. princeps furst" and „frucht der baum" - the old OCR turned these into „fürst" and „bäum". Stray dots like that don't look like errors in a text; they just read as a word. The new version has 4,767 headwords instead of 4,681 and reproduces the diacritics as they appear on the page.
⚠️ One thing to know when looking things up: the Bohemian is in the pre-Hussite spelling of
1470, using digraphs rather than diacritics - cz for č, ss for š,
w for v. „šíje" appears as ssyge, „barva" as barwa, „zeť" as
zeth. Searching in modern Czech spelling finds nothing. That is faithful to the
source, not an OCR failure.
The Kraken transcription is noticeably better than the previously available generic OCR versions, but not error-free - the engine occasionally confuses visually similar letters (e.g. b/d). As with any OCR digitization of a historical print: use it as a reading aid and research starting point, and cross-check important passages against an original or printed edition.
For the two Diefenbach volumes, only what is described above has been fixed, not everything. Elsewhere the old OCR still confuses „e" with „c" (for instance „dcr" for „der") and breaks „m" into „ni" or „iii"; this is denser in the front matter, with its larger display type, than in the dictionary proper. There is no closed rule for those errors, so they remain.