FLAME is a Python tool with command-line and graphical interfaces for finding formulaic language and text reuse in historical corpora, especially medieval charters. It uses a Leave-N-Out (LNO) n-gram approach, which matches variant forms of an expression across scribal variation, dialect and orthographic change. Normalization rules are learned from the corpus itself (ligatures, special characters), a BPE tokenizer absorbs rare words and morphological variants, the vocabulary size is suggested from the corpus's morphology, the similarity cutoff is set by Otsu's method, and an optional self-supervised auto-tune engine searches the window parameters.
A downloadable demo of the HTML output can be found in the repository (text_comparisons.html).
For a sequence of words (an n-gram), FLAME generates variants by omitting tokens, so underlying similarity survives when the surface forms do not.
Take the charter opening "In nomine sancte et individue trinitatis amen":
- Generate n-grams: slide a window of
ngramwords across the text. - Create LNO variants: remove
n_outtokens from each window. From the 5-gram[In, nomine, sancte, et, individue]come[_, nomine, sancte, et, individue],[In, _, sancte, et, individue], and so on. - Hash: each variant becomes an integer via a vectorised polynomial rolling hash.
- Similarity: cosine similarity between documents over their shared TF-IDF-weighted feature hashes.
- Visualise: interactive reports highlight the matches in their original context.
The LNO-gram method balances context preservation and flexibility better than plain n-grams or skip-grams for historical text.
| Method | Input Text | Subword Tokens (Example) | Generated Patterns (Examples) | Match Score | Notes |
|---|---|---|---|---|---|
| N-gram | In nomine sancte et individue |
[' In', ' nomine', ' sancte', ' et', ' individue'] |
[In nomine sancte et individue] |
1.0 | Rigid. Fails if a single word changes (e.g., dei for nomine). |
| (n=5) | In dei nomine sancte et |
[' In', ' dei', ' nomine', ' sancte', ' et'] |
[In dei nomine sancte et] |
0.0 | No tolerance for variation. |
| Skip-gram | In nomine sancte et individue |
[' In', ' nomine', ' sancte', ' et', ' individue'] |
[In sancte], [nomine et] |
~0.4 | Loses word order and creates noisy, out-of-context pairs. |
| (n=2, k=1) | In dei nomine sancte et |
[' In', ' dei', ' nomine', ' sancte', ' et'] |
[In nomine], [dei sancte] |
~0.3 | High noise, low contextual accuracy. |
| FLAME | In nomine sancte et individue |
[' In', ' nom', 'ine', ' sanct', 'e', ' et', ' in', 'di', 'vid', 'ue'] |
[nomine sancte et _], [In _ sancte et individue]... |
~0.95 | High flexibility. Captures whole-word and sub-word variations. |
| (LNO-gram + Subword) | In dei nomine sancte et |
[' In', ' dei', ' nom', 'ine', ' sanct', 'e', ' et'] |
[dei nomine sancte et _]... |
~0.90 | Robust. Matches novel words or spellings by comparing their constituent parts. |
Where n is the window size, k is the number of skips, and r is the number of removed tokens. Match scores are illustrative.
- Advanced LNO-gram Analysis: generates partial matches by removing combinations of tokens from traditional n-grams.
- Autonomous Parameter Auto-Tuning: a self-supervised "trial digging" engine injects synthetic transcription and dialect noise into a sample of your text to find the
ngramandn_outsetup that suits it. - Adaptive Character Normalization: learns normalization rules (e.g.
é->e, MUFI ligatures) as vectorized NumPy lookup views over the Unicode Basic Multilingual Plane. - Bigram Normalization: optional layer collapsing doubled consonants and diphthongs (
ss→s,ie→i,au→u) via configurable multi-character replacement rules. - Phonetic Reduction Layer: optional rule-based character mapping (
b→p,c→k,v→f) reducing the alphabet to a configurable target set, in the spirit of Metaphone/Soundex. - BPE Subword Tokenization: a Byte-Pair Encoding tokenizer whose vocabulary size is suggested from corpus morphology, absorbing rare words and orthographic variants into shared subword units.
- Inter-Corpus Comparison: two-directory mode (
input_path2) for cross-collection similarity analysis. - Flexible Corpus Selection: takes a directory (searched recursively) or a glob pattern, so a subset of a large corpus can be analysed without copying files — e.g.
./fsdb/DE-LANRWR**/*.htr.txt. - Unambiguous File Identity: reports name every text by its path relative to the input root, so nested corpora that reuse one filename per folder (as fsdb does) stay distinguishable. Flat corpora are unaffected.
- Duplicate Detection (optional): corpora assembled from editions, archival copies and re-downloads hold the same charter under several names. By default every file found is compared, as a scholar expects from what is on disk; with
deduplicateFLAME skips files whose text it has already loaded, keeps the shortest-named copy, and prints what it dropped and why. - Automatic Threshold Detection: Otsu's method on non-zero sparse data picks the similarity cutoff, removing manual guesswork.
- Clustering (recurring legal formulas): groups related pairs by the formula they share — the shared wording found by gapped local alignment — so a cluster answers "these charters use the same legal formula or record the same legal transaction", not merely "these charters are similar". Grouping runs on the spelling-folded wording, so
vnd/undand other medieval variants of one formula stay together. See Clustering. - Multi-Format Reporting: side-by-side HTML comparisons with dynamic fuzzy sliders, a similarity heatmap, a TSV summary of related documents, and a linguistic-variations TSV of alternative spellings and substitutions.
- Modern Tabbed Interface (GUI): a tabbed layout (
ttk.Notebook) separating data configuration, philological fine-tuning, and execution reporting. - High Performance & Scalability: memory-efficient sparse matrices and fast matrix-vector products for heavy historical corpora.
requirements: python 3.9 min, pip 26.1 min
A Python virtual environment is strongly recommended.
-
Clone the repository:
git clone https://github.com/kreeedit/FLAME cd FLAME -
Create and activate a virtual environment:
python -m venv venv source venv/bin/activate # On Windows, use `venv\Scripts\activate`
-
Install dependencies from
requirements.txt:pip install -r requirements.txt
-
Download NLTK data: run the following in a Python interpreter.
import nltk nltk.download('punkt') nltk.download('punkt_tab')
FlameApp.spec packages the GUI and the pipeline behind it into a self-contained folder, so the tool
can be handed to a philologist without a Python environment:
pip install pyinstaller
pyinstaller FlameApp.specThe result is dist/FlameApp/, a onedir bundle (COLLECT) rather than a single file, so the
executable does not unpack itself on every start. Two spec details are deliberate: Plotly's
package_data is bundled as data, which the offline heatmap needs at runtime, and console=True is
kept on so a failure inside a packaged run is readable in the launching terminal — set it to False
for a windowless build once the packaging is known good.
Run the analysis through either interface.
To launch the graphical interface, run:
python flame_gui.py
The window is organised into three tabs:
- Data & Core Setup — the primary corpus path and an optional secondary one for inter-corpus runs (each accepts a folder or a glob pattern),
file_suffix,keep_texts,min_text_length,ngram/n_out, the similarity threshold (a number orauto) with its method dropdown, and the duplicate-file switch. - Normalization & Philology — the character-normalization strategy and its minimum frequency, the bigram and phonetic reduction layers, the BPE vocabulary settings, and
fuzz_threshold,max_gap_words,cluster_threshold,cluster_minandmin_core_specificity. - Auto-Tune & Outputs — the auto-tune switch and its sample size, one checkbox per generated report, the
cluster_linkagedropdown, and the run button.
The pipeline runs on a worker thread, so the window stays responsive. Its stdout and stderr are
streamed live into the Pipeline Execution Standard Output Log pane, where the run's summary lines
appear (Anchoring: …, Duplicate links: …, the clustering statistics), and when the run finishes
one button per generated report opens it directly.
Not every parameter has a GUI field. The GUI carries a widget for 36 of the 43 parameters; the
seven it does not expose are exactly the clustering internals — core_anchors, core_anchor_window,
core_max_tokens, core_gap_tolerance, core_min_tokens, core_identity_threshold and
min_duplicate_tokens. Those keep their DEFAULT_PARAMS value, so the GUI and the CLI agree on the
defaults; changing one means the command line or editing flame.py. The boolean gen_* switches are
the mirror image: the command line can turn those on, never off — see
Command-Line Interface.
Arguments use a single leading dash (-input_path, not --input_path), in either
the -param value or -param=value form.
Boolean switches are presence switches: -auto_tune on its own turns the option on, and so do
-auto_tune False, -auto_tune false and -auto_tune 0 — the value is ignored. A boolean that
defaults to True (the gen_* report switches) therefore cannot be switched off from the command
line; use the GUI checkbox or change the default in DEFAULT_PARAMS in flame.py.
To see all available options and their defaults, run:
python flame.py -h
Example (Using Auto-Tune):
python flame.py -input_path ./path/to/texts -auto_tune True -similarity_threshold auto
Example (A subset of a nested corpus, given as a pattern):
python flame.py -input_path './fsdb/DE-LANRWR**/*.htr.txt' -similarity_threshold 0.60
input_path (and input_path2) accepts either a directory or a glob pattern:
- A directory is searched recursively for every file ending in
file_suffix(*.txtby default), at any depth. - A glob pattern is expanded as-is and decides alone which files are read —
file_suffixis deliberately not applied on top of it, so a pattern like*.htris not silently emptied by a.txtfilter.
This matters for corpora that nest one folder per document and reuse the same filename everywhere (an
fsdb tree of DE-LANRWR001/text.htr.txt, DE-LANRWR002/text.htr.txt, …). A bare filename would be
ambiguous there, so every report names a text by its path relative to the input root, giving
DE-LANRWR001/text.htr.txt. When the input is flat, that relative path is the filename, so existing
outputs are unchanged.
Two further properties:
- The file list is sorted, so corpus order — and with it the row/column order of the distance matrix and the heatmap — is reproducible rather than following the filesystem.
**is only recursive as a standalone path component../fsdb/DE-LANRWR**/*.htr.txtreads the**exactly like*, descending a single level, which is what the fsdb layout needs. For unbounded depth write./fsdb/**/*.htr.txt. FLAME prints a note when it sees a doubled star that is not a standalone component.
| Parameter | Default | Description |
|---|---|---|
input_path |
'' |
Required. Path to the primary corpus: a directory, or a glob pattern such as ./fsdb/DE-LANRWR**/*.htr.txt. |
input_path2 |
'' |
Optional second corpus (directory or glob pattern) for cross-corpus comparison. |
file_suffix |
.txt |
File extension of text documents to process. Ignored when input_path is a glob pattern, which then decides alone which files are read. |
keep_texts |
10000 |
Maximum number of texts to load from each directory. |
deduplicate |
False |
Skip files whose text is identical to one already loaded, keeping the shortest name. Off by default: whether two identical files are two witnesses or one is a scholarly judgement. The two sides of an inter-corpus run are deduplicated independently, so a document shared by both is still compared across them. |
ngram |
6 |
The size of the n-gram window for feature generation. |
n_out |
1 |
Number of tokens to "leave out" (drop) from each n-gram window. |
min_text_length |
150 |
Minimum character length for a file to be included in the corpus. |
similarity_threshold |
'auto' |
Similarity cutoff score. Can be a float (e.g., 0.75) or 'auto'. |
auto_threshold_method |
'otsu' |
Method for auto-thresholding: 'otsu' or 'percentile'. |
char_norm_alphabet |
abcdef... |
String of allowed lowercase base characters for normalization. |
char_norm_strategy |
'normalize' |
Strategy for handling unknown out-of-alphabet characters. |
char_norm_min_freq |
1 |
Minimum frequency for the adaptive normalizer to register an automated Unicode rule. |
phonetic_reduction_enabled |
False |
Enables the phonetic reduction layer (character mapping to a reduced alphabet). |
phonetic_reduction_alphabet |
aefiklmno... |
Target reduced alphabet for phonetic mapping. Characters outside this set become spaces. |
phonetic_reduction_rules |
b>p,c>k,... |
Comma-separated phonetic mapping rules in src>dst format (e.g., b>p,c>k). |
bigram_normalization_enabled |
False |
Enables bigram normalization (multi-character to single-character replacement). |
bigram_normalization_rules |
ss>s,ff>f,... |
Comma-separated bigram rules in src>dst format (e.g., ss>s,ie>i,au>u). Source must be 2+ chars, destination exactly 1. |
vocab_size |
'auto' |
Target subword vocabulary size. An integer, or 'auto' to calculate via morphology. |
vocab_min_word_freq |
5 |
Minimum frequency for a word to be evaluated for affix candidates. |
vocab_coverage |
0.85 |
Desired morphological coverage percentage of the corpus when vocab_size is 'auto'. |
fuzz_threshold |
0.70 |
Orthographic-variant vs. bridge cutoff (0-1): words between two matches whose spelling-folded forms score at least this high are an Orthographic Variant, below it a genuine Bridge. Measured on real two-copies comparisons, genuine divergences score ≤ 0.40 and spelling variants ≥ 0.72. |
max_gap_words |
5 |
Maximum structural token length allowed inside an individual non-matching gap segment. Larger gaps are left unmarked rather than classified. |
cluster_threshold |
0.70 |
Edit-similarity cutoff (0-1) at which two pairs' shared formulas count as the same formula, judged on the formula rather than the whole charter. 1.0 gives verbatim-only clusters. 0.70 is the default rather than KONI's 0.85, because on the MOM corpus 0.85 split one papal privilege template into clusters of 24, 12, 6 and 2 charters over differing recipients, and 0.70 merges them with no false merge among clusters of four or more. 0.65 starts over-merging. |
cluster_min |
2 |
Minimum number of pairs a cluster must contain to be reported. |
cluster_linkage |
'louvain' |
How the edit-similarity graph becomes clusters. louvain (default) finds modularity communities — cores denser among themselves than with the rest — keeping one template whole while refusing to chain two formulas that merely share a phrase; it needs networkx. clique splits each group into maximal cliques, so every formula matches every other one, but a chained group is broken apart and a formula between two groups appears in both clusters. union is KONI's original union-find, which chains. Not a boolean on purpose: -flag arguments are presence switches, so a True-by-default boolean could not be turned off from the command line. |
core_gap_tolerance |
8 |
How wide a gap the core aligner may bridge, in tokens, on either side. Strict word-for-word matching breaks a formula in two at the first inserted word (uestro auctoritate against auctoritate was enough); eight tokens spans the widest slot measured in the corpus without bridging two genuinely different formulas. |
core_min_tokens |
12 |
Shortest aligned window (its shorter side) that still counts as a formula. Below it the pair falls back to the longest contiguous run; the run reports how many took that path. |
core_anchors |
'' |
Where a formula is looked for: the performative verbs that carry a charter's legal act (donamus, contulimus, confirmamus…), as comma-separated stems prefix-matched against the folded tokens. Without them the alignment maximises matched tokens, and the longest shared run in a mediaeval charter is the protocol — on the MOM corpus no core came out below 56 tokens (median 208) and 28 of 40 clusters opened on a protocol phrase. A pair whose shared text carries no anchor keeps the unanchored window. Empty by default, and there is no built-in list — see Clustering for the measurements behind that. -core_anchors=donau,contul,tradid supplies them; -core_anchors= (the = form — fargv crashes on a bare empty argument) and -core_anchors none both select the unanchored extraction. |
core_anchor_window |
15 |
How far the window may extend past the anchor, in tokens, on either side. |
core_max_tokens |
50 |
Hard ceiling on the reported window, in tokens; the window is centred on the anchor and clamped inside the aligned span. A dispositio runs 15–40 words while a protocol panel runs to 200, so without a ceiling a long generic overlap outvotes a short specific one. 0 leaves the window uncapped. The near-duplicate test reads the pair's whole shared span, not this window. |
core_identity_threshold |
0.0 |
Minimum identity (2 × matched / (len_a + len_b), the same normalisation as cluster_threshold) for an aligned window to be accepted. 0.0 accepts every window the aligner returns; raise it to demand a tighter formula and send the loose pairs back to the contiguous rule. |
min_duplicate_tokens |
400 |
A cluster is reported as a near-duplicate charter rather than a formula when its shared window is at least this many words and covers at least 60% of the shorter charter — both conditions, see Clustering. |
min_core_specificity |
0.0 |
Drop clusters whose Specificity (CoreTokens × MeanIDF) is below this; 0.0 keeps everything. The knob for universal chancery phrases: a formula every charter carries has a low mean IDF whatever its length, so 10–15 keeps the specific formulas and discards the generic scaffolding. Filtering happens before cluster IDs are handed out, so the ClusterID columns of the other TSVs stay consistent. |
stopwords_file |
'' |
Path to a plain-text list of tokens, one per line. Listed tokens score IDF 0, so a formula built only from them scores Specificity 0 and min_core_specificity can remove it — the hand-written counterpart to the data-driven IDF. |
auto_tune |
False |
Enables self-supervised parameter discovery via temporary synthetic noise sweeps. |
auto_tune_sample_size |
30 |
Number of document vectors to isolate and sample when executing an auto_tune sweep. |
no_reports |
False |
If True, skips generating user-facing visual summaries and reports completely. |
gen_comparison_html |
True |
Generate interactive side-by-side HTML comparison reports. |
gen_summary_tsv |
True |
Generate a TSV summary of related documents and matching segments. |
gen_linguistic_tsv |
True |
Generate a granular TSV of linguistic variations (alternative spellings, substitutions). |
gen_heatmap |
True |
Generate an interactive Plotly similarity heatmap HTML file. Skipped, with a printed note, when the similarity matrix reaches 2000 documents in either dimension — a corpus that large is read better from the cluster and pair reports. |
gen_clusters |
True |
Generate the clustering (recurring legal formulas) report. Also adds a ClusterID column to similarity_summary.tsv and linguistic_variations.tsv — see the note under Outputs. |
A run produces up to six kinds of report in the directory where it was started; which appear depends
on the gen_* switches and, for the heatmap, on the corpus size. In all of them a document is
identified by its path relative to the input root (see
Input Paths, Patterns, and How Files Are Named).
dist_mat.npz: a SciPy sparse matrix of all pairwise similarity scores, for downstream validation without re-computing features. Where it lands depends on how FLAME was started: the GUI writes it, with the other intermediates (temp_corpus.txt,bpe_tokenizer.json), into the directory it was launched from; a CLI run keeps them in a temporary directory that is removed when the run ends. A command-line run therefore leaves the six reports below and nothing else — take the matrix from a GUI run, or from the engine's API with an explicit working directory.text_comparisons_XX.html: interactive side-by-side alignment reports, the primary visualisation for philological exploration.
- Synchronized scroll-locking and text matching cross-highlights.
- A Live Fuzzy Slider to adjust, in the browser, which structural bridges count as similar.
- A three-way classification of the words between two matches, by colour and on-page legend:
- Bridge word (yellow) — the two texts genuinely diverge here.
- Orthographic variant (green) — the same wording spelled differently, e.g.
deßhalb/deshalb. - Insertion (blue) — the wording is present in one of the two texts only.
- Directional layout control placing earlier documents on the left based on filename year markers.
similarity_heatmap.html: an interactive Plotly heatmap of the full pairwise similarity matrix. Generated only while both dimensions stay below 2000 documents; above that the run printsSkipping heatmap generation for large matrixand produces no file.similarity_summary.tsv: a spreadsheet summary of related matches, document frequencies, and prominent matching blocks (>4 words).linguistic_variations.tsv: a corpus-wide register of what sits inside identical formulaic expressions, with columnsFile_1,File_2,ClusterID,Variation_Type,Token_1,Token_2.Variation_Typeis one of:
Orthographic Variant— the same wording spelled differently on each side.Insertion— present in one of the two texts only (the other column carries-).Different Bridge Word— genuinely divergent wording between two matches.
clusters.html/clusters.tsv: the clustering report — one card and one row per recurring legal formula, naming the formula, the documents using it, the pairs with their two scores, and the run's statistics.Anchornames the performative verb the window was seeded on (empty when the shared text carries none),DistinctTextsthe number of different texts behind the document list, so a pair count inflated by repeated copies is visible. See Clustering.
The classification uses the same code path as the HTML report, so the register and the visualisation never disagree about what counts as a variant and what as a bridge.
Note — column layout changed.
similarity_summary.tsvandlinguistic_variations.tsvcarry aClusterIDcolumn whilegen_clustersis on (its default). A script reading those files by column position must be updated: the summary'sClusterIDis the second column (right afterDocumentFilename), the linguistic register's is the third (afterFile_2, beforeVariation_Type). Reading by header name is unaffected.clusters.tsvis a new file, with columnsClusterID, Size, Anchor, CoreFormula, CoreFolded, Documents, DistinctTexts, CoreTokens, MeanIDF, Specificity, Cohesion, Members, PairCosine, PairCoreRatio, SharedTokens, Coverage, Kind. Turning the option off restores the older layout, so long as it is turned off in the GUI or by setting the default inflame.py— the command line cannot switch a boolean off, since-flagarguments are presence switches (see Command-Line Interface).
Similarity says which charters resemble each other; clustering says why. For medieval charters the interesting "why" is usually legal: the same arenga, the same disposition formula, the same transaction recorded by two scribes. The component is a port of the three-stage clustering algorithm of the KONI project (see Acknowledgements), re-targeted from literary formulas to legal ones.
How it works. For every pair above the similarity threshold, FLAME extracts the pair's core by
gapped local alignment: a chain of matching word-runs whose gaps stay within core_gap_tolerance
tokens on either side. Gapped rather than contiguous, because a formula varies at its slots — a single
inserted or substituted word (uestro auctoritate against auctoritate) used to cut one formula into
two runs, and the strict rule reported whichever half was longer. On the FLAME corpus that fragmented
one papal confirmatio into six micro-clusters with a median core of 27 words; the aligned core spans
the same formula at a median of 106 words. Measured against a token-level Smith-Waterman
implementation on the same pairs, the chain picks the same windows at a fraction of the cost.
Cores are taken on spelling-folded tokens, so vnd/und and czu/zu fold together and one
formula spelled two ways stays one formula. Identical cores collapse in one pass; the remaining
distinct cores are compared to each other — not pairs to pairs, which would repeat the same
comparison thousands of times — and merged by edit similarity at cluster_threshold. Where that
chain opens is the anchoring question, answered by core_anchors; by default the aligner takes the
chain carrying the most matched words, which in a charter is the protocol rather than the act.
Because the distinct cores are the graph's nodes, membership follows the core, not the pair. The
pair-level near-duplicate rule is therefore evaluated first and the pairs that clear it are wired into
the graph as forced edges. The run prints both routes (Anchoring: …, Duplicate links: …), so a
report never leaves the reader guessing which built a cluster.
Whole-charter cosine decides only which pairs are worth looking at, never who shares a cluster with
whom. Both numbers are printed per pair (PairCosine is the whole-charter score that admitted the
pair, PairCoreRatio its formula similarity to the cluster's headline formula), and Cohesion — the
weakest core-to-core score inside a cluster — says how tight the cluster really is.
Reading a cluster. A cluster is not a similarity clique. In charters it means one of:
- the same legal formula — an arenga, a corroboratio, a penalty clause, word for word across unrelated documents;
- the same legal transaction — one act, two or more surviving copies or confirmations;
- a shared chancery template — a formulary or Formularbuch in use by one scribe, notary or issuing office.
CoreFormula is the formula as the charters actually spell it; CoreFolded is what the clustering
compared, and the evidence when two charters land in one cluster despite visibly different spelling.
Grouping (cluster_linkage). Edit similarity at a fixed threshold chains: A resembles B, B
resembles C, A and C do not. clique is the opposite extreme — each group is re-scored pairwise and
split into maximal cliques, so every member matches every other one and Cohesion is at or above
cluster_threshold by construction. Neither extreme fits a template, whose agreement edges do not
close into triangles everywhere: on the FLAME corpus the papal confirmatio survived as one connected
component that strict cliques then cut into five overlapping pieces. The default louvain asks the
question the diplomatist is asking — which cores are denser among themselves than with the rest —
so the template stays in one cluster while two formulas that merely share a phrase stay apart. union
restores KONI's transitive merging, the right choice when looking for families of related formulas
rather than exact recurrence.
louvain and union return disjoint clusters, so ClusterID is a single number and no pair appears
twice. clique overlaps by design and FLAME keeps the overlap rather than resolving it: a formula
between two groups appears in both clusters, marked shares pair(s) with cluster N, and
ClusterID in the summary TSVs becomes a comma-separated list.
Formulas versus copies. Not every long shared window is a formula. When it covers most of the
shorter charter, the two documents are not witnesses to a formula but the same charter written out
again, and the report lists those clusters in their own section so they neither pad nor distort the
formula clusters above. Both conditions are required, because either alone misreads a real case: the
share alone calls a short charter a duplicate when its text is mostly one formula (measured here:
eighteen distinct charters whose 160-word papal confirmatio is 99% of each), and the length alone
calls a long shared section a duplicate. A near-duplicate therefore needs a window of at least
min_duplicate_tokens words (default 400, longer than the longest formula in this corpus, 285 words)
and at least 60% of the shorter charter. Every card prints both numbers; clusters.tsv carries
them in SharedTokens and Coverage and marks each row formula or near-duplicate in Kind.
Both numbers are read from one pair — the widest overlap that clears the rule, or the widest
overlap overall when none does. They were once two independent maxima, so the two figures on a card
described no single pair: on the MOM corpus that happened in 11 of the 40 clusters, once pairing a
Coverage of 100% with a SharedTokens from a different, longer pair. No Kind flipped, but the
verdict is assembled from those two numbers, so both now come from the pair that carries the label.
(Coverage 1.000 with Kind: formula is not a contradiction on its own — the rule needs the word
count as well.)
The rule also decides cluster membership, not just the label, because a core narrowed to the legal
act can match nothing even between two copies of one charter: anchored runs on the MOM corpus left six
charter documents sharing 400+ words in no cluster at all, their pair alone below cluster_min.
A family of one pair stays below cluster_min and out of the cluster report; the run prints how many,
because a lone pair is a real find that only the pairs report can carry.
How specific is a formula? Length alone does not make a formula interesting: salutem et apostolicam benedictionem and the scripti patrocinio communimus clause are both recurring formulas,
but only the second says anything about the charter it is in. Each cluster's headline formula is
scored by length × rarity:
CoreTokens— how many words the formula has;MeanIDF— the average inverse document frequency of its distinct words, in nats, smoothed aslog((1 + N) / (1 + df))so a word present in every charter scores0rather than going negative;Specificity—CoreTokens × MeanIDF.
min_core_specificity cuts below a value you choose, and stopwords_file zeroes out tokens you
already know are uninformative. Both act before cluster IDs are assigned, so cross-references from
the other TSVs stay valid.
Watch the pair counts on a corpus with repeated files. A corpus assembled from editions and
archival copies holds the same charter under several names — a shelfmark, an edition, a (1)
re-download. Since a pair is a pair of files, one comparison then appears as nine, and a cluster can
read "9 pair(s), 6 document(s)" while holding two texts. DistinctTexts and the
(N distinct text(s)) note expose that; deduplicate removes the repeats upstream. The two are
independent: the note reports what is in the corpus, deduplicate decides what gets loaded.
The summary block at the top of clusters.html answers on what input were these clusters
computed: the mode and path or glob pattern, file_suffix, how many charters each side contributed
and how many files were found, loaded, skipped as too short, skipped as duplicates or cut off by
keep_texts, whether the threshold was automatic and by which method, the cluster threshold and
linkage, a timestamp, and the command line.
Reading the charters themselves. Unless the corpus is too large to embed (the report says so and
falls back to name lists), every charter in every cluster appears as a collapsible block with its
complete text and the shared formula in <mark>. The highlight maps the same (start, size) span
the clustering used back through the tokeniser, so it marks the words the algorithm matched rather
than a re-search for the string; a document carrying two related formulas has both marked.
Tuning. cluster_threshold sets how strictly "the same formula" is judged: 1.0 clusters only
verbatim-identical cores, 0.70 (the default) also tolerates the substitution of the template's
variable slots — a different recipient, a different place — so one template stays one cluster, and
0.65 or below starts merging formulas that merely share a phrase. The default was lowered from
0.85 after measuring both on the MOM corpus: at 0.85 a single papal privilege template whose
recipients differed fragmented into four clusters, at 0.70 it is one cluster and no other cluster of
four or more members merged. The core's shape is tuned separately: core_gap_tolerance widens or
narrows the slot the aligner may bridge, and core_identity_threshold filters the loose end of the
distribution. Both are reported in the input summary.
Anchoring. The window can be seeded on a performative verb through core_anchors — donamus,
contulimus, confirmamus — grown in both directions and capped at core_max_tokens, so a donation
is clustered on donamus … duas partes decimarum rather than on the papal greeting above it, and the
Anchor column names the verb. Supplied on the MOM corpus, anchoring moves the median core from 208
tokens to 82 (not to the 50-token ceiling: 153 of 441 pairs carry no anchor in their shared text and
keep the long window) and the formula clusters from 29 to 32, with the near-duplicate section growing
from 43 to 50 charter documents.
It ships off because the anchor set cannot be derived from the corpus. A hand-written
Latin-and-German list was tried first and does not hold up on the corpus it was written for: of its 38
stems, 11 (vendidimus, vendimus, emancipauimus, concessimus, promittimus, assignauimus,
commutauimus, verleihen, verkaufen, ubergeben, …) do not occur in the corpus at all, and
153 of the 441 admitted pairs (34.7%) share no word of it, so a third of the corpus fell back to the
unanchored window anyway. Automatic discovery was then tried and failed four ways:
| Attempt | Result |
|---|---|
| Distributional rules: document frequency, occurrences per document, context entropy, positional spread, neighbour IDF, and every conjunction of them | best rule selected 8676 tokens to find 45 verbs — precision 0.5%, recall 34.9% |
| Phrase-level rarity: ranking a pair's shared windows by the document frequency of their 4-grams | the act lands at relative rank 0.467 — chance |
| Rarity as the window's objective: mean, summed and rarest-word unigram IDF, and phrase DF | all anti-select the act (median relative rank 0.57–0.70) |
| Morphology, with the ending inventory of the known verbs handed over as an oracle | -mus / -nt / -it still selected 2215 tokens to find 50 verbs (2.3%) |
The verbs are 0.3% of the recurring vocabulary and statistically ordinary — contulimus has document
frequency 416 beside iuris at 1093, in a mid-frequency band holding 1241 tokens of which 16 are
verbs — and rarity as an objective is the wrong sign, because the act formula is shared by its whole
act family while the average shared passage is a one-off list of names. Identifying a performative
verb is therefore a lexical or morphological task, and the anchor set is per-corpus input rather than
a constant in the code. Two notes for supplying one: stems match as prefixes, so a stem that is too
broad swallows the nouns derived from the verb (dona matches donatio, turning a donation act into
a donation word), and IDF is not the seed's objective — it only chooses between anchors the act is
already known to be near.
cluster_min filters out formulas that occur only once. On a corpus where the same charter is filed
under several names, deduplicate is usually the more useful first move: it removes the repeated
comparisons before they inflate the pair counts.
For direct text copying, transmission lineages and structural plagiarism with minimal alterations.
DEFAULT_PARAMS = {
'input_path': './corpus',
'file_suffix': '.txt',
'keep_texts': 100000,
'ngram': 15, # Target long sequential strings
'n_out': 1, # Enforce rigid matching (only 1 drop allowed)
'min_text_length': 150,
'similarity_threshold': 0.85, # High threshold bar for rigid matches
'auto_threshold_method': 'otsu',
'char_norm_alphabet': "abcdefghijklmnopqrstuvwxyz",
'char_norm_strategy': 'normalize',
'char_norm_min_freq': 2,
'vocab_size': 8000, # Large vocabulary to force whole-word evaluation units
'vocab_min_word_freq': 3,
'vocab_coverage': 0.85,
}Balanced windows with widened token dropping, to see through heavy lexical change.
DEFAULT_PARAMS = {
'input_path': './corpus',
'file_suffix': '.txt',
'keep_texts': 100000,
'ngram': 10, # Phrase-level matching properties
'n_out': 3, # Higher tolerance for token shifts and insertions
'min_text_length': 150,
'similarity_threshold': 0.60, # Moderate cutoff to surface paraphrased reuse
'auto_threshold_method': 'otsu',
'char_norm_alphabet': "abcdefghijklmnopqrstuvwxyz",
'char_norm_strategy': 'normalize',
'char_norm_min_freq': 2,
'vocab_size': 'auto', # Handled flexibly via subwords
'vocab_min_word_freq': 3,
'vocab_coverage': 0.85,
}The standard optimized configuration, with autonomous thresholds and gapped footprints.
DEFAULT_PARAMS = {
'input_path': './corpus',
'file_suffix': '.txt',
'keep_texts': 100000,
'ngram': 6, # Standard formulaic anchor bounds
'n_out': 1, # Adaptive single gap indexing
'similarity_threshold': 'auto', # Calibrate cut-offs automatically via Otsu
'auto_threshold_method': 'otsu',
'char_norm_alphabet': "abcdefghijklmnopqrstuvwxyz",
'char_norm_strategy': 'normalize',
'char_norm_min_freq': 1,
'vocab_size': 'auto', # Morphologically optimized on-the-fly
'vocab_min_word_freq': 5,
'vocab_coverage': 0.85,
}Groups the related pairs by the formula they share rather than by overall resemblance, so the result
reads as a list of formulas — each with the charters that use it — instead of a list of similar pairs.
The default cluster_threshold of 0.70 tolerates ordinary scribal variation inside a formula and
the substitution of its variable slots; raise it to 1.0 to see only verbatim copies.
DEFAULT_PARAMS = {
'input_path': './corpus',
'file_suffix': '.txt',
'keep_texts': 100000,
'ngram': 6,
'n_out': 1,
'similarity_threshold': 'auto', # controls which pairs even reach the clustering
'auto_threshold_method': 'otsu',
'cluster_threshold': 0.70, # controls which of them count as one formula
'cluster_min': 2,
'cluster_linkage': 'louvain', # density-based; 'clique' for the strict guarantee
'core_anchors': '', # '' = unanchored; give stems to seed on the legal act
'core_anchor_window': 15, # how far past the verb the window may reach
'core_max_tokens': 50, # ceiling on the reported window; 0 = uncapped
'min_core_specificity': 0.0, # raise to ~10-15 to drop universal panels
'stopwords_file': '', # optional hand-written diplomatic stop list
'gen_clusters': True,
'char_norm_alphabet': "abcdefghijklmnopqrstuvwxyz",
'char_norm_strategy': 'normalize',
'char_norm_min_freq': 1,
'vocab_size': 'auto',
'vocab_min_word_freq': 5,
'vocab_coverage': 0.85,
}similarity_threshold and cluster_threshold answer two different questions and are tuned together:
the first decides which charters are compared at all, the second which of those are judged to share a
formula. The report's statistics line tells you how much work the gate did (lev_compares,
gate_skips), the quickest way to tell whether a threshold is too strict.
Corpora such as fsdb nest one folder per document and reuse the same filename everywhere. Point
input_path at a pattern to select a series; the reports still tell the documents apart, because each
is named by its path relative to the pattern's leading directory.
DEFAULT_PARAMS = {
'input_path': './fsdb/DE-LANRWR**/*.htr.txt', # one series, one level down
'file_suffix': '.txt', # ignored: the pattern decides
'keep_texts': 100000,
'ngram': 6,
'n_out': 1,
'min_text_length': 150,
'similarity_threshold': 'auto',
'auto_threshold_method': 'otsu',
'char_norm_alphabet': "abcdefghijklmnopqrstuvwxyz",
'char_norm_strategy': 'normalize',
'char_norm_min_freq': 1,
'vocab_size': 'auto',
'vocab_min_word_freq': 5,
'vocab_coverage': 0.85,
}Reports then read DE-LANRWR001/text.htr.txt rather than an ambiguous text.htr.txt. To compare two
series against each other, put one pattern in input_path and the other in input_path2; each side is
named relative to its own root.
The character normalization components build on principles found in Anguelos Nicolaou's library, whose efficient character mapping was a valuable reference for this project.
The clustering component is a port of the clustering algorithm from the KONI project (Korpus Nyelvtechnológiai Infrastruktúra), which groups literary formulas in Hungarian corpora. FLAME keeps KONI's three-stage structure and its scoring metric, and re-targets it at the legal formulas of medieval charters.
Kovács, T. (2025). FLAME: Formulaic Language Analysis in Medieval Expressions (Version 1.0.0) [Computer software]. GitHub. https://github.com/kreeedit/FLAME
@software{Kovacs_FLAME_2025,
author = {Kovács, Tamás},
title = {{FLAME: Formulaic Language Analysis in Medieval Expressions}},
version = {1.0.0},
publisher = {Zenodo},
year = {2025},
doi = {10.5281/zenodo.15805449},
url = {[https://github.com/kreeedit/FLAME](https://github.com/kreeedit/FLAME)}
}
This project is licensed under the Apache 2.0 License.

