Quantitative study of writing style to attribute, compare, or characterize texts by linguistic fingerprint.
Stylometry measures writing style to attribute, compare, or characterize texts through recurring linguistic patterns. It analyzes existing writing rather than generating text in a chosen style. Applications include authorship attribution, forensic document analysis, literary research, and comparison of language-model outputs.
A stylometric pipeline represents texts using features such as character or word sequences, function-word frequencies, sentence lengths, punctuation habits, or grammatical patterns. It then compares those feature distributions with statistical distances, visualization methods, or classifiers. For language models, analysts may pool many responses per system and control for topic and length before testing whether each model has a consistent linguistic fingerprint.
Simple features are inexpensive and interpretable, but they can mistake topic, genre, or verbosity for style. Fine-grained features can reveal subtle habits yet require larger corpora and may not transfer across languages or domains. Results therefore depend strongly on corpus design, feature selection, comparison metrics, and protections against multiple-comparison errors.
It remains unclear which stylometric signals survive paraphrasing, targeted prompting, fine-tuning, or changes in a model's deployment settings. Separating style from stance, content, and training-data overlap is also difficult. Open questions include how stable model fingerprints are over time and when attribution tools become reliable enough for high-stakes use.
Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Stylometry already loaded, so edit it or scan as is.