Articles
| Open Access |
DOI:
https://doi.org/10.37547/supsci-ojp-06-04-27
EMPIRICAL VALIDATION OF LINGUISTIC ANNOTATION CRITERIA IN LITERARY TEXT MATERIALS: A MULTI-LEVEL PILOT STUDY
Mohinur Alikulova ,Abstract
The multifaceted nature of literary texts and their implicit meanings require a rigorous empirical verification of linguistic annotation criteria when converting them into a digital format. This article presents the results of a pilot empirical study aimed at verifying the practical viability of multi-level linguistic annotation criteria applied to literary materials. A pilot corpus consisting of approximately 52,000 word-forms, selected from 20th-century Uzbek prose and poetry, was analyzed by four specially trained annotators across seven levels: lexico-morphological, syntactic, semantic, pragmatic, stylistic, discursive, and conceptual-cognitive. The study was based on five core management criteria: explicitness, layer independence, cross-level flexibility, interpretive transparency, and scalability. The inter-annotator agreement rate was measured using Cohen's kappa, Krippendorff's alpha, and inter-segment F1 coefficients. The data obtained from the pilot study demonstrated a differentiated nature of agreement indicators, which serves as an objective basis for the targeted refinement of the initial criteria.
Keywords
linguistic annotation; literary text; criteria validation; inter-annotator agreement; multi-level analysis; Uzbek literature; digital humanities.
References
Artstein, R., & Poesio, M. (2008). Inter-Coder Agreement for Computational Linguistics. Computational Linguistics, 34(4), 555–596.
McEnery, T., & Hardie, A. (2012). Corpus Linguistics: Method, Theory and Practice. Cambridge: Cambridge University Press. — 294 p.
Ide, N., & Pustejovsky, J. (Eds.). (2017). Handbook of Linguistic Annotation. Dordrecht: Springer. — 1295 p.
Krippendorff, K. (2018). Content Analysis: An Introduction to Its Methodology (4th ed.). Thousand Oaks: SAGE. — 472 p.
Pragglejaz Group. (2007). MIP: A Method for Identifying Metaphorically Used Words in Discourse. Metaphor and Symbol, 22(1), 1–39.
Steen, G. J., Dorst, A. G., Herrmann, J. B., Kaal, A. A., Krennmayr, T., & Pasma, T. (2010). A Method for Linguistic Metaphor Identification: From MIP to MIPVU. Amsterdam: John Benjamins. — 238 p.
Halliday, M. A. K., & Matthiessen, C. M. I. M. (2014). Halliday's Introduction to Functional Grammar (4th ed.). London: Routledge. — 808 p.
Semino, E., & Short, M. (2004). Corpus Stylistics: Speech, Writing and Thought Presentation in a Corpus of English Writing. London: Routledge. — 248 p.
Mahlberg, M. (2013). Corpus Stylistics and Dickens's Fiction. London: Routledge. — 232 p.
Stockwell, P. (2020). Cognitive Poetics: An Introduction (2nd ed.). London: Routledge. — 222 p.
Lakoff, G., & Johnson, M. (2003). Metaphors We Live By (with new afterword). Chicago: University of Chicago Press. — 276 p.
Fauconnier, G., & Turner, M. (2002). The Way We Think: Conceptual Blending and the Mind's Hidden Complexities. New York: Basic Books. — 464 p.
Pustejovsky, J. (1995). The Generative Lexicon. Cambridge, MA: MIT Press. — 312 p.
Mann, W. C., & Thompson, S. A. (1988). Rhetorical Structure Theory: Toward a Functional Theory of Text Organization. Text, 8(3), 243–281.
Burnard, L. (2014). What is the Text Encoding Initiative? How to Add Intelligent Markup to Digital Resources. Marseille: OpenEdition Press. — 142 p.
Hardie, A. (2014). Modest XML for Corpora: Not a Standard, but a Suggestion. ICAME Journal, 38, 73–103.
Hoover, D. L., Culpeper, J., & O'Halloran, K. (2014). Digital Literary Studies: Corpus Approaches to Poetry, Prose, and Drama. London: Routledge. — 272 p.
Leech, G. (2005). Adding Linguistic Annotation. In M. Wynne (Ed.), Developing Linguistic Corpora: A Guide to Good Practice (pp. 17–29). Oxford: Oxbow Books.
Article Statistics
Downloads
Copyright License
Copyright (c) 2026 Mohinur Alikulova

This work is licensed under a Creative Commons Attribution 4.0 International License.