Published on: 2026-09-22
Source: Novosibirsk State University –
An important disclaimer is at the bottom of this article.
A systematic comparison of the principles of linguistic glossing and machine annotation based on the Chinese language was conducted for the first time in domestic sinology by an undergraduate student specializing in “Oriental Studies and African Studies”. Humanitarian Institute of Novosibirsk State University Anna Potapova as part of her thesis “The Specifics of Markup Systems for Machine Text Processing Based on the Chinese Language” under the scientific supervision of a Candidate of Philological Sciences, senior lecturer of the Department of Oriental Studies Maria Durova. Anna Potapova found out which existing annotation systems are better suited for solving various tasks. The results of her research may be useful to specialists engaged in the field of machine translation, automatic summarization, and information extraction for training large language models, as well as for creating intelligent search systems.
Markup – this is a process in which the text is broken down into small parts – tokens – and each of them is assigned certain labels, for example, grammatical ones – for instance, this word is a noun, in the sentence it serves such-and-such a role, it has such-and-such grammatical characteristics – of this kind, case, etc.
Glossing — is a linguistic method of presenting examples in foreign or rare languages, in which an intermediate line with a word-by-word or morphemic analysis explaining the grammatical meaning and structure of each part of the word is placed between the original line and the translated line.
— In my thesis, I examined how well existing text annotation systems are adapted to the Chinese language, because they are mainly designed for languages of a different type, for example Indo-European languages, which have developed grammar and inflection, while Chinese is an isolating language, it has no case inflection or grammatical gender; the category of number is usually not expressed by obligatory inflection. When these systems are applied to Indo-European languages, we can assign various grammatical features to each word: case, number, gender, etc. In Chinese, such features are absent — words do not change. Their grammatical relations are largely determined by word order, function words, particles, and context. To apply these annotation systems to Chinese texts, the linguistic specificity needs to be taken into account and language-specific rules applied. The tags must be standardized, but the method of their assignment should be different. I analyzed how different text annotation systems handle this task. — Anna Potapova said.
A young researcher identified and analyzed specific principles of machine linguistic annotation systems for Chinese text in comparison with the tradition of linguistic glossing. Linguistic annotation of text, as the process of adding special tags to linguistic units to record their morphological, syntactic, and semantic characteristics, serves as a key tool for preparing texts for automatic processing, cross-linguistic comparison, and the creation of electronic resources. The quality and efficiency of natural language processing (NLP) and artificial intelligence technologies development directly depend on the volume and linguistic accuracy of annotated data. High-quality annotated corpora are used in the development, training, and evaluation of various natural language processing models, as well as in the creation of syntactic parsing systems, information extraction, and other specialized NLP tools. Currently, the problem of formalized representation of linguistic data is especially relevant due to the rapid development of digital humanities and corpus linguistics.
— The Chinese language belongs to the isolating language type. It is very different from the languages for which existing markup systems for machine text processing were created. It is an isolating language. In Chinese texts, there are no spaces; externally, they appear as a collection of characters. However, the difficulty in text markup lies not only in the absence of spaces between words but also in the poorly developed morphology. Hypothetically, if in the language some symbols always marked the beginning of a word and others its end, then word segmentation would be a relatively simple procedure even without spaces. The complexity is that in Chinese there are no unambiguous indicators of word boundaries. Moreover, most characters can occupy different positions within a word, so the character itself cannot be a reliable marker of the word boundary. Universal markup standards are forced to adapt to these features, which raises methodological questions that to date remain insufficiently addressed in domestic Sinology., — explained Anna Potapova.
To ensure high accuracy in automatic processing of Chinese texts, adaptation of universal annotation standards, originally primarily oriented towards fusional languages (in which words change form using endings and suffixes to show relationships between them in a sentence. These include Latin and Russian, as well as many Indo-European languages) and agglutinative languages (in which new words and forms are created by sequentially attaching suffixes and prefixes to an unchanging root. Each such element corresponds exactly to one specific meaning, and the boundaries between them are always clear. These include Turkic and Finno-Ugric languages), is necessary. The Chinese language, possessing an isolating structure, lacking clear morphological markers, and having continuous hieroglyphic writing, presents specific challenges to machine annotation systems that have no direct analogues in languages with developed inflection.
— Under the conditions dictated by the peculiarities of the Chinese language, universal annotation standards are forced to adapt: token boundaries in this case are defined not by orthography but by syntactic and lexico-statistical criteria; grammatical categories are encoded not by affixes but by word order, function words, and contextual position. The development of Chinese text machine annotation systems contributes to the creation of new linguistic resources and the expansion of the functionality of digital humanities platforms. In these conditions, the theoretical understanding of the principles of machine annotation and the identification of their correspondence to real tasks of Chinese language processing acquire not only scientific-theoretical but also important practical significance. A comparative analysis of machine annotation systems based on Chinese language material makes it possible to reveal how formal requirements of computational linguistics transform traditional linguistic categories, as well as to show in which aspects machine annotation diverges from human-oriented glossing and in which aspects it inevitably approximates it., — explained the young researcher.
The specifics of using machine linguistic annotation are largely determined by the structural features of the language to which a particular scheme is applied. Universal standards set a general set of categories and a unified way of presenting data, but the specific implementation of annotation depends on the typological features of the language. For the Chinese language, these features are its analytic structure, the absence of developed inflection, the lack of spaces between words, the significant role of function words, and homonymy.
In her work, Anna Potapova presented an overview of the most significant machine annotation systems for the Chinese language, depending on the principles applied in these systems. This approach allows one to see how theoretical premises are embodied in concrete practices and how different schemes reflect various conceptions of the structure of a Chinese sentence. She systematized existing approaches to machine annotation of Chinese text and identified their methodological foundations. Next followed a comparative analysis of the principles of Leipzig glossing and machine annotation according to key parameters: unit of analysis, method of categorization, encoding of grammatical information, resolution of ambiguity, representation format. As a result, specific adaptations of universal annotation standards to the isolating structure of the Chinese language were identified and criteria for choosing an annotation strategy depending on the research goal were formulated.
The systematization of existing annotation systems revealed that the choice of annotation scheme directly depends on the research goal: the Penn Chinese Treebank corpus provides the possibility for syntactic analysis, the Universal Dependencies standard ensures cross-linguistic comparability and cross-linguistic training of language models, and lexically-oriented standards are suitable for frequency analysis and information retrieval. A comparative analysis of five key parameters demonstrated a fundamental difference in the foundations of the two traditions: glossing prioritizes semantic transparency and linear clarity, while machine annotation prioritizes structural unambiguity and algorithmic reproducibility.
During her research, Anna Potapova found that the typological specifics of the Chinese language do not prevent the application of universal annotation standards, but require their adaptation. Token boundaries should be determined by syntactic-distributive criteria, grammatical categories should be encoded through syntactic dependencies, and the resolution of homonymy should be based on contextual features.
The student made an important conclusion that linguistic glossing and machine annotation do not compete, but complement each other — the former remains the optimal tool for philological analysis and scientific publications, while the latter provides scalability and a technical basis for automatic processing.
Anna Potapova also determined that the choice of annotation strategy should be driven by the research objective: Universal Dependencies is preferred for typological comparisons, Penn Chinese Treebank for syntactic analysis, lexically-oriented standards for lexicographic tasks, and Leipzig glossing for illustrative examples in publications.
Material prepared by: Elena Panfilo, NSU press service
Please note; This information is raw content obtained directly from the information source. It represents an accurate report of what the source claims and does not necessarily reflect the position of MIL-OSI or its clients.