• Contenu principal
  • Menu
OpenEdition Books
  • Accueil
  • Catalogue de 16501 livres
  • Éditeurs
  • Auteurs
  • Facebook
  • X
  • Partager
    • Facebook

    • X

    • Accueil
    • Catalogue de 16501 livres
    • Éditeurs
    • Auteurs
  • Ressources numériques en sciences humaines et sociales

    • OpenEdition
  • Nos plateformes

    • OpenEdition Books
    • OpenEdition Journals
    • Hypothèses
    • Calenda
  • Bibliothèques

    • OpenEdition Freemium
  • Suivez-nous

  • Lettre d’information
OpenEdition Search

Redirection vers OpenEdition Search.

À quel endroit ?
  • Accademia University Press
  • ›
  • Collana dell'Associazione Italiana di Li...
  • ›
  • Proceedings of the Eighth Italian Confer...
  • ›
  • Contributed Papers: Long Papers
  • ›
  • A Methodology for Large-Scale, Disambigu...
  • Accademia University Press
  • Accademia University Press
    Accademia University Press
    Informations sur la couverture
    Table des matières
    Liens vers le livre
    Informations sur la couverture
    Table des matières
    Formats de lecture

    Plan

    Plan détaillé Texte intégral 1. Introduction 2. Background and Related Work 3. The Multilingual Word Alignment 4. Implementation 5. Conclusions and Future Work Bibliographie Notes de bas de page Auteurs

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Ce livre est recensé par

    Précédent Suivant
    Table des matières

    A Methodology for Large-Scale, Disambiguated and Unbiased Lexical Knowledge Acquisition Based on Multilingual Word Alignment

    Francesca Grasso et Luigi Di Caro

    p. 169-175

    Résumé

    In order to be concretely effective, many NLP applications require the availability of lexical resources providing varied, broadly shared, and language-unbounded lexical information. However, state-of-the-art knowledge models rarely adopt such a comprehensive and cross-lingual approach to semantics. In this paper, we propose a novel automatable methodology for knowledge modeling based on a multilingual word alignment mechanism that enhances the encoding of unbiased and naturally disambiguated lexical knowledge. Results from a simple implementation of the proposal show relevant outcomes that are not found in other resources.

    Texte intégral Bibliographie Notes de bas de page Auteurs

    Texte intégral

    1. Introduction

    1Lexical resources constitute a key instrument for many NLP tasks such as Word Sense Disambiguation and Machine Translation. However, their potential may vary widely depending on the nature of the lexical-semantic knowledge they encode, as well as on how the linguistic data are stored and linked within the network (Zock and Biemann 2020). The resources that are presently available, such as WordNet (Miller 1995), typically encode lexical-semantic knowledge mainly in terms of word senses, defined by textual (i.e. dictionary) definitions, and lexical entries are linked and put in context through lexical-semantic relations. These relations, being only of a paradigmatic nature, are characterized by a sharing of the same defining properties between the words and a requirement that the words be of the same syntactic class (Morris and Hirst 2004). Typically related words are therefore not represented due to the absence of syntagmatic links. Additionally, word senses suffer from a lack of explicit common-sense knowledge and context-dependent information. Finally, the well-known fine granularity of word senses in WordNet (Palmer, Dang, and Fellbaum 2007) is due to the lack of a meaning encoding system capable of representing concepts in a flexible way. Other kinds of resources such as FrameNet (Baker, Fillmore, and Lowe 1998) and ConceptNet (Speer, Chin, and Havasi 2017) present the same issue, while returning different types and degrees of structural semantic information and disambiguation capabilities.

    2In this contribution, we provide a novel methodology for the retrieval and representation of unbiased and naturally disambiguated lexical information that relies on a multilingual word alignment mechanism. In particular, we exploit textual resources in different languages1 in order to acquire and align varied lexical-semantic material of the form <target-concept, {related words}k> that are common and shared by all the k languages involved. As we demonstrate through a simple implementation, our method allows to create new lexical-semantic relations between words that are not always available in other resources, as well as to perform an automatic word sense disambiguation process. This system therefore enhances the encoding of prototypical semantic information of concepts that is also likely to be free from strong cultural-linguistic and lexicographic biases.

    3The benefits provided by our novel multilingual word alignment mechanism are thus fourfold: (i) a linguistic and lexicographic de-biasing of lexical knowledge; (ii) naturally-disambiguated aligned lexical entries; (iii) the discovery of novel lexical-semantic relations; and (iv) the representation of prototypical semantic information of concepts in different languages.

    2. Background and Related Work

    2.1 Bias Types

    4Due to its complex and fluid nature, lexical semantics needs to undergo a process of abstraction and simplification in order to be encoded into a formal model. As a result, lexical knowledge provided by lexical resources - especially when monolingual - will inherently carry different types of biases. In particular, i) linguistic and ii) lexicographic biases affect the encoding, consumption, and exploitation of lexical knowledge in downstream tasks.

    Linguistic bias

    5Lexical information encoded in a language’s lexicon, as well as the potential contexts in which a given lexeme can occur, inevitably reflect the socio-cultural background of the speakers of that language. Lexical resources used for the compilation of lexical knowledge are often conceived as monolingual, therefore they mostly return culture-bounded semantic information which does not account for more shared knowledge.

    Lexicographic bias

    6The nuclear components extracted from textual definitions can be different depending on the resource used, even within a single language (Kiefer, n.d.). For example, the definition of “cow” reported by the Oxford Dictionary is “a large animal kept on farms to produce milk or beef” while the Merriam-Webster Dictionary reports “the mature female of cattle”. Both endogenous and exogenous properties can be subjectively reported (Woods 1975), such as the term “large” and the milk production respectively.

    2.2 Related Work

    7On one side, lexicons are built on top of synsets2 and contextualize meanings (or senses) mainly in terms of paradigmatic relations. WordNet (Miller 1995) and BabelNet (Navigli and Ponzetto 2010) can be seen as the cornerstone and the summit in that respect. However, if on the one hand WordNet’s dense network of taxonomic relationships allows a high degree of systematization, on the other hand, a key unsolved issue with “wordnets” is the fine granularity of their inventories. Note that multilingualism in BabelNet is provided as an indexing service rather than as an alignment and unbiasing systematization method.

    8Extensions of these resources also include Common-Sense Knowledge (CSK), which refers to some (to a certain extent) widely-accepted and shared information. CSK describes the kind of general knowledge material that humans use to define, differentiate and reason about the conceptualizations they have in mind (Ruggeri, Di Caro, and Boella 2019). ConceptNet (Speer, Chin, and Havasi 2017) is one of the largest CSK resources, collecting and automatically integrating data starting from the original MIT Open Mind Common Sense project3. However, terms in ConceptNet are not disambiguated. Property norms (McRae et al. 2005; Devereux et al. 2014) represent a similar kind of resource, which is more focused on the cognitive and perception-based aspects of word meaning. Norms, in contrast with ConceptNet, are based on semantic features empirically-constructed via questionnaires producing lexical (often ambiguous) labels associated with target concepts, without any systematic methodology of knowledge collection and encoding.

    9Another widespread modeling approach is based on vector space models of lexical knowledge. Vectors are automatically learnt from large corpora utilizing a wide range of statistical techniques, all centered on Harris’ distributional assumption (Harris 1954), i.e. words that occur in the same contexts tend to have similar meanings. Well-known models include word embeddings (Mikolov et al. 2013; Pennington, Socher, and Manning 2014; Bojanowski et al. 2016), sense embeddings (Huang et al. 2012; Iacobacci, Pilehvar, and Navigli 2015; Kumar et al. 2019), and contextualized embeddings (Scarlini, Pasini, and Navigli 2020). However, the relations holding between vector representations are not typed, nor are they organized systematically.

    10Among the several other modeling strategies proposed, lexicographic-centered resources have been focused on the contextualization of lexical items within syntactic structures, e.g. Corpus Pattern Analysis (CPA) (Hanks 2004), situation frames such as FrameNet (Fillmore 1977; Baker, Fillmore, and Lowe 1998) and conceptual frames (Moerdijk, Tiberius, and Niestadt 2008; Leone et al. 2020). Words are not taken in isolation and the meaning they are attributed is connected to prototypical patterns or typed slots. However, these theories and methods for building semantic resources remain linked to the lexical basis and do not manage the mentioned biases.

    3. The Multilingual Word Alignment

    11As is known, a single word form can be associated with more than one related sense, causing what is referred to as semantic ambiguity, or polysemy. This phenomenon, however, manifests itself differently across languages, since each language encodes meaning into words in its own particular way. We can therefore assume that, while a given polysemous word may be ambiguous in a certain context, a semantically corresponding word in another language will possibly not. Based on this assumption, it is possible to exploit this cross-language property to disambiguate a given word using its semantic equivalent in another language when they both occur in the same context. Such disambiguation process can take place because the two words feature different semantic - specifically, polysemous - behaviours. Accordingly, we developed a knowledge acquisition methodology that features the power of word sense disambiguation, relying on a multilingual <target-concept, {related words}k> alignment mechanism.

    12After providing a brief illustration of the languages we have selected for this first trial, we describe more in detail the methodology by using a basic example. Afterwards, a simple implementation of the proposed mechanism is presented.

    3.1 Languages Involved

    13Among the benefits provided by the multilingual word alignment methodology we propose, one is that it prevents the represented lexical information from containing strong cultural-linguistic biases. This objective is pursued through the use of three different languages, reflecting in turn three diverse backgrounds. For this first trial we involved English, German and Italian. These languages were chosen primarily because we are proficient in them, therefore we are able to exert control over the data of our trial, as well as to interpret the results properly. Concurrently, given the nature of the methodology, it was necessary to select a set of languages with a certain degree of similarity in terms of shared lexical-semantic material. Indeed, the alignment mechanism can work and be effective as long as the lexical-semantic systems of the languages involved reflect a somewhat similar cultural-linguistic background. For example, we might expect languages to agree on the meanings of “carp”, “cottage” and “sled” as long as speakers of these languages have comparable exposure to the relevant data. We would not expect a language spoken in a place without carps to have a word corresponding to “carp”. The purpose of this project is not to forcibly identify universally valid semantic relationships, rather to not report biased information deriving from the use of data coming from a single linguistic context. For this reason, in our case the choice fell on European languages4 (two Germanic languages and a Romance one).

    3.2 Method

    14We now describe in detail the alignment mechanism through a basic example. Consider the following word forms: wool (EN); Wolle (DE); lana (IT), expressing a single target concept5.

    15For each of the three lexical forms we collect a set of related words in terms of paradigmatic (e.g. synonyms) and syntagmatic (e.g. co-occurrences) relations. The target-related words can possibly be modifiers, verbs, or substantives. We thus obtain three different lists of words, one for each of the languages involved. The retrieved terms in the lists are still potentially ambiguous, since they refer to a lexical form rather than to a contextually defined concept. Table 1 provides a small excerpt of such unordered lists of related words.

    Table 1: Unordered lists of single-language related words for <wool (EN), Wolle (DE), lana (IT)>.

    wool

    Wolle

    lana

    sheep

    Schal

    cotone

    cotton

    spinnen

    Biella

    synthetic

    Baumwolle

    sintetica

    spin

    Rudolf

    sciarpa

    scarf

    synthetisch

    pecora

    mitten

    Schafe

    filare

    16The lexical data in the lists are subsequently compared and filtered in order to select only the semantic items that occur in all the lists, i.e., those shared by the three languages6, in the reported example. The resulting words are thus aligned with their semantic counterparts, generating a set of aligned triplets, as shown in Table 2.

    Table 2: Examples of aligned concept-related words for <wool (EN), Wolle (DE), lana (IT)>.

    wool

    Wolle

    lana

    sheep

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    Schafe

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    pecora

    cotton

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    Baumwolle

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    cotone

    synthetic

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    syntetisch

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    sintetica

    spin

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    spinnen

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    filare

    scarf

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    Schal

    Image 10000000000000110000000B332D3458BC368F0C.jpg

    sciarpa

    17This multilingual word alignment provides, as a consequence, an automatic Word Sense Disambiguation system. Once the triplets are formed, their members will be indeed associated with a likely unique sense, i.e. the one coming from the intersection of all possible language-specific senses related to the three words. In other terms, the target-related words, once aligned, naturally identify (and provide) a common semantic context. As a consequence, potentially polysemous words are disambiguated through such context, without any support from sense repositories. For example, the context-consistent sense of the verb to spin (EN), which is a highly polysemous word in English, can be identified by selecting the only sense that is also shared by the other two aligned words, i.e. “turn fibres into thread”. In fact, neither spinnen (DE) nor filare (IT) can possibly mean e.g. “rotate”.

    18This mechanism generates a twofold effect: besides performing word sense disambiguation, it also provides lexical knowledge in the form of (paradigmatic and syntagmatic) lexical-semantic relations between words that is also language-unbounded. In the first place, the uncontrolled character of the data retrieval and alignment process offers the generation of novel lexical-semantic relations that are likely not available in other structured resources. Additionally, since the resulting set of words related to the target can be only the one shared by multiple languages, the lexical knowledge it encodes does not reflect a single cultural/linguistic background, rather a common and shared one. For example, in Table 1 the presence of the word “Biella” among the list of words related to “lana”, probably refers to the fact that the Italian city Biella is (locally) famous for its wool, therefore the two words may co-occur frequently. Similarly, if we consider the alignment <cat (EN), Katze (DE), gatto (IT)>, a lexeme related to the English word form would be “rain”, due to the well-known idiom “it’s raining cats and dogs”. However, neither “Biella” nor corresponding words for “rain” can possibly result in the lists of related words of the respective other languages, being language-specific items within those contexts. Therefore, the lexical information provided by the alignment mechanism will be free from strong cultural-linguistic biases. Finally, as illustrated in the next section, by exploiting multiple and differently built resources, we are able to reduce arbitrariness and lexicographic biases within the lexical knowledge represented.

    4. Implementation

    19In this section we describe details and results of a simple implementation of the proposed alignment mechanism for the acquisition of disambiguated and unbiased lexical information. In particular, the system is composed of two main modules: a context generation and an alignment procedure. We finally report the results of an evaluation to highlight mainly (i) the autonomous disambiguation power of the approach, (ii) the quality of the alignments and their unbiased and syntagmatic nature, and (iii) the amount of unveiled lexical-semantic relations not covered by existing state-of-the-art resources such as BabelNet.

    Table 3: 10 automatic alignments (out of 74) for the target concept <scale (EN), bilancia (IT), Waage (DE)> (BabelNet synset:00069470n)

    POS

    scale

    bilancia

    Waage

    noun

    accuracy

    precisione

    Genauigkeit

    noun

    balance

    equilibrio

    Balance

    noun

    bulk

    massa

    Masse

    noun

    control

    controllo

    Kontrolle

    noun

    device

    dispositivo

    Gerät

    noun

    figure

    cifra

    Zahl

    adj

    accurate

    preciso

    genau

    adj

    smart

    intelligente

    intelligent

    verb

    indicate

    indicare

    zeigen

    verb

    set

    regolare

    einstellen

    4.1 Context for Multilingual Alignment

    20To retrieve the concept-related words for the multilingual alignment we made use of two textual resources: Sketch Engine (Kilgarriff et al. 2014) and the Leipzig Corpora Collection (Quasthoff, Goldhahn, and Eckart 2014). Through the former, we searched for related words with its tool named “Word Sketch” on the TenTen Corpus Family7. In particular, we were able to automatically collect words appearing in the following grammatical relations: “modifiers of w”, “adj. predicates of w”, “verbs with w as subject” and “verbs with w as object”. The retrieved concept-related words are then lemmatized and marked with the suitable POS tags. Finally, we utilized the Leipzig Corpora Collection portal for searching additional context words in terms of left and right (POS-tagged) co-occurrences.

    4.2 Multilingual Alignment

    21The Google Translate API was used for finding translations of related words in the three languages8. In particular, given a certain term tL1 in a language L1, we opted for retrieving all its possible translations into the other two languages (L2, L3). We then tried to match each translated item with the previously-retrieved sets of related words in L2, L3. Whenever the [tL1 Image 10000000000000110000000B332D3458BC368F0C.jpgtL2]; [tL1 Image 10000000000000110000000B332D3458BC368F0C.jpgtL3] match succeeded, we finally checked any possible [tL2 Image 10000000000000110000000B332D3458BC368F0C.jpgtL3] match. If a [tL1 Image 10000000000000110000000B332D3458BC368F0C.jpgtL2 Image 10000000000000110000000B332D3458BC368F0C.jpgtL3] semantic equivalence occurs, then the alignment can take place. Table 3 shows an excerpt of automatic alignments for the concept scale (bn:00069470n).

    4.3 Evaluation

    22Our aim is not to overcome state-of-the-art resources but rather to incorporate new and unbiased semantic relations from a novel multilingual alignment mechanism. In particular, we wanted to verify to what extent our knowledge acquisition method is able to unveil lexical relations yet uncovered by a state-of-the-art resource (BabelNet).

    23Thus, we first generated sets of related words from BabelNet in order to compare them with those produced and aligned by our (automatized) methodology. In particular, through the BabelNet API, we obtained the English, Italian, and German lexicalizations of the synsets connected to it, together with the words included in their glosses9.

    24As test cases, we randomly picked 500 concepts constituting polysemous words in at least one of the three languages, obtaining non-empty alignments for 456 of them. In Table 4 we report the results of the alignment on six concepts.

    25Despite its limitations, our first implementation of the proposed methodology was able to discover a total of 76,152 multilingual alignments over the 456 concepts, with (on average) more than 80% novel semantic relations with respect to what is currently encoded in BabelNet across the three languages. Still, the extracted data represent mostly unbiased and disambiguated knowledge, leading towards the construction of a new large-scale and multilingual prototypical lexical database.

    5. Conclusions and Future Work

    26In this paper we proposed an original methodology for acquiring and encoding lexical knowledge through a novel yet simple mechanism of multilingual alignment. The aim was to represent varied, disambiguated, and language-unbounded lexical knowledge by minimizing strong linguistic and lexicographic biases. A simple implementation and experimentation on 456 concepts carried to unveil around 76K aligned lexical-semantic features, of which more than 80% resulted new when compared with a current state-of-the-art resource such as BabelNet. Future directions include the use of more languages and large-scale runs over thousands of main concepts (Bentivogli et al. 2004; Di Caro and Ruggeri 2019; Camacho-Collados and Navigli 2017).

    Bibliographie

    Des DOI sont automatiquement ajoutés aux références bibliographiques par Bilbo, l’outil d’annotation bibliographique d’OpenEdition. Ces références bibliographiques peuvent être téléchargées dans les formats APA, Chicago et MLA.

    Format

    Proceedings of the Workshop on Multilingual Linguistic Ressources - MLR ’04. (2004). Presented at the the Workshop. https://doi.org/10.3115/1706238
    Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). (2015). Presented at the Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). https://doi.org/10.3115/v1/p15-1
    Proceedings of the HLT-NAACL Workshop on Computational Lexical Semantics - CLS ’04. (2004). Presented at the the HLT-NAACL Workshop. https://doi.org/10.3115/1596431
    Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). (2014). Presented at the Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). https://doi.org/10.3115/v1/d14-1
    Text Mining. (2014). In C. Biemann & A. Mehler (Eds.), Theory and Applications of Natural Language Processing. Springer International Publishing. https://doi.org/10.1007/978-3-319-12655-5
    Scarlini, B., Pasini, T., & Navigli, R. (2020). SensEmBERT: Context-Enhanced Sense Embeddings for Multilingual Word Sense Disambiguation. Association for the Advancement of Artificial Intelligence (AAAI). https://doi.org/10.1609/aaai.v34i05.6402
    Speer, R., Chin, J., & Havasi, C. (2017). ConceptNet 5.5: An Open Multilingual Graph of General Knowledge. Association for the Advancement of Artificial Intelligence (AAAI). https://doi.org/10.1609/aaai.v31i1.11164
    Woods, W. A. (1988). WHAT’S IN A LINK: Foundations for Semantic Networks. Elsevier. https://doi.org/10.1016/b978-1-4832-1446-7.50014-5
    “Proceedings of the Workshop on Multilingual Linguistic Ressources - MLR ’04”. []. Association for Computational Linguistics, 2004. doi:10.3115/1706238.
    “Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)”. []. Association for Computational Linguistics, 2015. doi:10.3115/v1/p15-1.
    “Proceedings of the HLT-NAACL Workshop on Computational Lexical Semantics - CLS ’04”. []. Association for Computational Linguistics, 2004. doi:10.3115/1596431.
    “Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)”. []. Association for Computational Linguistics, 2014. doi:10.3115/v1/d14-1.
    Biemann, Chris, and Alexander Mehler, eds. Text Mining. Theory and Applications of Natural Language Processing. Springer International Publishing, 2014. doi:10.1007/978-3-319-12655-5.
    Scarlini, Bianca, Tommaso Pasini, and Roberto Navigli. “SensEmBERT: Context-Enhanced Sense Embeddings for Multilingual Word Sense Disambiguation”. Proceedings of the AAAI Conference on Artificial Intelligence. Association for the Advancement of Artificial Intelligence (AAAI), April 3, 2020. doi:10.1609/aaai.v34i05.6402.
    Speer, Robyn, Joshua Chin, and Catherine Havasi. “ConceptNet 5.5: An Open Multilingual Graph of General Knowledge”. Proceedings of the AAAI Conference on Artificial Intelligence. Association for the Advancement of Artificial Intelligence (AAAI), February 12, 2017. doi:10.1609/aaai.v31i1.11164.
    Woods, William A. “WHAT’S IN A LINK: Foundations for Semantic Networks”. Readings in Cognitive Science. Elsevier, 1988. doi:10.1016/b978-1-4832-1446-7.50014-5.
    Proceedings of the Workshop on Multilingual Linguistic Ressources - MLR ’04. [], Association for Computational Linguistics, 2004. Crossref, https://doi.org/10.3115/1706238.
    Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). [], Association for Computational Linguistics, 2015. Crossref, https://doi.org/10.3115/v1/p15-1.
    Proceedings of the HLT-NAACL Workshop on Computational Lexical Semantics - CLS ’04. [], Association for Computational Linguistics, 2004. Crossref, https://doi.org/10.3115/1596431.
    Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). [], Association for Computational Linguistics, 2014. Crossref, https://doi.org/10.3115/v1/d14-1.
    Biemann, Chris, and Alexander Mehler, editors. Text Mining. Theory and Applications of Natural Language Processing, Springer International Publishing, 2014. Crossref, https://doi.org/10.1007/978-3-319-12655-5.
    Scarlini, Bianca, et al. “SensEmBERT: Context-Enhanced Sense Embeddings for Multilingual Word Sense Disambiguation”. Proceedings of the AAAI Conference on Artificial Intelligence, vols. 34, nos. 05, Association for the Advancement of Artificial Intelligence (AAAI), 3 Apr. 2020, pp. 8758-65. Crossref, https://doi.org/10.1609/aaai.v34i05.6402.
    Speer, Robyn, et al. “ConceptNet 5.5: An Open Multilingual Graph of General Knowledge”. Proceedings of the AAAI Conference on Artificial Intelligence, vols. 31, no. 1, Association for the Advancement of Artificial Intelligence (AAAI), 12 Feb. 2017. Crossref, https://doi.org/10.1609/aaai.v31i1.11164.
    Woods, William A. “WHAT’S IN A LINK: Foundations for Semantic Networks”. Readings in Cognitive Science, Elsevier, 1988, pp. 102-25. Crossref, https://doi.org/10.1016/b978-1-4832-1446-7.50014-5.

    Cette bibliographie a été enrichie de toutes les références bibliographiques automatiquement générées par Bilbo en utilisant Crossref.

    Collin F. Baker, Charles J Fillmore, and John B Lowe. 1998. “The Berkeley Framenet Project.” In 36th Annual Meeting of the Association for Computational Linguistics and 17th International Conference on Computational Linguistics, Volume 1, 86–90.

    Luisa Bentivogli, Pamela Forner, Bernardo Magnini, and Emanuele Pianta. 2004. “Revising the Wordnet Domains Hierarchy: Semantics, Coverage and Balancing.” In Proceedings of the Workshop on Multilingual Linguistic Resources, 94–101.

    10.3115/1706238 :

    Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2016. “Enriching Word Vectors with Subword Information.” arXiv Preprint arXiv:1607.04606.

    Jose Camacho-Collados and Roberto Navigli. 2017. “BabelDomains: Large-Scale Domain Labeling of Lexical Resources.” In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, 223–28.

    Barry J. Devereux, Lorraine K Tyler, Jeroen Geertzen, and Billi Randall. 2014. “The Cslb Concept Property Norms.” Behavior Research Methods 46 (4): 1119–27.

    Luigi Di Caro and Alice Ruggeri. 2019. “Unveiling Middle-Level Concepts Through Frequency Trajectories and Peaks Analysis.” In Proceedings of the 34th Acm/Sigapp Symposium on Applied Computing, 1035–42.

    Charles J. Fillmore. 1977. “Scenes-and-Frames Semantics.” Linguistic Structures Processing 59: 55–88.

    Patrick Hanks. 2004. “Corpus Pattern Analysis.” In Euralex Proceedings, 1:87–98.

    Zellig S. Harris. 1954. “Distributional Structure.” Word 10 (2-3): 146–62.

    Eric H. Huang, Richard Socher, Christopher D Manning, and Andrew Y Ng. 2012. “Improving Word Representations via Global Context and Multiple Word Prototypes.” In Proc. Of Acl, 873–82.

    Ignacio Iacobacci, Mohammad Taher Pilehvar, and Roberto Navigli. 2015. “SensEmbed: Learning Sense Embeddings for Word and Relational Similarity.” In Proceedings of Acl, 95–105.

    10.3115/v1/P15-1 :

    Ferenc Kiefer. n.d. “Linguistic, Conceptual and Encyclopedic Knowledge: Some Implications for Lexicography.” In Proceedings of the 3rd Euralex International Congress, edited by T. Magay and J. Zigány, 1–10. Budapest, Hungary: Akadémiai Kiadó.

    Adam Kilgarriff, Vít Baisa, Jan Bušta, Miloš Jakubíček, Vojtěch Kovář, Jan Michelfeit, Pavel Rychlý, and Vít Suchomel. 2014. “The Sketch Engine: Ten Years on.” The Lexicography 1(1): 7–36.

    Sawan Kumar, Sharmistha Jat, Karan Saxena, and Partha Talukdar. 2019. “Zero-Shot Word Sense Disambiguation Using Sense Definition Embeddings.” In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 5670–81.

    Valentina Leone, Giovanni Siragusa, Luigi Di Caro, and Roberto Navigli. 2020. “Building Semantic Grams of Human Knowledge.” In Proceedings of the 12th Language Resources and Evaluation Conference, 2991–3000.

    Ken McRae, George S Cree, Mark S Seidenberg, and Chris McNorgan. 2005. “Semantic Feature Production Norms for a Large Set of Living and Nonliving Things.” Behav. R. M. 37 (4): 547–59.

    Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. “Distributed Representations of Words and Phrases and Their Compositionality.” In Advances in Neural Information Processing Systems, 3111–9.

    George A. Miller. 1995. “WordNet: A Lexical Database for English.” Communications of the ACM 38 (11): 39–41.

    Fons Moerdijk, Carole Tiberius, and Jan Niestadt. 2008. “Accessing the Anw Dictionary.” In Proc. Of the Workshop on Cognitive Aspects of the Lexicon, 18–24.

    Jane Morris and Graeme Hirst. 2004. “Non-Classical Lexical Semantic Relations.” In Proceedings of the Computational Lexical Semantics Workshop at HLT-NAACL 2004, 46–51. Boston, Massachusetts, USA: Association for Computational Linguistics. https://aclanthology.org/W04-2607.

    10.3115/1596431 :

    Roberto Navigli and Simone Paolo Ponzetto. 2010. “BabelNet: Building a Very Large Multilingual Semantic Network.” In Proc. Of Acl, 216–25. Association for Computational Linguistics.

    Martha Palmer, Hoa Trang Dang, and Christiane Fellbaum. 2007. “Making Fine-Grained and Coarse-Grained Sense Distinctions, Both Manually and Automatically.” Nat.Lan.Eng. 13 (02): 137–63.

    Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. “Glove: Global Vectors for Word Representation.” In EMNLP, 14:1532–43.

    10.3115/v1/D14-1 :

    Uwe Quasthoff, Dirk Goldhahn, and Thomas Eckart. 2014. “Building Large Resources for Text Mining: The Leipzig Corpora Collection.” In Text Mining, 3–24. Springer.

    10.1007/978-3-319-12655-5 :

    Alice Ruggeri, Luigi Di Caro, and Guido Boella. 2019. “The Role of Common-Sense Knowledge in Assessing Semantic Association.” Journal on Data Semantics 8 (1): 39–56.

    Bianca Scarlini, Tommaso Pasini, and Roberto Navigli. 2020. “SensEmBERT: Context-Enhanced Sense Embeddings for Multilingual Word Sense Disambiguation.” In Proceedings of the 34th Conference on Artificial Intelligence. Association for the Advancement of Artificial Intelligence.

    10.1609/aaai.v34i05.6402 :

    Robert Speer, Joshua Chin, and Catherine Havasi. 2017. “Conceptnet 5.5: An Open Multilingual Graph of General Knowledge.” In Thirty-First Aaai Conference on Artificial Intelligence.

    10.1609/aaai.v31i1.11164 :

    William A. Woods. 1975. “What’s in a Link: Foundations for Semantic Networks.” In Representation and Understanding, 35–82. Elsevier.

    10.1016/B978-1-4832-1446-7.50014-5 :

    Michael Zock and Chris Biemann. 2020. “Comparison of Different Lexical Resources with Respect to the Tip-of-the-Tongue Problem.” Journal of Cognitive Science 21 (2): 193–252.

    Notes de bas de page

    1 In this work, we start with the combination of three languages: English, German and Italian.

    2 Words considered as synonyms in specific contexts.

    3 https://www.media.mit.edu/projects/open-mind-common-sense/overview/

    4 By “European” we refer to the European linguistic area.

    5 An absolute monosemy is, of course, realistically unreachable.

    6 This implies the presence of a translation step.

    7 https://www.sketchengine.eu/documentation/tenten-corpora

    8 No surrounding syntactic context for the words to align was available for more advanced Machine Translation.

    9 We used the SpaCy library to analyze, extract and lemmatize the text - https://spacy.io.

    Auteurs

    • Francesca Grasso

      University of Turin, Department of Computer Science – fr.grasso@unito.it

    • Luigi Di Caro

      University of Turin, Department of Computer Science – luigi.dicaro@unito.it

    Précédent Suivant
    Table des matières

    Creative Commons - Attribution - Pas d'Utilisation Commerciale - Pas de Modification 4.0 International - CC BY-NC-ND 4.0

    Le texte seul est utilisable sous licence Creative Commons - Attribution - Pas d'Utilisation Commerciale - Pas de Modification 4.0 International - CC BY-NC-ND 4.0. Les autres éléments (illustrations, fichiers annexes importés) sont « Tous droits réservés », sauf mention contraire.

    Voir plus de livres
    Proceedings of the Second Italian Conference on Computational Linguistics CLiC-it 2015

    Proceedings of the Second Italian Conference on Computational Linguistics CLiC-it 2015

    3-4 December 2015, Trento

    Cristina Bosco, Sara Tonelli et Fabio Massimo Zanzotto (dir.)

    2015

    Proceedings of the Third Italian Conference on Computational Linguistics CLiC-it 2016

    Proceedings of the Third Italian Conference on Computational Linguistics CLiC-it 2016

    5-6 December 2016, Napoli

    Anna Corazza, Simonetta Montemagni et Giovanni Semeraro (dir.)

    2016

    EVALITA. Evaluation of NLP and Speech Tools for Italian

    EVALITA. Evaluation of NLP and Speech Tools for Italian

    Proceedings of the Final Workshop 7 December 2016, Naples

    Pierpaolo Basile, Franco Cutugno, Malvina Nissim et al. (dir.)

    2016

    Proceedings of the Fourth Italian Conference on Computational Linguistics CLiC-it 2017

    Proceedings of the Fourth Italian Conference on Computational Linguistics CLiC-it 2017

    11-12 December 2017, Rome

    Roberto Basili, Malvina Nissim et Giorgio Satta (dir.)

    2017

    Proceedings of the Fifth Italian Conference on Computational Linguistics CLiC-it 2018

    Proceedings of the Fifth Italian Conference on Computational Linguistics CLiC-it 2018

    10-12 December 2018, Torino

    Elena Cabrio, Alessandro Mazzei et Fabio Tamburini (dir.)

    2018

    EVALITA Evaluation of NLP and Speech Tools for Italian

    EVALITA Evaluation of NLP and Speech Tools for Italian

    Proceedings of the Final Workshop 12-13 December 2018, Naples

    Tommaso Caselli, Nicole Novielli, Viviana Patti et al. (dir.)

    2018

    EVALITA Evaluation of NLP and Speech Tools for Italian - December 17th, 2020

    EVALITA Evaluation of NLP and Speech Tools for Italian - December 17th, 2020

    Proceedings of the Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian Final Workshop

    Valerio Basile, Danilo Croce, Maria Maro et al. (dir.)

    2020

    Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020

    Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020

    Bologna, Italy, March 1-3, 2021

    Felice Dell'Orletta, Johanna Monti et Fabio Tamburini (dir.)

    2020

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Milan, Italy, 26-28 January, 2022

    Elisabetta Fersini, Marco Passarotti et Viviana Patti (dir.)

    2022

    Voir plus de livres
    1 / 9
    Proceedings of the Second Italian Conference on Computational Linguistics CLiC-it 2015

    Proceedings of the Second Italian Conference on Computational Linguistics CLiC-it 2015

    3-4 December 2015, Trento

    Cristina Bosco, Sara Tonelli et Fabio Massimo Zanzotto (dir.)

    2015

    Proceedings of the Third Italian Conference on Computational Linguistics CLiC-it 2016

    Proceedings of the Third Italian Conference on Computational Linguistics CLiC-it 2016

    5-6 December 2016, Napoli

    Anna Corazza, Simonetta Montemagni et Giovanni Semeraro (dir.)

    2016

    EVALITA. Evaluation of NLP and Speech Tools for Italian

    EVALITA. Evaluation of NLP and Speech Tools for Italian

    Proceedings of the Final Workshop 7 December 2016, Naples

    Pierpaolo Basile, Franco Cutugno, Malvina Nissim et al. (dir.)

    2016

    Proceedings of the Fourth Italian Conference on Computational Linguistics CLiC-it 2017

    Proceedings of the Fourth Italian Conference on Computational Linguistics CLiC-it 2017

    11-12 December 2017, Rome

    Roberto Basili, Malvina Nissim et Giorgio Satta (dir.)

    2017

    Proceedings of the Fifth Italian Conference on Computational Linguistics CLiC-it 2018

    Proceedings of the Fifth Italian Conference on Computational Linguistics CLiC-it 2018

    10-12 December 2018, Torino

    Elena Cabrio, Alessandro Mazzei et Fabio Tamburini (dir.)

    2018

    EVALITA Evaluation of NLP and Speech Tools for Italian

    EVALITA Evaluation of NLP and Speech Tools for Italian

    Proceedings of the Final Workshop 12-13 December 2018, Naples

    Tommaso Caselli, Nicole Novielli, Viviana Patti et al. (dir.)

    2018

    EVALITA Evaluation of NLP and Speech Tools for Italian - December 17th, 2020

    EVALITA Evaluation of NLP and Speech Tools for Italian - December 17th, 2020

    Proceedings of the Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian Final Workshop

    Valerio Basile, Danilo Croce, Maria Maro et al. (dir.)

    2020

    Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020

    Proceedings of the Seventh Italian Conference on Computational Linguistics CLiC-it 2020

    Bologna, Italy, March 1-3, 2021

    Felice Dell'Orletta, Johanna Monti et Fabio Tamburini (dir.)

    2020

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Milan, Italy, 26-28 January, 2022

    Elisabetta Fersini, Marco Passarotti et Viviana Patti (dir.)

    2022

    Voir plus de chapitres

    From a Lexical to a Semantic Distributional Hypothesis

    Luigi Di Caro, Guido Boella, Alice Ruggeri et al.

    Voir plus de chapitres

    From a Lexical to a Semantic Distributional Hypothesis

    Luigi Di Caro, Guido Boella, Alice Ruggeri et al.

    Accès ouvert

    Accès ouvert

    ePub

    PDF

    PDF du chapitre

    1 In this work, we start with the combination of three languages: English, German and Italian.

    2 Words considered as synonyms in specific contexts.

    3 https://www.media.mit.edu/projects/open-mind-common-sense/overview/

    4 By “European” we refer to the European linguistic area.

    5 An absolute monosemy is, of course, realistically unreachable.

    6 This implies the presence of a translation step.

    7 https://www.sketchengine.eu/documentation/tenten-corpora

    8 No surrounding syntactic context for the words to align was available for more advanced Machine Translation.

    9 We used the SpaCy library to analyze, extract and lemmatize the text - https://spacy.io.

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    X Facebook Email

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Ce livre est cité par

    • Daraghmi, Eman. Atwe, Lour. Jaber, Areej. (2025) A Comparative Study of PEGASUS, BART, and T5 for Text Summarization Across Diverse Datasets. Future Internet, 17. DOI: 10.3390/fi17090389

    Ce chapitre est cité par

    • Grasso, Francesca. Rulfi, Vladimiro Lovera. Caro, Luigi Di. (2024) Communications in Computer and Information Science Metadata and Semantic Research. DOI: 10.1007/978-3-031-65990-4_1
    • Grasso, Francesca. Lovera Rulfi, Vladimiro. Di Caro, Luigi. (2022) Lecture Notes in Computer Science Knowledge Engineering and Knowledge Management. DOI: 10.1007/978-3-031-17105-5_3

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Vous pouvez vous connecter à votre bibliothèque à l’adresse suivante : https://freemium.openedition.org/oebooks

    Suggérer l’acquisition à votre bibliothèque

    Si vous avez des questions, vous pouvez nous écrire à access[at]openedition.org

    Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021

    Vérifiez si votre bibliothèque a déjà acquis ce livre : authentifiez-vous à OpenEdition Freemium for Books.

    Vous pouvez suggérer à votre bibliothèque d’acquérir un ou plusieurs livres publiés sur OpenEdition Books. N’hésitez pas à lui indiquer nos coordonnées : access[at]openedition.org

    Vous pouvez également nous indiquer, à l’aide du formulaire suivant, les coordonnées de votre bibliothèque afin que nous la contactions pour lui suggérer l’achat de ce livre. Les champs suivis de (*) sont obligatoires.

    Veuillez, s’il vous plaît, remplir tous les champs.

    La syntaxe de l’email est incorrecte.

    Référence numérique du chapitre

    Format

    Grasso, F., & Di Caro, L. (2022). A Methodology for Large-Scale, Disambiguated and Unbiased Lexical Knowledge Acquisition Based on Multilingual Word Alignment. In E. Fersini, M. Passarotti, & V. Patti (éds.), Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021. Torino: Accademia University Press. https://doi.org/10.4000/books.aaccademia.10653
    Grasso, Francesca, et Luigi Di Caro. « A Methodology for Large-Scale, Disambiguated and Unbiased Lexical Knowledge Acquisition Based on Multilingual Word Alignment ». In Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-It 2021, édité par Elisabetta Fersini, Marco Passarotti, et Viviana Patti. Torino: Accademia University Press, 2022. doi:10.4000/books.aaccademia.10653.
    Grasso, Francesca, et Luigi Di Caro. « A Methodology for Large-Scale, Disambiguated and Unbiased Lexical Knowledge Acquisition Based on Multilingual Word Alignment ». Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-It 2021, édité par Elisabetta Fersini et al., Accademia University Press, 2022, https://doi.org/10.4000/books.aaccademia.10653.

    Référence numérique du livre

    Format

    Fersini, E., Passarotti, M., & Patti, V. (éds.). (2022). Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-it 2021. Torino: Accademia University Press. https://doi.org/10.4000/books.aaccademia.10417
    Fersini, Elisabetta, Marco Passarotti, et Viviana Patti, éd. Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-It 2021. Torino: Accademia University Press, 2022. doi:10.4000/books.aaccademia.10417.
    Fersini, Elisabetta, et al., éditeurs. Proceedings of the Eighth Italian Conference on Computational Linguistics CliC-It 2021. Accademia University Press, 2022, https://doi.org/10.4000/books.aaccademia.10417.
    Compatible avec Zotero Zotero

    1 / 3

    Accademia University Press

    Accademia University Press

    • Plan du site
    • Se connecter

    Suivez-nous

    • Facebook
    • Flux RSS

    URL : http://www.aaccademia.it/

    Email : info@aaccademia.it

    OpenEdition
    • Candidater à OpenEdition Books
    • Connaître le programme OpenEdition Freemium
    • Commander des livres
    • S’abonner à la lettre d’OpenEdition
    • CGU d’OpenEdition Books
    • Accessibilité : partiellement conforme
    • Données personnelles
    • Gestion des cookies
    • Système de signalement