连续空间中单词的学习表示可能是NLP中最基本的任务,但是单词以比向量点产品相似性提供的富裕方式相互作用。单词之间的许多关系可以从理论上表达为设置,例如形容词 - 名称化合物(例如“红色汽车” $ \ subseteq $“ Cars”)和同符(例如,“舌头” $ \ cap $应该是与“口”相似,而“舌头” $ \ cap $“语言”应该与“方言”相似)具有自然的理论解释。盒子嵌入是一种新型基于区域的表示,可提供执行这些设定理论操作的能力。在这项工作中,我们提供了对盒子嵌入的模糊集解释,并使用设定理论训练目标学习单词的框表示。我们在各种单词相似性任务上,尤其是在较不常见的单词上表现出改善的性能,并执行定量和定性分析,以探讨Word2box提供的其他独特表达性。
translated by 谷歌翻译
查询嵌入(QE) - 旨在嵌入实体和一阶逻辑(FOL)查询在低维空间中 - 在知识图表中的多跳推理中显示出强大的功率。最近,嵌入实体和具有几何形状的查询成为有希望的方向,因为几何形状可以自然地代表它们之间的答案和逻辑关系。然而,现有的基于几何的模型难以建模否定查询,这显着限制了它们的适用性。为了解决这一挑战,我们提出了一种新型查询嵌入模型,即锥形嵌入式(锥形),即锥形嵌入式(锥形),它是可以处理所有的基于几何的QE模型,包括所有FOL操作,包括结合,分离和否定。具体而言,锥形代表实体和查询作为二维锥体的笛卡尔产品,其中锥体的交叉和联合自然地模拟了结合和分离操作。通过进一步注意到,锥体的补充仍然存在锥体,我们在嵌入空间中设计几何补充运算符进行否定操作。实验表明,锥体在基准数据集上显着优于现有的现有技术。
translated by 谷歌翻译
在大规模不完整的知识图(kgs)上回答复杂的一阶逻辑(fol)查询是一项重要但挑战性的任务。最近的进步将逻辑查询和KG实体嵌入了相同的空间,并通过密集的相似性搜索进行查询。但是,先前研究中设计的大多数逻辑运算符不满足经典逻辑的公理系统,从而限制了其性能。此外,这些逻辑运算符被参数化,因此需要许多复杂的查询作为训练数据,在大多数现实世界中,这些数据通常很难收集甚至无法访问。因此,我们提出了Fuzzqe,这是一种基于模糊逻辑的逻辑查询嵌入框架,用于回答KGS上的查询。 Fuzzqe遵循模糊逻辑以原则性和无学习的方式定义逻辑运算符,在这种方式中,只有实体和关系嵌入才需要学习。 Fuzzqe可以从标记为训练的复杂逻辑查询中进一步受益。在两个基准数据集上进行的广泛实验表明,与最先进的方法相比,Fuzzqe在回答FOL查询方面提供了明显更好的性能。此外,只有KG链接预测训练的Fuzzqe可以实现与经过额外复杂查询数据训练的人的可比性能。
translated by 谷歌翻译
translated by 谷歌翻译
We propose a unified neural network architecture and learning algorithm that can be applied to various natural language processing tasks including: part-of-speech tagging, chunking, named entity recognition, and semantic role labeling. This versatility is achieved by trying to avoid task-specific engineering and therefore disregarding a lot of prior knowledge. Instead of exploiting man-made input features carefully optimized for each task, our system learns internal representations on the basis of vast amounts of mostly unlabeled training data. This work is then used as a basis for building a freely available tagging system with good performance and minimal computational requirements.
translated by 谷歌翻译
基于变压器的语言模型最近在许多自然语言任务中取得了显着的结果。但是,通常通过利用大量培训数据来实现排行榜的性能,并且很少通过将明确的语言知识编码为神经模型。这使许多人质疑语言学对现代自然语言处理的相关性。在本文中,我介绍了几个案例研究,以说明理论语言学和神经语言模型仍然相互关联。首先,语言模型通过提供一个客观的工具来测量语义距离,这对语言学家很有用,语义距离很难使用传统方法。另一方面,语言理论通过提供框架和数据源来探究我们的语言模型,以了解语言理解的特定方面,从而有助于语言建模研究。本论文贡献了三项研究,探讨了语言模型中语法 - 听觉界面的不同方面。在论文的第一部分中,我将语言模型应用于单词类灵活性的问题。我将Mbert作为语义距离测量的来源,我提供了有利于将单词类灵活性分析为方向过程的证据。在论文的第二部分中,我提出了一种方法来测量语言模型中间层的惊奇方法。我的实验表明,包含形态句法异常的句子触发了语言模型早期的惊喜,而不是语义和常识异常。最后,在论文的第三部分中,我适应了一些心理语言学研究,以表明语言模型包含了论证结构结构的知识。总而言之,我的论文在自然语言处理,语言理论和心理语言学之间建立了新的联系,以为语言模型的解释提供新的观点。
translated by 谷歌翻译
translated by 谷歌翻译
Natural Language Understanding has seen an increasing number of publications in the last few years, especially after robust word embeddings models became prominent, when they proved themselves able to capture and represent semantic relationships from massive amounts of data. Nevertheless, traditional models often fall short in intrinsic issues of linguistics, such as polysemy and homonymy. Any expert system that makes use of natural language in its core, can be affected by a weak semantic representation of text, resulting in inaccurate outcomes based on poor decisions. To mitigate such issues, we propose a novel approach called Most Suitable Sense Annotation (MSSA), that disambiguates and annotates each word by its specific sense, considering the semantic effects of its context. Our approach brings three main contributions to the semantic representation scenario: (i) an unsupervised technique that disambiguates and annotates words by their senses, (ii) a multi-sense embeddings model that can be extended to any traditional word embeddings algorithm, and (iii) a recurrent methodology that allows our models to be re-used and their representations refined. We test our approach on six different benchmarks for the word similarity task, showing that our approach can produce state-of-the-art results and outperforms several more complex state-of-the-art systems.
translated by 谷歌翻译
Recent methods for learning vector space representations of words have succeeded in capturing fine-grained semantic and syntactic regularities using vector arithmetic, but the origin of these regularities has remained opaque. We analyze and make explicit the model properties needed for such regularities to emerge in word vectors. The result is a new global logbilinear regression model that combines the advantages of the two major model families in the literature: global matrix factorization and local context window methods. Our model efficiently leverages statistical information by training only on the nonzero elements in a word-word cooccurrence matrix, rather than on the entire sparse matrix or on individual context windows in a large corpus. The model produces a vector space with meaningful substructure, as evidenced by its performance of 75% on a recent word analogy task. It also outperforms related models on similarity tasks and named entity recognition.
translated by 谷歌翻译
translated by 谷歌翻译
translated by 谷歌翻译
Continuous word representations, trained on large unlabeled corpora are useful for many natural language processing tasks. Popular models that learn such representations ignore the morphology of words, by assigning a distinct vector to each word. This is a limitation, especially for languages with large vocabularies and many rare words. In this paper, we propose a new approach based on the skipgram model, where each word is represented as a bag of character n-grams. A vector representation is associated to each character n-gram; words being represented as the sum of these representations. Our method is fast, allowing to train models on large corpora quickly and allows us to compute word representations for words that did not appear in the training data. We evaluate our word representations on nine different languages, both on word similarity and analogy tasks. By comparing to recently proposed morphological word representations, we show that our vectors achieve state-of-the-art performance on these tasks.
translated by 谷歌翻译
translated by 谷歌翻译
We present a model that generates natural language descriptions of images and their regions. Our approach leverages datasets of images and their sentence descriptions to learn about the inter-modal correspondences between language and visual data. Our alignment model is based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding. We then describe a Multimodal Recurrent Neural Network architecture that uses the inferred alignments to learn to generate novel descriptions of image regions. We demonstrate that our alignment model produces state of the art results in retrieval experiments on Flickr8K, Flickr30K and MSCOCO datasets. We then show that the generated descriptions significantly outperform retrieval baselines on both full images and on a new dataset of region-level annotations.
translated by 谷歌翻译
translated by 谷歌翻译
Machine Learning for Source Code (ML4Code) is an active research field in which extensive experimentation is needed to discover how to best use source code's richly structured information. With this in mind, we introduce JEMMA, an Extensible Java Dataset for ML4Code Applications, which is a large-scale, diverse, and high-quality dataset targeted at ML4Code. Our goal with JEMMA is to lower the barrier to entry in ML4Code by providing the building blocks to experiment with source code models and tasks. JEMMA comes with a considerable amount of pre-processed information such as metadata, representations (e.g., code tokens, ASTs, graphs), and several properties (e.g., metrics, static analysis results) for 50,000 Java projects from the 50KC dataset, with over 1.2 million classes and over 8 million methods. JEMMA is also extensible allowing users to add new properties and representations to the dataset, and evaluate tasks on them. Thus, JEMMA becomes a workbench that researchers can use to experiment with novel representations and tasks operating on source code. To demonstrate the utility of the dataset, we also report results from two empirical studies on our data, ultimately showing that significant work lies ahead in the design of context-aware source code models that can reason over a broader network of source code entities in a software project, the very task that JEMMA is designed to help with.
translated by 谷歌翻译
The ability to compare the semantic similarity between text corpora is important in a variety of natural language processing applications. However, standard methods for evaluating these metrics have yet to be established. We propose a set of automatic and interpretable measures for assessing the characteristics of corpus-level semantic similarity metrics, allowing sensible comparison of their behavior. We demonstrate the effectiveness of our evaluation measures in capturing fundamental characteristics by evaluating them on a collection of classical and state-of-the-art metrics. Our measures revealed that recently-developed metrics are becoming better in identifying semantic distributional mismatch while classical metrics are more sensitive to perturbations in the surface text levels.
translated by 谷歌翻译
最近,越来越多的努力用于学习符号知识库(KB)的持续表示。但是,这些方法要么仅嵌入数据级知识(ABOX),要么在处理概念级知识(Tbox)时受到固有的局限性,即它们不能忠实地对KBS中存在的逻辑结构进行建模。我们提出了Boxel,这是一种几何KB嵌入方法,可以更好地捕获描述逻辑EL ++中的逻辑结构(即Abox和Tbox Axioms)。 Boxel模型在Kb中作为轴平行框,适用于建模概念交叉点,作为点内部的实体以及概念/实体之间的关系作为仿射转换。我们展示了Boxel的理论保证(声音),以保存逻辑结构。也就是说,有损耗0的框嵌入模型是KB​​的(逻辑)模型。实验结果(合理)补充推理和用于蛋白质 - 蛋白质预测的现实世界应用的结果表明,Boxel的表现优于传统知识图嵌入方法以及最先进的EL ++嵌入方法。
translated by 谷歌翻译
For natural language understanding (NLU) technology to be maximally useful, it must be able to process language in a way that is not exclusive to a single task, genre, or dataset. In pursuit of this objective, we introduce the General Language Understanding Evaluation (GLUE) benchmark, a collection of tools for evaluating the performance of models across a diverse set of existing NLU tasks. By including tasks with limited training data, GLUE is designed to favor and encourage models that share general linguistic knowledge across tasks. GLUE also includes a hand-crafted diagnostic test suite that enables detailed linguistic analysis of models. We evaluate baselines based on current methods for transfer and representation learning and find that multi-task training on all tasks performs better than training a separate model per task. However, the low absolute performance of our best model indicates the need for improved general NLU systems.
translated by 谷歌翻译
translated by 谷歌翻译