4.3CLJun 23, 2022
AST-Probe: Recovering abstract syntax trees from hidden representations of pre-trained language modelsJosé Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado et al.
The objective of pre-trained language models is to learn contextual representations of textual data. Pre-trained language models have become mainstream in natural language processing and code modeling. Using probes, a technique to study the linguistic properties of hidden vector spaces, previous works have shown that these pre-trained language models encode simple linguistic properties in their hidden representations. However, none of the previous work assessed whether these models encode the whole grammatical structure of a programming language. In this paper, we prove the existence of a syntactic subspace, lying in the hidden representations of pre-trained language models, which contain the syntactic information of the programming language. We show that this subspace can be extracted from the models' representations and define a novel probing method, the AST-Probe, that enables recovering the whole abstract syntax tree (AST) of an input code snippet. In our experimentations, we show that this syntactic subspace exists in five state-of-the-art pre-trained language models. In addition, we highlight that the middle layers of the models are the ones that encode most of the AST information. Finally, we estimate the optimal size of this syntactic subspace and show that its dimension is substantially lower than those of the models' representation spaces. This suggests that pre-trained language models use a small part of their representation spaces to encode syntactic information of the programming languages.
0.8SEJun 28
Connecting the Models: A Global Mega-model of MDE Projects on GitHubJesús Sánchez Cuadrado
A key element of Model-Driven Engineering is the construction of domain-specific modelling environments to improve productivity and quality. In theory, dedicated technologies like EMF, ATL, Epsilon, Xtext, etc. would boost the construction of high-quality environments with a relatively modest effort by chaining the output of one tool to the input of another. However, there is little empirical evidence of how this idea has fared in reality and many open research questions remain, such as how MDE tools are used and combined, whether the resulting environments are maintained or not, which tools are used more frequently, etc. In this paper, we aim to build a foundation for studying how MDE is used in practice. First, we constructed a dataset by mining 7,436 Github projects comprising over 325,000 MDE artefacts. These artefacts encompass representative Eclipse EMF-related technologies, namely Ecore, Emfatic, OCL, ATL, Epsilon, QVTo, Henshin, Acceleo, Xtext, Emftext, GMF and Sirius. We also integrated into the dataset repository-level information extracted from the Git repositories and the GitHub API. From this dataset, we devised a technique to recover the mega-model of each project in order to represent the relationships between its artefacts. Then, we built a global mega-model relating the different MDE projects by performing an analysis of near-duplicates across all artefacts and grouping duplicate artefacts into single nodes and rewiring the connections. This global mega-model can be used to derive additional information like inter-project dependencies or studying connected subgraphs of artefacts. Finally, we propose a number of research questions that could be answered with the provided dataset, which we hope will foster empirical analysis of how MDE is applied.
5.3SEAug 26, 2020
MAR: A structure-based search engine for modelsJosé Antonio Hernández López, Jesús Sánchez Cuadrado
The availability of shared software models provides opportunities for reusing, adapting and learning from them. Public models are typically stored in a variety of locations, including model repositories, regular source code repositories, web pages, etc. To profit from them developers need effective search mechanisms to locate the models relevant for their tasks. However, to date, there has been little success in creating a generic and efficient search engine specially tailored to the modelling domain. In this paper we present MAR, a search engine for models. MAR is generic in the sense that it can index any type of model if its meta-model is known. MAR uses a query-by-example approach, that is, it uses example models as queries. The search takes the model structure into account using the notion of bag of paths, which encodes the structure of a model using paths between model elements and is a representation amenable for indexing. MAR is built over HBase using a specific design to deal with large repositories. Our benchmarks show that the engine is efficient and has fast response times in most cases. We have also evaluated the precision of the search engine by creating model mutants which simulate user queries. A REST API is available to perform queries and an Eclipse plug-in allows end users to connect to the search engine from model editors. We have currently indexed more than 50.000 models of different kinds, including Ecore meta-models, BPMN diagrams and UML models. MAR is available at http://mar-search.org.