Christopher J. MacLellan

AI
h-index7
5papers
30citations
Novelty54%
AI Score29

5 Papers

2.1CLDec 22, 2022
Efficient Induction of Language Models Via Probabilistic Concept Formation

Christopher J. MacLellan, Peter Matsakis, Pat Langley · gatech

This paper presents a novel approach to the acquisition of language models from corpora. The framework builds on Cobweb, an early system for constructing taxonomic hierarchies of probabilistic concepts that used a tabular, attribute-value encoding of training cases and concepts, making it unsuitable for sequential input like language. In response, we explore three new extensions to Cobweb -- the Word, Leaf, and Path variants. These systems encode each training case as an anchor word and surrounding context words, and they store probabilistic descriptions of concepts as distributions over anchor and context information. As in the original Cobweb, a performance element sorts a new instance downward through the hierarchy and uses the final node to predict missing features. Learning is interleaved with performance, updating concept probabilities and hierarchy structure as classification occurs. Thus, the new approaches process training cases in an incremental, online manner that it very different from most methods for statistical language learning. We examine how well the three variants place synonyms together and keep homonyms apart, their ability to recall synonyms as a function of training set size, and their training efficiency. Finally, we discuss related work on incremental learning and directions for further research.

1.9CLSep 19, 2024
Incremental and Data-Efficient Concept Formation to Support Masked Word Prediction

Xin Lian, Nishant Baglodi, Christopher J. MacLellan

This paper introduces Cobweb4L, a novel approach for efficient language model learning that supports masked word prediction. The approach builds on Cobweb, an incremental system that learns a hierarchy of probabilistic concepts. Each concept stores the frequencies of words that appear in instances tagged with that concept label. The system utilizes an attribute value representation to encode words and their surrounding context into instances. Cobweb4L uses the information theoretic variant of category utility and a new performance mechanism that leverages multiple concepts to generate predictions. We demonstrate that with these extensions it significantly outperforms prior Cobweb performance mechanisms that use only a single node to generate predictions. Further, we demonstrate that Cobweb4L learns rapidly and achieves performance comparable to and even superior to Word2Vec. Next, we show that Cobweb4L and Word2Vec outperform BERT in the same task with less training data. Finally, we discuss future work to make our conclusions more robust and inclusive.

7.3AIMay 23, 2024
HTN-Based Tutors: A New Intelligent Tutoring Framework Based on Hierarchical Task Networks

Momin N. Siddiqui, Adit Gupta, Jennifer M. Reddig et al.

Intelligent tutors have shown success in delivering a personalized and adaptive learning experience. However, there exist challenges regarding the granularity of knowledge in existing frameworks and the resulting instructions they can provide. To address these issues, we propose HTN-based tutors, a new intelligent tutoring framework that represents expert models using Hierarchical Task Networks (HTNs). Like other tutoring frameworks, it allows flexible encoding of different problem-solving strategies while providing the additional benefit of a hierarchical knowledge organization. We leverage the latter to create tutors that can adapt the granularity of their scaffolding. This organization also aligns well with the compositional nature of skills.

11.1AIMay 30, 2025
Taxonomic Networks: A Representation for Neuro-Symbolic Pairing

Zekun Wang, Ethan L. Haarer, Nicki Barari et al.

We introduce the concept of a \textbf{neuro-symbolic pair} -- neural and symbolic approaches that are linked through a common knowledge representation. Next, we present \textbf{taxonomic networks}, a type of discrimination network in which nodes represent hierarchically organized taxonomic concepts. Using this representation, we construct a novel neuro-symbolic pair and evaluate its performance. We show that our symbolic method learns taxonomic nets more efficiently with less data and compute, while the neural method finds higher-accuracy taxonomic nets when provided with greater resources. As a neuro-symbolic pair, these approaches can be used interchangeably based on situational needs, with seamless translation between them when necessary. This work lays the foundation for future systems that more fundamentally integrate neural and symbolic computation.

1.2NAOct 29, 2015
Accelerated Magnetic Resonance Thermometry in Presence of Uncertainties

Reza Madankan, Wolfgang Stefan, Samuel Fahrenholtz et al.

An accelerated model-based information theoretic approach is presented to perform the task of Magnetic Resonance (MR) thermal image reconstruction from a limited number of observed samples on k-space. The key idea of the proposed approach is to utilize information theoretic techniques to optimally detect samples of k-space that are information rich with respect to a model of the thermal data acquisition. These highly informative k-space samples are then used to refine the mathematical model and reconstruct the image. The information theoretic reconstruction is demonstrated retrospectively in data acquired during MR guided Laser Induced Thermal Therapy (MRgLITT) procedures. The approach demonstrates that locations of high-information content with respect to a model based reconstruction of MR thermometry may be quantitatively identified. The predicted locations of high-information content are sorted and retrospectively extracted from the fully sampled k-space measurements data set. The effect of interactively increasing the predicted number of data points used in the subsampled reconstruction is quantified using the L2-norm of the distance between the subsampled and fully sampled reconstruction. Performance of the proposed approach is also compared with clinically available subsampling techniques (rectilinear subsampling and variable-density Poisson disk undersampling). It is shown that the proposed subsampling scheme results in accurate reconstructions using small fraction of k-space points and suggest that the reconstruction technique may be useful in improving the efficiency of the thermometry data temporal resolution.