Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New LanguagesShih-Cheng Huang, Pin-Zu Li, Yu-Chi Hsu et al.
Recently, the development of open-source large language models (LLMs) has advanced rapidly. Nevertheless, due to data constraints, the capabilities of most open-source LLMs are primarily focused on English. To address this issue, we introduce the concept of $\textit{chat vector}$ to equip pre-trained language models with instruction following and human value alignment via simple model arithmetic. The chat vector is derived by subtracting the weights of a pre-trained base model (e.g. LLaMA2) from those of its corresponding chat model (e.g. LLaMA2-chat). By simply adding the chat vector to a continual pre-trained model's weights, we can endow the model with chat capabilities in new languages without the need for further training. Our empirical studies demonstrate the superior efficacy of the chat vector from three different aspects: instruction following, toxicity mitigation, and multi-turn dialogue. Moreover, to showcase the adaptability of our approach, we extend our experiments to encompass various languages, base models, and chat vectors. The results underscore the chat vector's simplicity, effectiveness, and wide applicability, making it a compelling solution for efficiently enabling conversational capabilities in pre-trained language models. Our code is available at https://github.com/aqweteddy/ChatVector.
5.7SDJun 16, 2022
EPG2S: Speech Generation and Speech Enhancement based on Electropalatography and Audio Signals using Multimodal LearningLi-Chin Chen, Po-Hsun Chen, Richard Tzong-Han Tsai et al.
Speech generation and enhancement based on articulatory movements facilitate communication when the scope of verbal communication is absent, e.g., in patients who have lost the ability to speak. Although various techniques have been proposed to this end, electropalatography (EPG), which is a monitoring technique that records contact between the tongue and hard palate during speech, has not been adequately explored. Herein, we propose a novel multimodal EPG-to-speech (EPG2S) system that utilizes EPG and speech signals for speech generation and enhancement. Different fusion strategies based on multiple combinations of EPG and noisy speech signals are examined, and the viability of the proposed method is investigated. Experimental results indicate that EPG2S achieves desirable speech generation outcomes based solely on EPG signals. Further, the addition of noisy speech signals is observed to improve quality and intelligibility. Additionally, EPG2S is observed to achieve high-quality speech enhancement based solely on audio signals, with the addition of EPG signals further improving the performance. The late fusion strategy is deemed to be the most effective approach for simultaneous speech generation and enhancement.
8.7LGMar 1, 2022
Side Effects of Learning from Low-dimensional Data Embedded in a Euclidean SpaceJuncai He, Richard Tsai, Rachel Ward
The low-dimensional manifold hypothesis posits that the data found in many applications, such as those involving natural images, lie (approximately) on low-dimensional manifolds embedded in a high-dimensional Euclidean space. In this setting, a typical neural network defines a function that takes a finite number of vectors in the embedding space as input. However, one often needs to consider evaluating the optimized network at points outside the training distribution. This paper considers the case in which the training data is distributed in a linear subspace of $\mathbb R^d$. We derive estimates on the variation of the learning function, defined by a neural network, in the direction transversal to the subspace. We study the potential regularization effects associated with the network's depth and noise in the codimension of the data manifold. We also present additional side effects in training due to the presence of noise.
4.3CLAug 29, 2023
Large Language Models on the Chessboard: A Study on ChatGPT's Formal Language Comprehension and Complex Reasoning SkillsMu-Tien Kuo, Chih-Chung Hsueh, Richard Tzong-Han Tsai
While large language models have made strides in natural language processing, their proficiency in complex reasoning tasks requiring formal language comprehension, such as chess, remains less investigated. This paper probes the performance of ChatGPT, a sophisticated language model by OpenAI in tackling such complex reasoning tasks, using chess as a case study. Through robust metrics examining both the legality and quality of moves, we assess ChatGPT's understanding of the chessboard, adherence to chess rules, and strategic decision-making abilities. Our evaluation identifies limitations within ChatGPT's attention mechanism that affect its formal language comprehension and uncovers the model's underdeveloped self-regulation abilities. Our study also reveals ChatGPT's propensity for a coherent strategy in its gameplay and a noticeable uptick in decision-making assertiveness when the model is presented with a greater volume of natural language or possesses a more lucid understanding of the state of the chessboard. These findings contribute to the growing exploration of language models' abilities beyond natural language processing, providing valuable information for future research towards models demonstrating human-like cognitive abilities.
1.2NADec 25, 2017
Volumetric variational principles for a class of partial differential equations defined on surfaces and curvesJay Chu, Richard Tsai
In this paper, we propose simple numerical algorithms for partial differential equations (PDEs) defined on closed, smooth surfaces (or curves). In particular, we consider PDEs that originate from variational principles defined on the surfaces; these include Laplace-Beltrami equations and surface wave equations. The approach is to systematically formulate extensions of the variational integrals, derive the Euler-Lagrange equations of the extended problem, including the boundary conditions} that can be easily discretized on uniform Cartesian grids or adaptive meshes. In our approach, the surfaces are defined implicitly by the distance functions or by the closest point mapping. As such extensions are not unique, we investigate how a class of simple extensions can influence the resulting PDEs. In particular, we reduce the surface PDEs to model problems defined on a periodic strip and the corresponding boundary conditions, and use classical Fourier and Laplace transform methods to study the well-posedness of the resulting problems. For elliptic and parabolic problems, our boundary closure mostly yields stable algorithms to solve nonlinear surface PDEs. For hyperbolic problems, the proposed boundary closure is unstable in general, but the instability can be easily controlled by either adding a higher order regularization term or by periodically but infrequently "reinitializing" the computed solutions. Some numerical examples for each representative surface PDEs are presented.
Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese HokkienSin-En Lu, Bo-Han Lu, Chao-Yi Lu et al.
In natural language processing (NLP), code-mixing (CM) is a challenging task, especially when the mixed languages include dialects. In Southeast Asian countries such as Singapore, Indonesia, and Malaysia, Hokkien-Mandarin is the most widespread code-mixed language pair among Chinese immigrants, and it is also common in Taiwan. However, dialects such as Hokkien often have a scarcity of resources and the lack of an official writing system, limiting the development of dialect CM research. In this paper, we propose a method to construct a Hokkien-Mandarin CM dataset to mitigate the limitation, overcome the morphological issue under the Sino-Tibetan language family, and offer an efficient Hokkien word segmentation method through a linguistics-based toolkit. Furthermore, we use our proposed dataset and employ transfer learning to train the XLM (cross-lingual language model) for translation tasks. To fit the code-mixing scenario, we adapt XLM slightly. We found that by using linguistic knowledge, rules, and language tags, the model produces good results on CM data translation while maintaining monolingual translation quality.
1.2NAMar 6, 2017
An extrapolative approach to integration over hypersurfaces in the level set frameworkCatherine Kublik, Richard Tsai
We provide a new approach for computing integrals over hypersurfaces in the level set framework. The method is based on the discretization (via simple Riemann sums) of the classical formulation used in the level set framework, with the choice of specific kernels supported on a tubular neighborhood around the interface to approximate the Dirac delta function. The novelty lies in the choice of kernels, specifically its number of vanishing moments, which enables accurate computations of integrals over a class of closed, continuous, piecewise smooth, curves or surfaces; e.g. curves in two dimensions that contain finite number of corners. We prove that for smooth interfaces, if the kernel has enough vanishing moments (related to the dimension of the embedding space), the analytical integral formulation coincides exactly with the integral one wishes to calculate. For curves with corners and cusps, the formulation is not exact but we provide an analytical result relating the severity of the corner or cusp with the width of the tubular neighborhood. We show numerical examples demonstrating the capability of the approach, especially for integrating over piecewise smooth interfaces and for computing integrals where the integrand is only Lipschitz continuous or has an integrable singularity.
1.2NANov 18, 2015
Parareal methods for highly oscillatory dynamical systemsGil Ariel, Seong Jun Kim, Richard Tsai
We introduce a new strategy for coupling the parallel in time (parareal) iterative methodology with multiscale integrators. Following the parareal framework, the algorithm computes a low-cost approximation of all slow variables in the system using an appropriate multiscale integrator, which is refined using parallel fine scale integrations. Convergence is obtained using an alignment algorithm for fast phase-like variables. The method may be used either to enhance the accuracy and range of applicability of the multiscale method in approximating only the slow variables, or to resolve all the state variables. The numerical scheme does not require that the system is split into slow and fast coordinates. Moreover, the dynamics may involve hidden slow variables, for example, due to resonances. We propose an alignment algorithm for almost-periodic solution, in which case convergence of the parareal iterations is proved. The applicability of the method is demonstrated in numerical examples.
1.2NAApr 1, 2016
An implicit boundary integral method for interfaces evolving by Mullins-Sekerka dynamicsChieh Chen, Catherine Kublik, Richard Tsai
We present an algorithm for computing the nonlinear interface dynamics of the Mullins-Sekerka model for interfaces that are defined implicitly (e.g. by a level set function) using integral equations . The computation of the dynamics involves solving Laplace's equation with Dirichlet boundary conditions on multiply connected and unbounded domains and propagating the interface using a normal velocity obtained from the solution of the PDE at each time step. Our method is based on a simple formulation for implicit interfaces, which rewrites boundary integrals as volume integrals over the entire space. The resulting algorithm thus inherits the benefits of both level set methods and boundary integral methods to simulate the nonlocal front propagation problem with possible topological changes. We present numerical results in both two and three dimensions to demonstrate the effectiveness of the algorithm.
1.2NAFeb 7, 2018
θ-parareal schemesGil Ariel, Hieu Nguyen, Richard Tsai
A weighted version of the parareal method for parallel-in-time computation of time dependent problems is presented. Linear stability analysis for a scalar weighing strategy shows that the new scheme may enjoy favorable stability properties with marginal reduction in accuracy at worse. More complicated matrix-valued weights are applied in numerical examples. The weights are optimized using information from past iterations, providing a systematic framework for using the parareal iterations as an approach to multiscale coupling. The advantage of the method is demonstrated using numerical examples, including some well-studied nonlinear Hamiltonian systems.
3.8LGJul 5, 2023
Linear Regression on Manifold Structured Data: the Impact of Extrinsic Geometry on SolutionsLiangchen Liu, Juncai He, Richard Tsai
In this paper, we study linear regression applied to data structured on a manifold. We assume that the data manifold is smooth and is embedded in a Euclidean space, and our objective is to reveal the impact of the data manifold's extrinsic geometry on the regression. Specifically, we analyze the impact of the manifold's curvatures (or higher order nonlinearity in the parameterization when the curvatures are locally zero) on the uniqueness of the regression solution. Our findings suggest that the corresponding linear regression does not have a unique solution when the embedded submanifold is flat in some dimensions. Otherwise, the manifold's curvature (or higher order nonlinearity in the embedding) may contribute significantly, particularly in the solution associated with the normal directions of the manifold. Our findings thus reveal the role of data manifold geometry in ensuring the stability of regression models for out-of-distribution inferences.
1.2NAJun 12, 2018
A Multiscale Domain Decomposition Algorithm For Boundary Value Problems For Eikonal EquationsLindsay Martin, Richard Tsai
In this paper, we present a new multiscale domain decomposition algorithm for computing solutions of static Eikonal equations. The new method is an iterative two-scale method that uses a parareal-like update scheme in combination with standard Eikonal solvers. The purpose of the two scales is to accelerate convergence and maintain accuracy. We adapt a weighted version of the parareal method for stability, and the optimal weights are studied via a model problem. Numerical examples are given to demonstrate the method.
SMUTF: Schema Matching Using Generative Tags and Hybrid FeaturesYu Zhang, Mei Di, Haozheng Luo et al.
We introduce SMUTF (Schema Matching Using Generative Tags and Hybrid Features), a unique approach for large-scale tabular data schema matching (SM), which assumes that supervised learning does not affect performance in open-domain tasks, thereby enabling effective cross-domain matching. This system uniquely combines rule-based feature engineering, pre-trained language models, and generative large language models. In an innovative adaptation inspired by the Humanitarian Exchange Language, we deploy "generative tags" for each data column, enhancing the effectiveness of SM. SMUTF exhibits extensive versatility, working seamlessly with any pre-existing pre-trained embeddings, classification methods, and generative models. Recognizing the lack of extensive, publicly available datasets for SM, we have created and open-sourced the HDXSM dataset from the public humanitarian data. We believe this to be the most exhaustive SM dataset currently available. In evaluations across various public datasets and the novel HDXSM dataset, SMUTF demonstrated exceptional performance, surpassing existing state-of-the-art models in terms of accuracy and efficiency, and improving the F1 score by 11.84% and the AUC of ROC by 5.08%. Code is available at https://github.com/fireindark707/Python-Schema-Matching.
5.5IRJan 29, 2019Code
Revised JNLPBA Corpus: A Revised Version of Biomedical NER Corpus for Relation Extraction TaskMing-Siang Huang, Po-Ting Lai, Richard Tzong-Han Tsai et al.
The advancement of biomedical named entity recognition (BNER) and biomedical relation extraction (BRE) researches promotes the development of text mining in biological domains. As a cornerstone of BRE, robust BNER system is required to identify the mentioned NEs in plain texts for further relation extraction stage. However, the current BNER corpora, which play important roles in these tasks, paid less attention to achieve the criteria for BRE task. In this study, we present Revised JNLPBA corpus, the revision of JNLPBA corpus, to broaden the applicability of a NER corpus from BNER to BRE task. We preserve the original entity types including protein, DNA, RNA, cell line and cell type while all the abstracts in JNLPBA corpus are manually curated by domain experts again basis on the new annotation guideline focusing on the specific NEs instead of general terms. Simultaneously, several imperfection issues in JNLPBA are pointed out and made up in the new corpus. To compare the adaptability of different NER systems in Revised JNLPBA and JNLPBA corpora, the F1-measure was measured in three open sources NER systems including BANNER, Gimli and NERSuite. In the same circumstance, all the systems perform average 10% better in Revised JNLPBA than in JNLPBA. Moreover, the cross-validation test is carried out which we train the NER systems on JNLPBA/Revised JNLPBA corpora and access the performance in both protein-protein interaction extraction (PPIE) and biomedical event extraction (BEE) corpora to confirm that the newly refined Revised JNLPBA is a competent NER corpus in biomedical relation application. The revised JNLPBA corpus is freely available at iasl-btm.iis.sinica.edu.tw/BNER/Content/Revised_JNLPBA.zip.
Enhancing Taiwanese Hokkien Dual Translation by Exploring and Standardizing of Four Writing SystemsBo-Han Lu, Yi-Hsuan Lin, En-Shiun Annie Lee et al.
Machine translation focuses mainly on high-resource languages (HRLs), while low-resource languages (LRLs) like Taiwanese Hokkien are relatively under-explored. The study aims to address this gap by developing a dual translation model between Taiwanese Hokkien and both Traditional Mandarin Chinese and English. We employ a pre-trained LLaMA 2-7B model specialized in Traditional Mandarin Chinese to leverage the orthographic similarities between Taiwanese Hokkien Han and Traditional Mandarin Chinese. Our comprehensive experiments involve translation tasks across various writing systems of Taiwanese Hokkien as well as between Taiwanese Hokkien and other HRLs. We find that the use of a limited monolingual corpus still further improves the model's Taiwanese Hokkien capabilities. We then utilize our translation model to standardize all Taiwanese Hokkien writing systems into Hokkien Han, resulting in further performance improvements. Additionally, we introduce an evaluation method incorporating back-translation and GPT-4 to ensure reliable translation quality assessment even for LRLs. The study contributes to narrowing the resource gap for Taiwanese Hokkien and empirically investigates the advantages and limitations of pre-training and fine-tuning based on LLaMA 2.
4.9CLSep 24, 2025
SiniticMTError: A Machine Translation Dataset with Error Annotations for Sinitic LanguagesHannah Liu, Junghyun Min, Ethan Yue Heng Cheung et al.
Despite major advances in machine translation (MT) in recent years, progress remains limited for many low-resource languages that lack large-scale training data and linguistic resources. Cantonese and Wu Chinese are two Sinitic examples, although each enjoys more than 80 million speakers around the world. In this paper, we introduce SiniticMTError, a novel dataset that builds on existing parallel corpora to provide error span, error type, and error severity annotations in machine-translated examples from English to Mandarin, Cantonese, and Wu Chinese. Our dataset serves as a resource for the MT community to utilize in fine-tuning models with error detection capabilities, supporting research on translation quality estimation, error-aware generation, and low-resource language evaluation. We report our rigorous annotation process by native speakers, with analyses on inter-annotator agreement, iterative feedback, and patterns in error type and severity.
1.2NAJul 19, 2025
Numerical Artifacts in Learning Dynamical SystemsBing-Ze Lu, Richard Tsai
In many applications, one needs to learn a dynamical system from its solutions sampled at a finite number of time points. The learning problem is often formulated as an optimization problem over a chosen function class. However, in the optimization procedure, it is necessary to employ a numerical scheme to integrate candidate dynamical systems and assess how their solutions fit the data. This paper reveals potentially serious effects of a chosen numerical scheme on the learning outcome. In particular, our analysis demonstrates that a damped oscillatory system may be incorrectly identified as having "anti-damping" and exhibiting a reversed oscillation direction, despite adequately fitting the given data points.
1.9RONov 18, 2019
Strategy Synthesis for Surveillance-Evasion Games with Learning-Enabled Visibility OptimizationSuda Bharadwaj, Louis Ly, Bo Wu et al.
This paper studies a two-player game with a quantitative surveillance requirement on an adversarial target moving in a discrete state space and a secondary objective to maximize short-term visibility of the environment. We impose the surveillance requirement as a temporal logic constraint.We then use a greedy approach to determine vantage points that optimize a notion of information gain, namely, the number of newly-seen states. By using a convolutional neural network trained on a class of environments, we can efficiently approximate the information gain at each potential vantage point.Subsequent vantage points are chosen such that moving to that location will not jeopardize the surveillance requirement, regardless of any future action chosen by the target. Our method combines guarantees of correctness from formal methods with the scalability of machine learning to provide an efficient approach for surveillance-constrained visibility optimization.
1.1CLOct 11, 2015
Textual Analysis for Studying Chinese Historical Documents and Literary NovelsChao-Lin Liu, Guan-Tao Jin, Hongsu Wang et al.
We analyzed historical and literary documents in Chinese to gain insights into research issues, and overview our studies which utilized four different sources of text materials in this paper. We investigated the history of concepts and transliterated words in China with the Database for the Study of Modern China Thought and Literature, which contains historical documents about China between 1830 and 1930. We also attempted to disambiguate names that were shared by multiple government officers who served between 618 and 1912 and were recorded in Chinese local gazetteers. To showcase the potentials and challenges of computer-assisted analysis of Chinese literatures, we explored some interesting yet non-trivial questions about two of the Four Great Classical Novels of China: (1) Which monsters attempted to consume the Buddhist monk Xuanzang in the Journey to the West (JTTW), which was published in the 16th century, (2) Which was the most powerful monster in JTTW, and (3) Which major role smiled the most in the Dream of the Red Chamber, which was published in the 18th century. Similar approaches can be applied to the analysis and study of modern documents, such as the newspaper articles published about the 228 incident that occurred in 1947 in Taiwan.
1.2NAOct 15, 2015
Integration over curves and surfaces defined by the closest point mappingCatherine Kublik, Richard Tsai
We propose a new formulation for integrating over smooth curves and surfaces that are described by their closest point mappings. Our method is designed for curves and surfaces that are not defined by any explicit parameterization and is intended to be used in combination with level set techniques. However, contrary to the common practice with level set methods, the volume integrals derived from our formulation coincide exactly with the surface or line integrals that one wish to compute. We study various aspects of this formulation and provide a geometric interpretation of this formulation in terms of the singular values of the Jacobian matrix of the closest point mapping. Additionally, we extend the formulation - initially derived to integrate over manifolds of codimension one - to include integration along curves in three dimensions. Some numerical examples using very simple discretizations are presented to demonstrate the efficacy of the formulation.