John T. E. Richardson

CL
h-index69
3papers
4,339citations
Novelty30%
AI Score30

3 Papers

42.5CLAug 19, 2018Code
SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Taku Kudo, John Richardson

This paper describes SentencePiece, a language-independent subword tokenizer and detokenizer designed for Neural-based text processing, including Neural Machine Translation. It provides open-source C++ and Python implementations for subword units. While existing subword segmentation tools assume that the input is pre-tokenized into word sequences, SentencePiece can train subword models directly from raw sentences, which allows us to make a purely end-to-end and language independent system. We perform a validation experiment of NMT on English-Japanese machine translation, and find that it is possible to achieve comparable accuracy to direct subword training from raw sentences. We also compare the performance of subword training and segmentation with various configurations. SentencePiece is available under the Apache 2 license at https://github.com/google/sentencepiece.

29.7LGFeb 21, 2019Code
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

Jonathan Shen, Patrick Nguyen, Yonghui Wu et al.

Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models are composed of modular building blocks that are flexible and easily extensible, and experiment configurations are centralized and highly customizable. Distributed training and quantized inference are supported directly within the framework, and it contains existing implementations of a large number of utilities, helper functions, and the newest research ideas. Lingvo has been used in collaboration by dozens of researchers in more than 20 papers over the last two years. This document outlines the underlying design of Lingvo and serves as an introduction to the various pieces of the framework, while also offering examples of advanced features that showcase the capabilities of the framework.

3.0HCJul 12, 2018
Using the Value of Information (VoI) Metric to Improve Sensemaking

Mark Mittrick, John Richardson, Derrik E. Asher et al.

Sensemaking is the cognitive process of extracting information, creating schemata from knowledge, making decisions from those schemata, and inferring conclusions. Human analysts are essential to exploring and quantifying this process, but they are limited by their inability to process the volume, variety, velocity, and veracity of data. Visualization tools are essential for helping this human-computer interaction. For example, analytical tools that use graphical linknode visualization can help sift through vast amounts of information. However, assisting the analyst in making connections with visual tools can be challenging if the information is not presented in an intuitive manner. Experimentally, it has been shown that analysts increase the number of hypotheses formed if they use visual analytic capabilities. Exploring multiple perspectives could increase the diversity of those hypotheses, potentially minimizing cognitive biases. In this paper, we discuss preliminary research results that indicate an improvement in sensemaking over the traditional link-node visualization tools by incorporating an annotation enhancement that differentiates links connecting nodes. This enhancement assists by providing a visual cue, which represents the perceived value of reported information. We conclude that this improved sensemaking occurs because of the removal of the limitations of mentally consolidating, weighing, and highlighting data. This study aims to investigate whether line thickness can be used as a valid representation of VoI.