CL AIMar 13, 2024

Gemma: Open Models Based on Gemini Research and Technology

Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot

DeepMind

arXiv:2403.08295v450.11150 citationsh-index: 102

Originality Incremental advance

AI Analysis

This work provides open, state-of-the-art models for language tasks, enabling broader access and innovation in LLM development.

The authors introduced Gemma, a family of lightweight open models derived from Gemini research, which outperforms similarly sized open models on 11 out of 18 text-based tasks and includes evaluations for safety and responsibility.

This work introduces Gemma, a family of lightweight, state-of-the art open models built from the research and technology used to create Gemini models. Gemma models demonstrate strong performance across academic benchmarks for language understanding, reasoning, and safety. We release two sizes of models (2 billion and 7 billion parameters), and provide both pretrained and fine-tuned checkpoints. Gemma outperforms similarly sized open models on 11 out of 18 text-based tasks, and we present comprehensive evaluations of safety and responsibility aspects of the models, alongside a detailed description of model development. We believe the responsible release of LLMs is critical for improving the safety of frontier models, and for enabling the next wave of LLM innovations.

View on arXiv PDF

Similar