CVJul 2

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing

arXiv:2607.0229023.3Has Code
Predicted impact top 5% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers in image generation and editing, this work addresses the problem of generating knowledge-intensive diagrams that require disciplinary correctness, moving beyond aesthetic plausibility.

The paper introduces DisciplineGen-1M, a million-scale multidisciplinary dataset for text-to-image generation and editing, and a discipline-informed model that achieves substantial improvements over open-source baselines on discipline-related benchmarks (GenExam, GRADE) and general reasoning-informed benchmarks (WISE, RISE).

Recent image generation and editing models can produce visually appealing natural images, yet they remain unreliable when the target image is a knowledge-intensive diagram whose correctness depends on disciplinary concepts, symbolic structure, and precise spatial relations. We introduce DisciplineGen-1M, a million-scale multidisciplinary dataset that supports text-to-image generation and image editing. It contains 1.2M samples spanning mathematics, physics, chemistry, biology, geography, computer science, economics, history, music, and sports. To construct the dataset, we design a scalable framework that combines vector-graphics rendering, OCR-based editing, curated programmatic synthesis, and large-scale text-to-image filtering. These pipelines produce captions, editing instructions, structured annotations, and paired images with controllable semantic differences. Building on DisciplineGen-1M, we further introduce a discipline-informed reasoning-generation model for both text-to-image generation and image editing. Experiments on discipline-related benchmarks, GenExam and GRADE, show substantial improvements over open-source baselines, while evaluations on general reasoning-informed benchmarks, WISE and RISE, further indicate broader transfer. The results suggest that large-scale structured academic visual data is a key ingredient for moving image generation from aesthetic plausibility toward verifiable knowledge-grounded visual creation. We will publicly release our dataset, model, and source code of the data curation pipeline to ensure reproducibility and benefit future research.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes