CLJun 4, 2025

Around the World in 24 Hours: Probing LLM Knowledge of Time and Place

arXiv:2506.03984v16 citationsh-index: 28ACL
Originality Incremental advance
AI Analysis

This work addresses a gap in understanding language models' spatiotemporal reasoning for AI research, though it is incremental as it builds on prior isolated evaluations.

The paper tackled the problem of evaluating language models' joint reasoning over time and space, creating the GeoTemp dataset with 320k prompts across 289 cities, 217 countries, and 37 time zones, and found that models perform well on temporal tasks but struggle with connecting temporal and geographic information, with performance influenced by prompt formulation.

Reasoning over time and space is essential for understanding our world. However, the abilities of language models in this area are largely unexplored as previous work has tested their abilities for logical reasoning in terms of time and space in isolation or only in simple or artificial environments. In this paper, we present the first evaluation of the ability of language models to jointly reason over time and space. To enable our analysis, we create GeoTemp, a dataset of 320k prompts covering 289 cities in 217 countries and 37 time zones. Using GeoTemp, we evaluate eight open chat models of three different model families for different combinations of temporal and geographic knowledge. We find that most models perform well on reasoning tasks involving only temporal knowledge and that overall performance improves with scale. However, performance remains constrained in tasks that require connecting temporal and geographical information. We do not find clear correlations of performance with specific geographic regions. Instead, we find a significant performance increase for location names with low model perplexity, suggesting their repeated occurrence during model training. We further demonstrate that their performance is heavily influenced by prompt formulation - a direct injection of geographical knowledge leads to performance gains, whereas, surprisingly, techniques like chain-of-thought prompting decrease performance on simpler tasks.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes