SEJul 28

CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agents

arXiv:2607.2543124.3h-index: 4
Predicted impact top 2% in SE · last 90 daysOriginality Incremental advance
AI Analysis

This work addresses the problem of inefficient context serving for coding agents by providing a unified runtime that maintains multiple views across repository edits, reducing redundant computation and token usage.

CodeNib introduces a multi-view data system that serves reusable lexical, dense, and structural views per repository commit to coding agents, achieving 8.7× and 25.4× faster graph and vector updates at the median compared to independent rebuilds, and reducing trajectory tokens by 50–87% while preserving localization across five models.

Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, language servers, and task-local histories force repeated discovery and obscure lifecycle costs. CodeNib builds reusable lexical, dense, and structural views per repository commit, maps outputs to repository-relative source ranges, maintains selected views across edits, and serves ranked search, symbol navigation, and bounded context through one runtime. Across 100 snapshots, we map quality-cost frontiers across the repository-context lifecycle. When outputs match an independent rebuild, graph and vector updates are $8.7\times$ and $25.4\times$ faster at the median. On the static-navigation subset matching normalized live-server locations (63% of 1,000 requests), the median per-request live/static latency ratio is $4.7\times$. Across five models, selected context policies preserve localization with 50--87% fewer trajectory tokens than paired grep/read. Together, these results support multi-view repository-context serving with explicit, operation-specific validity boundaries.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes