Towards Anticipatory Databases Through Shared Data and Workload Semantics
cs.DB, cs.LG
Submitted: 2026-09-13
Updated: 2026-09-13
Comments: 8 pages, 6 Figures
Code: https://github.com/fzirak/semantic-dbms
License: http://creativecommons.org/licenses/by/4.0/
The gist: Database management systems increasingly serve dynamic and exploratory workloads, yet many of their decisions still rely on low-level signals such as recency, frequency, and address locality.
Terminology
Abstract
Database management systems increasingly serve dynamic and exploratory workloads, yet many of their decisions still rely on low-level signals such as recency, frequency, and address locality. These signals capture how data was accessed, but not what is being examined or how an analytical focus evolves. We argue for treating workload semantics as a first-class control signal for anticipatory decision making. Central to this view, we introduce semantic locality and semantic trajectories, which capture relationships among nearby queries and how those relationships evolve across a session. We propose a framework that represents semantic context at the data, query, and session levels, models its evolution over time, and translates it into task-specific utility estimates. We instantiate this framework in semantic prefetching and semantic cache eviction, which share a semantic layer to make two separate decisions. Prefetching uses semantic trajectories to anticipate future accesses beyond what address-based locality can capture, while eviction uses semantic relevance to inform block replacement. These systems provide initial evidence that shared semantic context can support multiple DBMS components. We further outline how this principle can extend to other decisions and data systems, and discuss key challenges in representation, cost, adaptation, and evaluation.
Sources
- QVCache: A Query-Aware Vector Cache
- Learned Cardinalities: Estimating Correlated Joins with Deep Learning
- Fast EXP3 Algorithms
Related papers
- Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries
- Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering
- Bridging Business Intent and Data: A Benchmark for Automatic Relational Data Product Generation
- DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation
- MaDI-Bench: An End-to-End Data Integration Benchmark
- Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning