What is relativedb?
relativedb answers questions about the future of your relational data. You declare the shape of your tables and links, wire small retriever callbacks over your existing storage, and write a predictive query:
PREDICT NOT EXISTS(orders.*) OVER (90 DAYS FOLLOWING) FROM customers
How it fits together
- RelQL. A SQL-flavored query language for predictions. Parsed and validated against your declared schema. See the RelQL docs.
- Retrievers. The engine never touches your database. All data access goes through callbacks you implement, GraphQL-style. See Retrievers.
- Temporal context assembly. The engine hops your relational graph to build a per-entity context, and guarantees nothing newer than the anchor time enters it. See Temporal correctness.
Installation
Python
Requires Python 3.10+.
pip install relativedb
Quickstart
Predict 90-day churn for every customer, as of July 1.
- Python
Declare the schema and wire callbacks over your own storage. If your data is in pandas, pandas remains an application dependency:
import pandas as pd
from relativedb import Engine, ExecutionInput, RetrieverWiring, Row, RtNativeBackend
customer_rows = [Row("customers", r.customer_id, {"age": float(r.age)})
for r in customers.itertuples()]
by_id = {row.id: row for row in customer_rows}
def fetch_customers(table, ids, bound):
return [by_id[row_id] for row_id in ids if row_id in by_id]
wiring = (RetrieverWiring.new_wiring()
.entities("customers", fetch_customers)
.entities("orders", fetch_orders)
.default_links(fetch_order_children)
.scanner("customers", scan_customers)
.build())
# Scoring requires a model backend. RtNativeBackend runs the RT-J relational
# model; it needs a cached RelativeDB/rt-j-fp16 checkpoint.
engine = Engine(schema, wiring, model_backend=RtNativeBackend(schema=schema))
result = engine.execute(ExecutionInput(
query="PREDICT NOT EXISTS(orders.*) OVER (90 DAYS FOLLOWING) FROM customers",
anchor_time=pd.Timestamp("2026-07-01").to_pydatetime()))
Next steps
- Learn the query language: RelQL tutorial
- Wire your own storage: Retrievers
- Fit a ranking adapter on your history: Fit a multiclass/ranking adapter
Retrievers
The engine never touches a database. It asks your code for rows through four small interfaces, each a plain Python callable:
| Interface | Signature (conceptual) | Role |
|---|---|---|
EntityRetriever | (table, ids, bound) → rows | Batched point lookup: seed rows, parents |
LinkRetriever | (link, parent_id, bound, limit) → rows | Children along one FK link, newest-first |
CohortRetriever (optional) | (table, anchor, bound, limit) → ids | Similar entities for in-context examples |
TableScanner (optional) | (table, bound) → row stream | Bulk streaming. Enables whole-table FROM and CSC mode |
(Normalization statistics are not a retriever: pass a ColumnStats to the
model backend.)
Rows
A Row carries typed cells, an optional timestamp, and parent edges
({fk_column: parent_id}). FK values are not cells. They surface as edges.
The primary key surfaces as identity, and additionally as a cell when the
schema declares it as a column.
A row whose table declares no feature columns emits no tokens. A token-less row
that others link through is a dead end: nothing below it reaches the prediction,
and every entity scores alike. The engine emits a ContextConnectivityWarning
when it detects this. Give the table a feature column, or declare its primary
key as one.
Keys as features
A primary key is identity by default. When the key carries meaning (a stock
code, an ISBN, an airport code), declare it as a column too, the same way you
declare a time_column. The engine then emits it as a feature cell:
TableDef.new_table("users").primary_key("user_id") # identity only
(TableDef.new_table("products") # ...and a feature
.column("stock_code", ValueType.TEXT).primary_key("stock_code"))
Leave synthetic keys out. Autoincrement IDs track insertion order, so the model reads one as a proxy for tenure, a signal that breaks on a new ID range.
Wiring
A RetrieverWiring binds retrievers to tables and links, with a
default_links catch-all. It is validated against the schema when the engine
is built, so a missing retriever fails at startup, before any query runs.
Temporal correctness
Temporal leakage (a "future" fact sneaking into the features) is the classic way predictive systems lie in backtests. relativedb makes leakage prevention an engine guarantee rather than something you have to remember.
The anchor time
Every execution has an anchor time t₀. The prediction target reads the window after t₀. The assembled context may only contain data at or before t₀.
Defense in depth
- Every retriever call carries a
TemporalBound: "return nothing newer than this". Rows without timestamps (static dimension tables) are always admitted. - The engine re-checks every returned row against the bound and drops violations. A buggy or malicious retriever cannot leak the future into context. Dedicated tests in all three libraries feed a deliberately broken retriever and assert the future row never appears.
Window direction is validated
Target windows must face the future (non-negative offsets). WHERE filter
windows face the past (negative or -INF starts). The validator rejects
queries that mix these up.
Backtesting for free
Because "as of" is an explicit input, evaluating yesterday's model is just running the same query with yesterday's anchor. No snapshot tables, no point-in-time joins.
Sampler modes
Context assembly walks the graph: seed entity → parents (always followed) → children (fanout-capped, newest-first) → optional cohort, until the hop limit or cell budget. Two interchangeable samplers drive this walk, and both produce identical contexts (asserted by tests).
RETRIEVER (default)
Pull-per-hop: the hop loop calls your retrievers for each expansion.
Use when data is remote, huge, or access-controlled. Nothing is copied, and your retrievers see every access.
CSC
The engine drains each TableScanner once into in-memory
compressed-sparse-column adjacency arrays (time-sorted neighbor lists), then
samples entirely in-process, so "latest w children ≤ anchor" is one binary
search plus a tail slice.
Use for latency-sensitive, repeated scoring over data that fits in memory. The index is a snapshot, built once when the engine is constructed and immutable thereafter. To pick up changed data, construct a new engine.
Context budgets
ContextPolicy supports two geometries:
- per-hop fanouts, e.g.
fanouts=(64, 64) - a uniform
bfs_width(default 32) under a globalmax_context_cellsbudget (default 2048)
All of it is configurable from the Python API. The model backend has a
matching token-sequence cap, max_seq_len (default 2048; the reference
evaluation runs at 8192):
from relativedb import ContextPolicy, Engine, RtNativeBackend
engine = Engine(schema, wiring,
context_policy=ContextPolicy(max_context_cells=8192, bfs_width=64),
model_backend=RtNativeBackend(schema=schema, max_seq_len=8192))
Larger budgets admit more history per entity and cost latency; keep
max_seq_len at or above max_context_cells so assembled context is not
truncated at the model boundary.
See Choose a sampler mode for a decision guide.
Supported output types
The checkpoint executes binary classification, regression,
multiclass classification, and ranking. RETURN CLASS, RETURN DISTRIBUTION, RETURN PROBABILITY, and RETURN EXPECTED VALUE work.
Multiclass reuses the checkpoint's text head. It decodes the masked target
cell to a 384-dim embedding, then matches that against the class labels' MiniLM
embeddings by cosine similarity. You get a predicted class and approximate class
probabilities: the argmax is reference-exact, while the probabilities are only a
softmax over cosine scores. Ranking scores each candidate parent ID with the existence
head, sigmoids it, and returns the top k. RETURN QUANTILES and
RETURN INTERVAL are not part of the language: the model exposes a single
point estimate, not a distribution, so they could never execute. A query using
them is rejected at parse time with a message naming them.
Choose a sampler mode
Both modes produce identical contexts. Choose by data locality.
| RETRIEVER (default) | CSC | |
|---|---|---|
| Data location | stays in your store | copied into an in-memory index |
| Freshness | live, per query | snapshot, fixed at construction |
| Requires | Entity + Link retrievers | TableScanner per table |
| Best for | remote, huge, or access-controlled data | repeated low-latency scoring |
Switching
from relativedb import Engine, ExecutionInput, SamplerMode
engine = Engine(schema, wiring, sampler_mode=SamplerMode.CSC)
result = engine.execute(ExecutionInput(query=query, anchor_time=t0))
What CSC buys you
On a synthetic churn workload (10,000 customers, 200,000 orders, history baseline, M-series laptop), scoring every customer:
| Approach | Time | Throughput |
|---|---|---|
| relativedb, CSC sampler | 0.66 s | ~15,000 entities/s |
| naive per-entity pandas loop | 57.4 s | ~174 entities/s |
The CSC index turns each "latest w children ≤ anchor" expansion into a binary search plus a tail slice, and its build cost is paid once per snapshot.
Rule of thumb
Start with RETRIEVER. Move to CSC when the same engine scores many queries or large populations and the data fits in memory.
Fit a multiclass/ranking adapter
The released checkpoint is zero-shot. Training moved out of relativedb: fit
adapters with the relational-transformers
package (fit_feature_head over frozen target features, RelationalTrainer
for full fine-tuning) and serve the resulting head here by passing it to the
backend:
from relational_transformers import fit_feature_head
head = fit_feature_head(features, labels, "ranking",
group_offsets=offsets, n_groups=n_groups)
head.save("head.safetensors")
tuned = Engine(schema, wiring, model_backend=RtNativeBackend(
schema=schema, wiring=wiring, head="head.safetensors"))
The head replaces the checkpoint's zero-shot head only for the task it was trained on, and inference on a trained head is plain CPU, so an adapter trained anywhere serves anywhere. Judge the result on held-out anchors, never on training loss.