Catalogian turns data into a record your LLM can question, then watches that record for changes. This page defines the concepts the whole product is built on: the record, the session, the key field, snapshots, delta events, editions, checks, and the sample.
A record is one versioned dataset. It is born the moment you paste a link or drop a file: rows are counted, columns are named, and the card flips to show it. A record can begin as an anonymous session on the homepage and later be claimed (kept) into an account, or be created directly in the dashboard wizard. Either way it is the same object: a name, a format (CSV, TSV, or JSONL), a key field, a status, and a tier.
Everything else scopes to the record: snapshots, delta events, editions, webhooks, and per-record limits. Row caps, check cadence, and history windows are properties of the record's tier, not of your account.
{
"id": "src_01j...",
"name": "West Coast Earthquakes",
"type": "http",
"format": "csv",
"keyField": "id",
"status": "active",
"tier": "free"
}An unkept record lives in an ephemeral session. The session has a hard 24-hour clock. The session token (cat_eph_...) is the only credential: whoever holds it can read the record and connect an LLM to it. When the clock runs out, an unkept record is gone; when you keep it, the record moves into your account intact.
The record also gets its own shareable address (/s/<token>) while the session runs.
The key field is the column that identifies each row. Catalogian diffs editions through it: a key present in the new data but not the old is a new row, a key present in both with different values is a changed row, a key that disappeared is a deleted row. On the homepage it is chosen for you (when one column stands out, it becomes the key); in the dashboard wizard you can pick it, with common key-name candidates like id, item_id, item_number, and upc suggested from your headers.
The key must be unique. Key values must be unique per row and stable across versions. An ingest whose key column has duplicate values is rejected with a structured 422 (DUPLICATE_KEY), and data where no column can serve as the key answers NO_KEY_FIELD. A row whose key changes between editions reads as a deletion plus an addition.
A snapshot is a point-in-time indexed copy of the record: each key maps to a row hash and the full row data. Every ingest produces a snapshot, which is what makes comparisons possible. The latest snapshot is always browsable: search it, filter it, profile it, query it, or export it, over MCP or the REST API.
A delta event is the diff between two consecutive snapshots. Every ingest that finds differences records one, with exact counts and the affected keys:
A check that finds nothing records a no-change event instead. Delta events carry summary counts plus up to 100 affected keys per category by default (the REST keysLimit parameter can raise that to 1,000). For full before/after row data, use the delta rows endpoint.
{
"id": "evt_01j...",
"detectedAt": "2026-09-28T20:00:00.000Z",
"newCount": 5,
"changedCount": 3,
"deletedCount": 1,
"unchangedCount": 193,
"totalCount": 202,
"isNoChange": false,
"newKeys": ["nc75443740", "nc75443741"],
"changedKeys": ["nc75443737", "nc75443736"],
"deletedKeys": ["nc75443001"]
}Every ingest of a record is an edition: a version of the record's data. The first ingest is edition one; each revision (a newer file or URL) adds another and diffs against the record, which is why the card can say things like "Revision 2: 40 new, 3 changed, 5 gone since the last one." Editions are listable and diffable over MCP (list_editions, get_edition_diff), and on a kept record the scheduled checks produce them automatically.
A kept record checks its source on a schedule. The floor depends on the record's tier, and the hourly cadence is a capability, not a price feature: when the origin serves ETags (Catalogian observes this on a fetch), the record earns faster conditional checks. Free records check daily; paid records check 4 times a day, hourly once the source shows ETag support. Every tier also gets a manual Check Now allowance per record: 3 a day on free, 10 a day on watch, unlimited on max.
An ingest larger than the row cap is never rejected. Catalogian indexes the first 10,000 rows and says so: the record, the API, and MCP all carry the truncation state, worded as an honest sample ("Sampled 10,000 of ~48,000 rows"). The sample is pinned to the first key set it saw, so later ingests diff within the same frame: rows that depart a fully scanned sample are real deletions ("left your catalog"), while a windowed scan only ever says rows "fell out of your sample". Watching the record unlocks the rest of the file on the next ingest, and a kept record can be re-pinned on demand with a resample: POST /v1/sources/:id/resample clears the pin and queues a check; there is no dedicated record-page UI for it.
Pricing is per record. Every account starts with 2 records that are always free (10,000 rows, sampled beyond); watch and max are paid rungs that raise the row cap, the cadence, and the history window. One paid record unlocks webhooks and API keys for the whole account. MCP is on every rung: anonymous sessions get 18 read-only tools with 50 calls per session, and account records keep the same tools. The exact ladder lives on Plans & Limits.
The same record speaks four surfaces: the dashboard (for people), MCP (for your LLM), the REST API (for your code), and webhooks (for push). How each one authenticates is covered in Authentication & Keys.
Ready to try it? Quick Start guide →