A record's source is where its data comes from: a URL Catalogian fetches, or a file you upload. Either way the record is born on ingest and every later ingest becomes a new edition, diffed row by row through the key field. See Core Concepts for the vocabulary.
On the homepage, the server fetches the file from the origin during the ingest and the card flips when the record is back. In the dashboard wizard, a URL record starts its schedule and the first check establishes the baseline. The URL can require authentication: basic auth, bearer tokens, and API keys in a header or query param are all supported. See Authentication & Keys.
Drag a file onto the homepage card or choose one, or upload through the dashboard. CSV, TSV, JSONL, and gzip-compressed variants of all three, up to 520 MB per ingest. Uploads show real progress (bytes per second, time remaining), then an indexing panel while the server parses. An oversize file answers a structured 413 that suggests gzip; nothing else is implied.
A file upload in the dashboard is staged, not instant: a 201 means the bytes are accepted and the first ingest is queued. The record then walks a visible lifecycle: staged (bytes in durable storage, a worker has not picked the job up yet), ingesting (a worker is processing), then live with a sample notice when the free cap applied, or error with the server's message. If one stage sits too long, the wizard says so and offers to keep waiting instead of spinning forever.
Records are versioned by re-ingesting. A kept record ingests on its check schedule, so editions accrue automatically and each one diffs against the record ("40 new, 3 changed, 5 gone since the last one"). Anonymous records are a read-and-connect preview: the record travels by its share link, and versioning begins when the record is kept.
If your pipeline owns the cadence instead of a scheduler, a kept record can run on demand: pause it, overwrite your file, trigger the check with POST /v1/sources/:id/check, and pull the delta. The full walkthrough is in Pipeline-Triggered Ingestion.
Every record has a key field: the column that uniquely identifies each row (id, item_id, and friends). On the homepage it is chosen for you from your headers; the dashboard wizard analyzes the first rows, scores each column for uniqueness, and suggests candidates. The key is what makes deltas possible, and duplicate key values reject an ingest. Details in Core Concepts.
CSV, TSV, JSONL, and gzip-compressed variants of all three. Max file size: 520 MB. See File Formats for delimiters, encoding, and edge cases.
A URL record's health panel shows what the last check saw, including whether the origin served an ETag. ETag support is what unlocks the hourly cadence on paid records: the origin answers conditional requests, so checks get cheap and frequent. A record whose free sample is pinned can also be resampled on demand by calling the resample endpoint (POST /v1/sources/:id/resample); there is no dedicated record-page UI for it.
Row caps, check cadence, and history windows are per record; they scale with the record's tier. View plans & limits →