MCP Integration

Catalogian supports the Model Context Protocol (MCP), letting AI agents query your data changes directly: no custom integration work required. Connect once and your agents can list records, query delta events, and retrieve full row-level field data. Catalogian exposes 18 tools, all read-only; the list_tools directory always reflects the live registry.

No account?

You do not need an account to use the MCP server. Drop a link to your data and get an MCP server in seconds:

curl -X POST https://api.catalogian.com/v1/ephemeral/ingest \
  -H "Content-Type: application/json" \
  -d '{"url": "https://your-bucket.com/data.csv"}'

The response includes a session token (cat_eph_...). Connect your LLM to https://api.catalogian.com/v1/mcp with Authorization: Bearer cat_eph_YOUR_SESSION_TOKEN and all 18 tools work, scoped to the session's record. Free: 10,000 rows indexed per record, 3 ingests per session, 24-hour session life, 50 MCP tool calls per session. Create an account to keep records and unlock more.

The endpoint takes a credential two ways: the Authorization: Bearer header (preferred, works with every credential type) and a self-contained URL https://api.catalogian.com/v1/mcp?token=YOUR_SESSION_TOKEN for clients that cannot set headers. The URL form exists for ephemeral session tokens; prefer the header wherever the client supports one, because URLs are logged and shared more easily than headers.

Claude.ai (OAuth — Recommended)

Claude.ai supports OAuth — no API key or config file needed. Works on web and desktop. Configure once and it syncs everywhere.

  1. Open Settings → Connectors
  2. Click Add custom connector
  3. Enter Catalogian for the name
  4. Enter https://api.catalogian.com/v1/mcp for the URL
  5. Click Add
  6. Find Catalogian in the connectors list and click Connect — you'll be redirected to Catalogian to authorize, then returned to Claude.

After connecting, enable it per-conversation by clicking + in the chat input → Connectors → toggle Catalogian on.

Note: Configure once on Claude.ai web — settings sync automatically to Claude desktop and mobile.

ChatGPT

ChatGPT connects through its developer-mode plugin flow (the same guided flow the record's Connect surface offers):

  1. In ChatGPT: Settings → Security and login → turn on Developer mode (once)
  2. At chatgpt.com/plugins, press the plus and paste https://api.catalogian.com/v1/mcp as the server URL
  3. Choose the auth option: OAuth for a Catalogian account

Note: developer mode is the documented way to add custom MCP servers in ChatGPT today. The record page's Connect surface has a guided Connect ChatGPT button that copies the URL for you.

Connecting

The MCP endpoint accepts standard Streamable HTTP transport:

POST https://api.catalogian.com/v1/mcp
Authorization: Bearer <your-api-key>

Use a Catalogian API key (cat_live_...) from the Keys page in your dashboard. Create it with Read only scope — MCP only reads data, and a read-scoped key prevents any accidental writes via the REST API.

Key scopes

Catalogian credentials come in two scopes. A record key opens the one record it was created on. An account key opens every record the account owns. OAuth connections (Claude, ChatGPT) carry account credentials too: approving the consent screen grants the app account-wide reach.

CredentialReachesTiers
Record key (cat_src_...)One recordread-only, download-only
Account key (cat_live_...)Every record in your accountread-write, read-only, download-only
OAuth connection (Claude, ChatGPT)Every record in your accountread, write, admin, mcp:access

Record key - the least-privilege default. Created on a record's page, it works on that record only: whatever an agent does with it stays inside one record. Use it for partners and single-record agents.

Account key - the whole-library credential. Created from the account page, it works on every record you own. Pick the tier when you create it: read-write can also change records through the REST API, read-only is the safe default for agents, and download-only can fetch files and nothing else.

OAuth connection - the same account-wide reach, granted to an app instead of a key. The consent screen says what the app will be able to do before you approve it, and you can revoke the connection any time from Settings, Connected Apps. The OAuth scopes are catalogian:read, catalogian:write, catalogian:admin, and mcp:access (the last one is required for MCP access).

Why the account scope matters: an account key lets one agent connection see every record at once. Eight feed records in one Claude context for cross-feed comparison is one connection with an account key - not eight separate single-record setups.

Claude Desktop / API Key (Advanced)

The config-file method also works for Cursor and other editors. Add a config to your claude_desktop_config.json on macOS at ~/Library/Application Support/Claude/claude_desktop_config.json. Restart Claude Desktop. Your records will appear as available tools.

Ephemeral (no account): use the session token from the anonymous ingest above as the credential. This is the quickest way to try Catalogian in Claude Desktop:

{
  "mcpServers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer cat_eph_YOUR_SESSION_TOKEN"
      ]
    }
  }
}

Note: the session token expires with its session (24-hour life, 3 ingests, 50 tool calls). When the session ends, create a new one and replace the token in the config.

With an account: use a durable API key (cat_live_...) instead, so the connection survives session expiry and reaches every record:

{
  "mcpServers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer cat_live_your_key_here"
      ]
    }
  }
}

Note: Claude Desktop requires the mcp-remote proxy for HTTP MCP servers — the type: "http" syntax is not yet supported in stable builds. For local or staging testing over HTTP, add "--allow-http" as an additional arg after the URL. Create your API key with Read only scope from the Keys page — MCP only needs read access.

IDE Setup (Cursor, Windsurf, VS Code)

All three editors use the same mcp-remote bridge to connect to Catalogian's Streamable HTTP endpoint via STDIO. Each config below shows the ephemeral session token first (no account needed) and the durable account key second. Account keys are created with Read only scope from the Keys page; a session token comes from the anonymous ingest in the No account? section above.

Cursor

Add to .cursor/mcp.json in your project root. Ephemeral session first:

{
  "mcpServers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer cat_eph_YOUR_SESSION_TOKEN"
      ]
    }
  }
}

Or with an account API key:

{
  "mcpServers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer cat_live_YOUR_KEY"
      ]
    }
  }
}

Windsurf

Add to ~/.codeium/windsurf/mcp_config.json. Ephemeral session first:

{
  "mcpServers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer cat_eph_YOUR_SESSION_TOKEN"
      ]
    }
  }
}

Or with an account API key:

{
  "mcpServers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer cat_live_YOUR_KEY"
      ]
    }
  }
}

VS Code

Add to .vscode/mcp.json in your project root. Ephemeral session first:

{
  "servers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer cat_eph_YOUR_SESSION_TOKEN"
      ]
    }
  }
}

Or with an account API key. VS Code supports an inputs array so your key isn't stored in plaintext:

{
  "servers": {
    "catalogian": {
      "command": "npx",
      "args": [
        "-y",
        "mcp-remote",
        "https://api.catalogian.com/v1/mcp",
        "--header",
        "Authorization: Bearer ${input:catalogianApiKey}"
      ]
    }
  },
  "inputs": [
    {
      "id": "catalogianApiKey",
      "type": "promptString",
      "description": "Catalogian API key (cat_live_...)",
      "password": true
    }
  ]
}

Troubleshooting: All three editors require Node.js 18+ with npx available in your PATH. If the server fails to start, verify with npx --version in your terminal. For local or staging testing over HTTP, add "--allow-http" as an additional arg after the URL.

Available Tools

list_records

List all records monitored for this account. Start here to discover available recordId and recordSlug values needed by every other tool. Returns id, slug, name, url, status, format, lastCheckedAt, and vanityUrl for each record. Use the slug with other tools for more readable calls; a record whose slug is null has no slug, so use its id. No parameters required.

Example response:

[
  {
    "id": "cmmfyryl30005ygld7m5596gm",
    "name": "West Coast Earthquakes",
    "slug": "west-coast-earthquakes",
    "url": "https://earthquake.usgs.gov/earthquakes/feed/v1.0/summary/all_day.csv",
    "status": "active",
    "format": "csv",
    "lastCheckedAt": "2026-09-28T20:00:00.000Z",
    "vanityUrl": "https://catalogian.com/record/alex/west-coast-earthquakes"
  }
]

get_record_by_slug

Look up a single record by its human-readable slug and get full details including id, name, url, type, status, format, keyField, and timestamps. Use this when you know the slug but need the internal UUID or other metadata. Most other tools accept recordSlug directly, so you may not need this. Use list_records first if you don't know the slug.

Input:

{ "recordSlug": "west-coast-earthquakes" }

Example response:

{
  "id": "cmmfyryl30005ygld7m5596gm",
  "name": "West Coast Earthquakes",
  "slug": "west-coast-earthquakes",
  "url": "https://earthquake.usgs.gov/earthquakes/feed/v1.0/summary/all_day.csv",
  "type": "http",
  "status": "active",
  "format": "csv",
  "keyField": "id",
  "lastCheckedAt": "2026-09-28T20:00:00.000Z",
  "createdAt": "2026-09-28T10:00:00.000Z"
}

get_delta

Change events for a record: how many rows were added, changed, or deleted, and when. Returns deltaEventId values needed for get_delta_rows. Use list_records first to find recordId or recordSlug. Use get_delta_rows to retrieve the actual changed row data.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
sinceISO 8601 timestamp — only return events after this time
limitMax results (default 50, max 200)

Example response:

[
  {
    "id": "evt_01j...",
    "snapshotId": "snap_01j...",
    "recordId": "cmmfyryl30005ygld7m5596gm",
    "detectedAt": "2026-09-28T20:00:00.000Z",
    "newCount": 5,
    "changedCount": 3,
    "deletedCount": 1,
    "unchangedCount": 193,
    "totalCount": 202,
    "newKeys": ["nc75443740", "nc75443741"],
    "changedKeys": ["nc75443737"],
    "deletedKeys": ["nc75443001"]
  }
]

get_delta_rows

Get the actual row-level field data for a specific delta event: the rows that were added, changed, or deleted. Returns complete field values, previous values for changed rows, and field-level diffs. Call get_delta first to get deltaEventId values. Supports filtering by change type and pagination.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
deltaEventIdThe delta event ID (from get_delta results)
changeTypenew | changed | deleted | all (default: all)
limitMax rows per page (default 100, max 500)
offsetPagination offset (default 0)

Example response:

{
  "deltaEventId": "evt_01j...",
  "rows": [
    {
      "key": "nc75443737",
      "changeType": "changed",
      "fields": {
        "id": "nc75443737",
        "place": "7 km WNW of Cobb, CA",
        "mag": "1.4",
        "status": "reviewed"
      },
      "previousData": {
        "id": "nc75443737",
        "place": "7 km WNW of Cobb, CA",
        "mag": "1.2",
        "status": "automatic"
      },
      "fieldDiff": {
        "mag": { "before": "1.2", "after": "1.4" },
        "status": { "before": "automatic", "after": "reviewed" }
      }
    },
    {
      "key": "nc75443740",
      "changeType": "new",
      "fields": {
        "id": "nc75443740",
        "place": "12 km NE of Ridgecrest, CA",
        "mag": "2.6",
        "status": "automatic"
      }
    }
  ],
  "total": 17,
  "hasMore": false
}

get_snapshot_rows

Read raw rows from a record's latest ingested snapshot. Use this to browse live data page by page. Returns rowKey, row data, total row count, and a nextCursor for pagination. For filtering or searching, use filter_snapshot_rows or search_snapshot instead. Use list_records first to find recordId or recordSlug.

Input parameters:

ParamDescription
recordIdRecord ID (use get_record_by_slug to resolve from slug)
recordSlugRecord slug (alternative to recordId)
limitRows per page (default 20, max 100)
cursorPagination cursor from previous response (omit for first page)

Example response:

{
  "rows": [
    {
      "rowKey": "nc75443737",
      "data": {
        "id": "nc75443737",
        "time": "2026-09-28T10:11:39.970Z",
        "place": "7 km WNW of Cobb, CA",
        "mag": "0.75"
      }
    },
    {
      "rowKey": "nc75443736",
      "data": {
        "id": "nc75443736",
        "time": "2026-09-28T09:45:12.110Z",
        "place": "8 km S of The Geysers, CA",
        "mag": "1.1"
      }
    }
  ],
  "total": 202,
  "hasMore": true,
  "nextCursor": "nc75443736",
  "snapshotId": "abc123...",
  "snapshotCreatedAt": "2026-09-28T20:00:00.000Z"
}

download_snapshot

Export the full current snapshot of a record as a downloadable CSV or JSON file. Returns a pre-signed URL valid for 1 hour, plus row count and format info. Use this when the user wants the complete record data as a file. For a filtered subset, use download_filtered_snapshot instead. Use list_records first to find recordId or recordSlug.

Lifecycle: the 1 hour is the signature, not the storage. For a record in an anonymous session, the session sweeper deletes that session's export files shortly after the session ends (24-hour life), so a link can stop working before its hour is up. For account records the export file stays in storage; only the signed URL expires.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
formatcsv | json (default: csv)

Example response:

Your file is ready: https://r2.example.com/exports/usr_01j.../src_01j.../20260928T120000Z.csv?X-Amz-...

Rows: 202 | Format: csv | Expires: 2026-09-28T13:00:00.000Z

snapshot_schema

Get the column names, null rates, and example values for a record's latest snapshot. Always call this before filter_snapshot_rows, query_snapshot, search_snapshot, or sample_snapshot: field names in those tools must match exactly. Returns a fields array with name, null percentage, and sample values for each column. Use list_records first to find recordId or recordSlug.

Input: recordId or recordSlug

Example response:

Snapshot schema for source cmmfyryl30005ygld7m5596gm
Snapshot: snap_01j... (created 2026-09-28T20:00:00.000Z)
Total rows: 202 | Sampled: 100

Fields (7):
  id (null 0%) - e.g. "nc75443737", "nc75443736"
  latitude (null 0%) - e.g. 38.83, 40.51
  longitude (null 0%) - e.g. -122.80, -124.10
  mag (null 1%) - e.g. 0.75, 1.4, 2.9
  place (null 0%) - e.g. "7 km WNW of Cobb, CA"
  status (null 0%) - e.g. "automatic", "reviewed"
  time (null 0%) - e.g. "2026-09-28T10:11:39.970Z"

query_snapshot

Run a single server-side aggregation on a record's snapshot data: handles records with 100K+ rows efficiently without downloading. Operations: 'distinct' (unique values for a field), 'count' (total rows, or rows with a field), 'group_by' (value breakdown with counts), 'min'/'max'/'avg'/'sum' (numeric aggregations), 'filter' (row subset matching a condition). Call snapshot_schema first to get exact field names. For multi-condition filters, use filter_snapshot_rows instead.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
operationdistinct | count | group_by | min | max | avg | sum | filter
fieldField name from snapshot_schema (required for all except count)
filterOperatoreq | neq | contains | gt | lt | gte | lte (for filter)
filterValueValue to compare against (for filter)
limitMax results (default 100, max 500)
cursorPagination cursor for filter operation

Example: group by status:

{
  "recordSlug": "west-coast-earthquakes",
  "operation": "group_by",
  "field": "status"
}

Response:

{
  "type": "group_by",
  "field": "status",
  "groups": [
    { "value": "automatic", "count": 187 },
    { "value": "reviewed", "count": 15 }
  ]
}

filter_snapshot_rows

Filter rows from a record's current snapshot using one or more AND-joined conditions. Handles records with 100K+ rows server-side: only matching rows are returned. Supports text (eq, contains, starts_with), numeric (gt, lt, gte, lte), and null (is_null, is_not_null) checks. Call snapshot_schema first to get exact field names. Returns matched rows with pagination. For a single simple filter, query_snapshot with operation='filter' also works. To export filtered results as a file, use download_filtered_snapshot.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
conditionsArray of filter conditions (AND-joined). Each has field, operator, and optional value
conditions[].operatoreq | neq | contains | starts_with | gt | lt | gte | lte | is_null | is_not_null
fieldsFields to include in output (default: all): use to reduce response size on wide records
limitRows per page (default 20, max 100)
cursorPagination cursor from previous response

Example: find rows with mag > 2.5:

{
  "recordSlug": "west-coast-earthquakes",
  "conditions": [
    { "field": "mag", "operator": "gt", "value": "2.5" }
  ],
  "fields": ["id", "time", "place", "mag"]
}

Response:

Found ~12 matching rows (showing first 20)

Row 1 (key: nc75443901):
  id: nc75443901
  time: 2026-09-28T06:42:19.300Z
  place: 12 km NE of Ridgecrest, CA
  mag: 3.2

Row 2 (key: nc75443897):
  id: nc75443897
  time: 2026-09-28T04:17:55.610Z
  place: 21 km W of Petrolia, CA
  mag: 2.8

...

More results available - use cursor: "clxyz..."

download_filtered_snapshot

Export only the rows matching filter conditions as a downloadable CSV or JSON file. Returns a pre-signed URL valid for 1 hour. Use the same conditions format as filter_snapshot_rows. Call snapshot_schema first to get exact field names. Use download_snapshot instead if you want the complete unfiltered data. The URL lifecycle is the one described under download_snapshot above.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
conditionsArray of filter conditions (AND-joined), same as filter_snapshot_rows
formatcsv | json (default: csv)

Example: export reviewed earthquakes as CSV:

{
  "recordSlug": "west-coast-earthquakes",
  "conditions": [
    { "field": "status", "operator": "eq", "value": "reviewed" }
  ],
  "format": "csv"
}

Response:

Your filtered export is ready: https://r2.example.com/exports/usr_.../src_.../filtered-20260928T120000Z.csv?X-Amz-...

Matched rows: 15 | Format: csv | Conditions: 1
Snapshot date: 2026-09-28T20:00:00.000Z
Link expires: 2026-09-28T13:00:00.000Z

compare_snapshots

Compare two snapshots of a record to see exactly what rows were added, removed, or modified between them. By default compares the two most recent snapshots. Use sinceDate to compare against a historical point in time, or sinceSnapshotId for a specific snapshot. Returns row-level diffs with before/after values. Different from get_delta: this compares full snapshot contents, while get_delta shows pre-computed change summaries.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
sinceDateISO date (e.g. "2026-03-07") — compare latest snapshot against nearest snapshot before this date
sinceSnapshotIdCompare against this specific snapshot ID
limitMax changed rows to show (default 50, max 200)

Example: what changed since last Tuesday:

{
  "recordSlug": "west-coast-earthquakes",
  "sinceDate": "2026-09-21"
}

Response:

Comparing snapshots: 2026-09-21T20:00:00.000Z -> 2026-09-28T20:00:00.000Z

Added: 12 rows
Removed: 3 rows
Modified: 47 rows

ADDED (12):
  nc75443901: {"place":"12 km NE of Ridgecrest, CA","mag":"3.2",...}
  ...

REMOVED (3):
  nc75443001: {"place":"9 km SW of Livermore, CA","mag":"1.4",...}
  ...

MODIFIED (47):
  nc75443737:
    before: {"place":"7 km WNW of Cobb, CA","mag":"1.2"}
    after:  {"place":"7 km WNW of Cobb, CA","mag":"1.4"}

search_snapshot

Full-text keyword search across all fields in a record's snapshot. Returns rows containing the search terms anywhere in their data, ranked by relevance score. Use this when you want to find rows by natural language terms (e.g. names, descriptions, brands) without knowing which field contains the value. For structured filtering by specific fields, use filter_snapshot_rows instead. Use list_records first to find recordId or recordSlug.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
querySearch terms (e.g. "Cobb" or "Ridgecrest")
limitResults per page (default 20, max 100)
cursorPagination cursor from previous response
fieldsFields to include in output (default: all)

Example: find rows mentioning Cobb:

{
  "recordSlug": "west-coast-earthquakes",
  "query": "Cobb",
  "limit": 5
}

Response:

Found rows matching "Cobb" (ranked by relevance):

Result 1 (key: nc75443737, relevance: 0.075):
  place: 7 km WNW of Cobb, CA
  mag: 0.75
  status: automatic

Result 2 (key: nc75443736, relevance: 0.061):
  place: 9 km SSW of Cobb, CA
  mag: 1.1
  status: automatic

More results available - use cursor: "nc75443736"

sample_snapshot

Get a random or stratified sample of rows from a record's snapshot. Use this as a starting point for AI analysis of large records: avoids downloading the entire data set. Without stratifyBy, returns n random rows. With stratifyBy, returns rowsPerGroup rows per unique combination of the specified fields, ensuring representative coverage across groups. Call profile_snapshot first to find good stratification fields. Call snapshot_schema first to get exact field names.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
nTotal rows to return (1–500, default 50). Ignored when stratifyBy is set.
stratifyByFields to stratify by (max 3). Returns rowsPerGroup rows per unique combination.
rowsPerGroupRows per unique combination when stratifyBy is set (1–20, default 2). Total capped at 500.
fieldsFields to include in output (default: all)

Example: stratified sample by network and status:

{
  "recordSlug": "west-coast-earthquakes",
  "stratifyBy": ["net", "status"],
  "rowsPerGroup": 2,
  "fields": ["id", "place", "mag", "status"]
}

Response:

Sampled 12 rows from 202 total, stratified by net x status, 6 groups x 2 rows/group

Row 1 (key: nc75443737):
  id: nc75443737
  place: 7 km WNW of Cobb, CA
  mag: 0.75
  status: automatic

Row 2 (key: uw61324871):
  id: uw61324871
  place: 14 km E of Granite Falls, WA
  mag: 1.6
  status: reviewed

profile_snapshot

Analyze the data quality and distribution of every field in a record's snapshot. Returns cardinality (unique value count), null rates, type hints, and top values for each field, grouped by cardinality level. Also provides stratification recommendations for sample_snapshot. Use this to understand the data's structure before querying, or to identify key/identifier columns, data quality issues, and good fields for filtering or grouping. Use list_records first to find recordId or recordSlug.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
topNTop N values to show for low/medium cardinality fields (1–50, default 10)

Example:

{
  "recordSlug": "west-coast-earthquakes",
  "topN": 5
}

Response:

Record Profile: west-coast-earthquakes (snapshot: 2026-09-28, 202 rows, 22 fields)

-- LOW CARDINALITY ----------------------
  status (2 values, 0% null, text)
    automatic: 187  reviewed: 15

  type (2 values, 0% null, text)
    earthquake: 200  quarry blast: 2

-- MEDIUM CARDINALITY -------------------
  net (12 values, 0% null, text)
  magType (14 values, 0% null, text)

-- HIGH CARDINALITY ---------------------
  latitude (202 values, 0% null, numeric)
  longitude (202 values, 0% null, numeric)
  time (198 values, 0% null, text)

-- UNIQUE / IDENTIFIER ------------------
  id

Stratification recommendation:
  Best fields: status, net, magType
  Try: sample_snapshot(recordSlug: "west-coast-earthquakes", stratifyBy: ["status","net"], rowsPerGroup: 3)

get_health

Get the health score and diagnostic indicators for a record: whether it's active, recently checked, and responding correctly. Returns a 0-100 score and individual indicator flags. Use list_records first to find recordId or recordSlug. Use this to diagnose why a record might not be updating.

Input: recordId or recordSlug

Example response:

{
  "score": 100,
  "indicators": {
    "isActive": true,
    "etagSupported": true,
    "hasBeenChecked": true,
    "checkIsFresh": true
  }
}

list_editions

List the editions of a record: every ingest that created a new version, newest last. Each edition shows its row count and what changed vs the previous edition. Editions are numbered from the first ingest (edition 1) and are returned oldest to newest, so the list reads as the record's history. Use get_edition_diff to see the exact rows that changed between two editions.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
limitEditions per page (default 20, max 100)
cursorPagination cursor from previous response

Example response:

{
  "recordId": "cmmfyryl30005ygld7m5596gm",
  "editions": [
    {
      "edition": 1,
      "snapshotId": "snap_01j...",
      "createdAt": "2026-09-28T10:00:00.000Z",
      "kind": "first",
      "rowCount": 198,
      "delta": null
    },
    {
      "edition": 2,
      "snapshotId": "snap_02j...",
      "createdAt": "2026-09-28T20:00:00.000Z",
      "kind": "revision",
      "rowCount": 202,
      "delta": {
        "deltaEventId": "evt_01j...",
        "newCount": 5,
        "changedCount": 3,
        "deletedCount": 1
      }
    }
  ],
  "totalEditions": 2,
  "hasMore": false,
  "nextCursor": null
}

Note: the first edition has no previous edition to diff against, so its delta is null. A later edition's delta is also null when the change event was removed by data retention.

get_edition_diff

Diff two editions of a record. Returns rows that were added, changed, or removed between the two versions, with field-level before and after. Address the pair by edition number (fromEdition and toEdition, from list_editions) or by snapshot UUID (fromSnapshotId and toSnapshotId). Use list_editions first to see the available editions. Use get_delta_rows when you already have a specific change event id.

Input parameters:

ParamDescription
recordIdThe record ID (use this or recordSlug)
recordSlugThe record slug (alternative to recordId)
fromEditionThe older edition number (must be less than toEdition). Required unless fromSnapshotId and toSnapshotId are used
toEditionThe newer edition number. Required unless fromSnapshotId and toSnapshotId are used
fromSnapshotIdSnapshot UUID of the older edition (from list_editions). Expert alternative to fromEdition
toSnapshotIdSnapshot UUID of the newer edition. Expert alternative to toEdition
changeTypeFilter: new, changed, deleted, or all (default)
limitRows per page (default 100, max 500)
cursorPagination cursor from previous response

Example: what changed between the first and second ingest:

{
  "recordSlug": "west-coast-earthquakes",
  "fromEdition": 1,
  "toEdition": 2,
  "changeType": "changed",
  "limit": 50
}

Response:

{
  "recordId": "cmmfyryl30005ygld7m5596gm",
  "fromEdition": 1,
  "toEdition": 2,
  "fromSnapshotId": "snap_01j...",
  "toSnapshotId": "snap_02j...",
  "events": 1,
  "rows": [
    {
      "key": "nc75443737",
      "changeType": "changed",
      "fields": { "place": "7 km WNW of Cobb, CA", "mag": "1.4" },
      "previousData": { "place": "7 km WNW of Cobb, CA", "mag": "1.2" },
      "fieldDiff": { "mag": { "before": "1.2", "after": "1.4" } }
    }
  ],
  "total": 1,
  "hasMore": false,
  "nextCursor": null
}

Note: when the two editions are not adjacent, the diff covers every ingest between them, and a row that changed more than once shows the values from its most recent change (the response carries a note and an events count in that case). No change data is stored for pairs whose change events were removed by retention.

list_tools

List all available Catalogian MCP tools with short descriptions. Use this to discover what operations are available and choose the right tool for a task. Returns tool names and summaries. For parameter details, refer to each tool's own schema. No parameters. Useful as the first call in a fresh client to see what the 18 tools cover.

Example response (abridged):

[
  { "name": "list_records", "description": "START HERE: list all records with their ids and slugs needed by every other tool" },
  { "name": "get_delta", "description": "Get change event summaries for a record: returns deltaEventIds for get_delta_rows" },
  { "name": "snapshot_schema", "description": "Get field names and null rates - CALL FIRST before filtering, querying, or searching" },
  ...
]

Example Agent Workflow

When a user asks "What changed in my data today?", an agent can:

  1. Call list_records to find available records
  2. Call get_delta with since set to the start of today to find recent events
  3. Call get_delta_rows on the latest event with changeType: "changed" to get actual field data
  4. Summarize: "12 rows changed today: 5 new earthquakes added, 6 magnitude updates, 1 removed."

Example get_delta_rows call:

{
  "recordSlug": "west-coast-earthquakes",
  "deltaEventId": "evt_01j...",
  "changeType": "changed",
  "limit": 50
}

Snapshot Workflow

When working with snapshot data, follow this natural sequence:

  1. snapshot_schema — discover available fields and their types
  2. filter_snapshot_rows — preview matching rows (e.g. mag > 2.5 AND status = reviewed)
  3. download_filtered_snapshot — export just the matches as CSV/JSON
  4. query_snapshot — run aggregations (count, group_by, avg, etc.)
  5. download_snapshot — export the full snapshot as CSV/JSON

Recommended System Prompt

Add this to your Claude Desktop system prompt for optimal tool usage:

When working with Catalogian sources: always call snapshot_schema first to understand the data structure before querying or filtering. Use filter_snapshot_rows for finding specific records, query_snapshot for aggregations, download_filtered_snapshot to export just the matching rows, and download_snapshot to export the full snapshot.

Tips

  • • Use get_record_by_slug when you know the slug but not the internal ID
  • • Use the since parameter on get_delta for incremental polling — only fetch what's new
  • • Use changeType filter on get_delta_rows to focus on specific change categories
  • • Use snapshot_schema first to discover available fields before running query_snapshot or filter_snapshot_rows
  • • Use filter_snapshot_rows for multi-condition filtering — supports text, numeric, and null checks with AND logic
  • • Use query_snapshot for server-side aggregations on large records: much faster than paginating through all rows
  • • Use get_snapshot_rows to read current record data: paginate with cursor for large records
  • • Use download_filtered_snapshot to export only matching rows as CSV/JSON — same conditions as filter_snapshot_rows
  • • Use download_snapshot to export the entire snapshot as a CSV or JSON file — returns a pre-signed URL valid for 1 hour
  • • Use compare_snapshots to diff two snapshots — see what rows were added, removed, or modified between two points in time
  • • Use list_editions to see a record's ingest history: each edition is one ingest, numbered from the first
  • • Use get_edition_diff to see the exact rows that changed between two editions, with field-level before and after
  • • Use search_snapshot for full-text keyword search: it finds rows by name, description, or any text content, ranked by relevance
  • • Use profile_snapshot before sampling — it reveals field cardinality, null rates, and recommends fields for stratified sampling
  • • Use sample_snapshot with stratifyBy for representative samples across categories — much better than random sampling for analysis
  • • Results are paginated via limit + offset (deltas) or cursor (snapshots) — use hasMore to detect when to fetch more