LangChain agents can use Catalogian as a tool provider through the OpenAI Responses API endpoint. This guide walks through setting up a LangChain agent that can list records, query delta events, and browse snapshot data.
cat_live_): available once any record is on a paid tier (Watch or Max)pip install langchain langchain-openai openaiCatalogian implements the OpenAI Responses API format, so you can use the standard OpenAI Python SDK pointed at Catalogian's endpoint:
import os
from openai import OpenAI
catalogian = OpenAI(
base_url="https://api.catalogian.com/v1",
api_key=os.environ["CATALOGIAN_KEY"], # cat_live_...
)Define the tools your agent can use. Here are the most common ones:
CATALOGIAN_TOOLS = [
{
"type": "function",
"function": {
"name": "list_records",
"description": "List all records monitored for this account. Start here to discover available recordId and recordSlug values needed by every other tool. Returns id, slug, name, url, status, format, lastCheckedAt, and vanityUrl for each record. Use the slug with other tools for more readable calls; a record whose slug is null has no slug, so use its id.",
"parameters": {"type": "object", "properties": {}},
},
},
{
"type": "function",
"function": {
"name": "get_delta",
"description": "Change events for a record: how many rows were added, changed, or deleted, and when. Returns deltaEventId values needed for get_delta_rows. Use list_records first to find recordId or recordSlug. Use get_delta_rows to retrieve the actual changed row data.",
"parameters": {
"type": "object",
"properties": {
"recordSlug": {
"type": "string",
"description": "The record's human-readable slug (from list_records slug field). Alternative to recordId. Some records have no slug (slug is null): use recordId for those.",
},
"limit": {
"type": "integer",
"description": "Max events to return (default 10)",
},
},
},
},
},
{
"type": "function",
"function": {
"name": "profile_snapshot",
"description": "Analyze the data quality and distribution of every field in a record's snapshot. Returns cardinality (unique value count), null rates, type hints, and top values for each field, grouped by cardinality level. Also provides stratification recommendations for sample_snapshot. Use this to understand the data's structure before querying, or to identify key/identifier columns, data quality issues, and good fields for filtering or grouping. Use list_records first to find recordId or recordSlug.",
"parameters": {
"type": "object",
"properties": {
"recordSlug": {
"type": "string",
"description": "The record's human-readable slug (from list_records slug field). Alternative to recordId. Some records have no slug (slug is null): use recordId for those.",
},
},
},
},
},
{
"type": "function",
"function": {
"name": "search_snapshot",
"description": "Full-text keyword search across all fields in a record's snapshot. Returns rows containing the search terms anywhere in their data, ranked by relevance score. Use this when you want to find rows by natural language terms (e.g. names, descriptions, brands) without knowing which field contains the value. For structured filtering by specific fields, use filter_snapshot_rows instead. Use list_records first to find recordId or recordSlug.",
"parameters": {
"type": "object",
"properties": {
"recordSlug": {"type": "string"},
"query": {"type": "string", "description": "Search term"},
},
"required": ["query"],
},
},
},
]Create a function that dispatches tool calls to Catalogian:
import json
def call_catalogian_tool(tool_name: str, arguments: dict) -> str:
"""Execute a Catalogian tool via the Responses API."""
response = catalogian.responses.create(
model="catalogian-1",
input=json.dumps(arguments) if arguments else "{}",
tools=[{"type": "function", "name": tool_name}],
tool_choice={"type": "function", "name": tool_name},
)
# Extract the text content from the response
for item in response.output:
if hasattr(item, "content"):
for part in item.content:
if hasattr(part, "text"):
return part.text
return "No result"Wrap the Catalogian tools as LangChain tools and create an agent:
from langchain_core.tools import StructuredTool
from langchain_openai import ChatOpenAI
from langchain.agents import AgentExecutor, create_openai_tools_agent
from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
# Wrap each Catalogian tool as a LangChain tool
def make_langchain_tool(tool_def):
func_def = tool_def["function"]
props = func_def.get("parameters", {}).get("properties", {})
def invoke(**kwargs):
return call_catalogian_tool(func_def["name"], kwargs)
return StructuredTool.from_function(
func=invoke,
name=func_def["name"],
description=func_def["description"],
)
tools = [make_langchain_tool(t) for t in CATALOGIAN_TOOLS]
# Create the agent with your preferred LLM
llm = ChatOpenAI(model="gpt-4o", temperature=0)
prompt = ChatPromptTemplate.from_messages([
("system", """You are a data analyst with access to Catalogian,
a data source monitoring service. Use the available tools to answer
questions about data changes, source health, and data quality.
Always call profile_snapshot before analyzing a record for the first time."""),
MessagesPlaceholder("chat_history", optional=True),
("human", "{input}"),
MessagesPlaceholder("agent_scratchpad"),
])
agent = create_openai_tools_agent(llm, tools, prompt)
executor = AgentExecutor(agent=agent, tools=tools, verbose=True)# Ask about recent changes
result = executor.invoke({
"input": "What changed in west-coast-earthquakes in the last 24 hours?"
})
print(result["output"])
# Search for a specific row
result = executor.invoke({
"input": "Find all rows matching 'Cobb' in west-coast-earthquakes"
})
print(result["output"])
# Get a data quality report
result = executor.invoke({
"input": "Profile west-coast-earthquakes and tell me about data quality issues"
})
print(result["output"])Catalogian exposes 18 tools through both MCP and the Responses API. See the full list in the MCP Integration docs: all tools are available via the Responses API as well.
Best practice: Always call profile_snapshot before querying a record for the first time. It returns field names, types, cardinality, and null rates, giving the LLM context it needs to write effective queries. See Agent Best Practices.
Building with CrewAI instead? CrewAI Guide →