Building a Safer Way to Troubleshoot CockroachDB from Claude, ChatGPT and Cursor

Abhishek Dutta
Abhishek Dutta
Building a Safer Way to Troubleshoot CockroachDB from Claude, ChatGPT and Cursor

A read-only MCP gateway built on NitroStack lets developers investigate live CockroachDB performance from Claude, ChatGPT and Cursor, without handing an AI agent control of the database.

What this enables

  • An engineer can ask "staging feels slow, can you take a look?" and have the agent check running queries, inspect a plan, and walk the table access path in one conversation.
  • Writes, DDL, DML, and query cancellation are blocked at the application layer, not left to a prompt instruction.
  • CockroachDB's own Agent Skills give the investigation a repeatable procedure instead of letting the model improvise each time.
  • The same gateway runs as one hosted endpoint reachable from Claude Desktop, Claude Code, Cursor, ChatGPT, and NitroStudio.

The question worth asking before the tools

The project started from an ordinary engineering question: if someone tells an AI assistant their CockroachDB cluster feels slow, how much of the actual investigation can happen inside that conversation? Not a demo where the model explains what an EXPLAIN plan is in the abstract, and not a chatbot fed a database handbook — the goal was an assistant that looks at what the database is actually doing, inspects a real query plan, and helps the engineer work through the problem from there.

CockroachDB already exposes SQL MCP capabilities and Cockroach Labs has been building Agent Skills around workflows like live SQL triage and statement fingerprint profiling, so the job wasn't inventing another database intelligence layer. It was working out how those existing capabilities should come together as an actual troubleshooting experience. That became the Cockroach Live Triage Gateway: a remote MCP server, built on NitroStack, read-only by design.

Not every API operation deserves to be a tool

Exposing more tools doesn't automatically make an MCP application more useful. CockroachDB can report on databases, schemas, active statements, query plans, and cluster information, and the gateway does expose some of that directly — lower-level crdb_sql_* tools stay available for developers who already know exactly what they're looking for.

But most investigations don't start that specifically. They start as "staging feels slow, can you take a look?" — and at that point the user doesn't know whether they need the running-query endpoint, an explain plan, table metadata, or cluster information. Figuring that out is part of the job. So a second layer of tools was built around the investigation itself, rather than around the raw API surface:

  • crdb_running_queries_snapshot — checks current activity.
  • crdb_explain_plan — inspects a specific query.
  • crdb_table_access_path — shows how a table is being reached.
  • crdb_compare_explain — compares alternative query shapes.
  • crdb_cluster_overview and crdb_regions_layout — bring in distributed-cluster context.
  • crdb_triage_summary — pulls the findings together at the end.

In the demo, a single prompt — "staging feels slow, read-only only, check what's running, inspect the plan for a query against the rides table, look at the access path, tell me what stands out" — can move through several of these in sequence. An empty first snapshot isn't a failed tool call; it's information, and the agent continues to the query plan and access path instead of stalling. That's closer to how an engineer actually troubleshoots: follow the evidence you have, not the evidence you expected.

The database stays the source of truth throughout. The model can explain a plan, compare two plans, and reason about what to investigate next — but it doesn't get to invent the plan. When comparing SELECT * FROM rides WHERE city = 'san francisco' against SELECT id FROM rides WHERE city = 'san francisco', the gateway pulls both plans from the actual optimizer first and lets the model reason from real output, rather than asking it to speculate from two SQL strings.

A skill is a way of working, not just access

This was the less obvious part of the build. Cockroach Labs has published Agent Skills such as triaging-live-sql-activity and profiling-statement-fingerprints, and the gateway uses them alongside its own tools — a distinction that only makes sense once you've used both. A tool gives the agent access to something; a skill gives it a way of working. The gateway can fetch currently running statements, but the triage skill gives the agent a repeatable method for what to look at once it has that data. The gateway can return statement or plan data; the profiling skill carries the operational procedure for interpreting recurring workload behavior.

That idea is exposed inside the MCP server directly, through a crdb_live_triage prompt and a crdb://triage/playbook resource, so a client can load the playbook and work through an investigation instead of improvising the whole process from scratch each time. Giving an agent access to a system is only half the problem — the useful part starts when it also has a reliable operating procedure for that system. CockroachDB owns the database facts, CockroachDB's skills carry the domain knowledge for investigating them, and NitroStack provides the MCP application layer that brings the combined workflow into the AI client.

The safety model lives in the gateway, not the prompt

The demo prompt says "read-only only," but that sentence was never meant to be the thing actually preventing a mutation. The gateway runs read-only by default with CRDB_BLOCK_WRITES=true. DDL and DML are blocked. Query cancellation is blocked. If the user asks for one of those actions — or the model decides mid-investigation that cancelling something would help — the gateway refuses it regardless.

That distinction matters: a prompt is guidance for the model, a policy boundary is a rule enforced by the application, and for a tool working this close to production infrastructure, the second one is what actually holds. The gateway is also built to be inspectable rather than opaque — tools exist for checking gateway status and health, crdb_gateway_recent_audit exposes recent activity, every call is policy-checked and logged, and resources describe the gateway's policy and operating limits directly. None of that is glamorous, but it's the difference between an agent integration that feels like a toy and one that could reasonably sit near real infrastructure. The goal was never an autonomous DBA — it was letting the AI go surprisingly far into an investigation without an equally surprising blast radius.

Some results are worse as prose

A list of running queries should look like a table. An EXPLAIN plan should look like a plan. Cluster information is easier to parse with visual hierarchy than as narrated paragraphs. So the gateway includes widgets for exactly those moments — a running-query table, an explain-plan tree, and a cluster overview card — connected through NitroStack's Widget SDK. If the connected MCP client supports the UI layer, those components render alongside the conversation; if it doesn't, the same tool still returns structured data and the workflow continues normally. That portability was deliberate: the capability lives in the MCP server once, rather than as a separate CockroachDB integration rebuilt per client, and different clients decide how richly to render it.

Where NitroStack fits

The gateway is deployed on NitroCloud over Streamable HTTP, giving it one remote endpoint reachable from every client — Claude Desktop, ChatGPT as a remote MCP connector, Cursor, and NitroStudio for development and testing of the same tools and widgets. NitroStack's SDK provided the framework for the tool and resource definitions, including the prompt and playbook resource that expose CockroachDB's skills inside the server itself; the Widget SDK provided the rendering layer for the running-query table, plan tree, and cluster card; and NitroCloud is what turns the whole thing into a single hosted surface instead of a per-client integration. The application logic stops belonging to whichever AI interface happens to be popular that month.

Where this lands

The finished gateway is intentionally narrower than "AI database agent" might suggest. A connected client can ask what's happening on the cluster, inspect running queries, explain a statement, compare plans, check table access paths, review the regional layout, and work through a triage playbook — and the agent can help connect those pieces and explain what deserves a closer look. What it cannot do is quietly turn that investigation into a production change.

Reusable takeaway

Infrastructure agents get more useful by narrowing the job, not by maximizing the number of actions available. The gateway isn't interesting because it exposes a lot of CockroachDB — it's interesting because the read-only boundary is enforced as an application policy rather than a prompt instruction, and because tools and skills were kept as separate layers: one for access, one for procedure. For a privileged system, knowing when to stop is the part worth keeping even if everything else gets rebuilt.

Teams building a similar investigation layer over a production system can look at the NitroStack SDK and NitroCloud for defining a policy-enforced, read-only MCP surface and hosting it as one endpoint across multiple AI clients.