How We Turned StepFun's Model Stack Into One MCP Gateway

Abhishek Dutta
Abhishek Dutta
How We Turned StepFun's Model Stack Into One MCP Gateway

We used NitroStack to build a single MCP application around StepFun's models, research workflows, computer use and audio, then deployed it in separate fixture and live environments so people could actually try the whole thing.

What this enables

  • An AI client gets one place to discover StepFun's capabilities: routing, generation, reasoning, code, research, computer use, and audio.
  • Heavy capabilities stay in their own services and are called over HTTP, so the gateway isn't forced to own everything.
  • A fixture environment gives deterministic demos, while a live environment reaches real StepFun models through NitroCloud.
  • Claude, ChatGPT, Cursor, and NitroStudio all connect to the same server instead of needing a separate integration each.

Why a few model calls wasn't enough

The obvious version of this project was to expose a handful of model calls and stop: generate text, reason through a problem, write code, put them behind MCP tools. Technically, StepFun would then be reachable from Claude, ChatGPT, Cursor, or any other MCP client.

But that would have been a thin version of what StepFun is. The interesting part isn't a single model endpoint. It's the routing between models, general generation, reasoning, coding, research, computer-use work, audio, and realtime experiences. An MCP layer built around that should feel like an entry point into the broader platform, not a wrapper around one chat-completions API.

The result is stepfun-mcp, a NitroStack application that gives AI clients one place to discover StepFun capabilities and call them through ordinary MCP tools. The gateway itself is built on @nitrostack/core.

The gateway doesn't try to own everything

An MCP server is tempting to make responsible for every capability behind it. This one deliberately isn't. The main StepFun MCP application owns the interface agents see: routing, intelligence tools, the Step to the Moon challenge, and the MCP-facing wrappers around the other capabilities. The research pipeline, computer-use engine, and audio system remain separate applications, and the gateway calls them over HTTP rather than importing their engines and pretending it is one codebase.

From a client's point of view that's still one MCP server. Behind it, each system stays responsible for what it's actually good at:

  • Intelligence: stepfun_route_task asks the router which model should handle a task, and stepfun_generate, stepfun_reason, and stepfun_code cover generation, reasoning, and code, without a different integration pattern for each.
  • Research: search and reasoning, source comparison, summarization, claim extraction, and report generation are exposed as MCP workflows, handled by the separate RAG pipeline.
  • Computer use: analyze_screen, understand_image, inspect_ui, and read_document sit in the MCP surface, while the computer-use service owns the work behind the HTTP boundary.
  • Audio: speech synthesis, transcription, and realtime-session creation remain the audio service's responsibility.

The payoff is one agent-facing product without forcing every underlying capability into one giant server.

Fixture and live environments, with the limits stated

The project runs two deployed MCP instances. One is built primarily for fixtures; the other is the live-test version that calls real StepFun models through NitroCloud. The distinction is kept visible rather than hidden to make the demo look cleaner.

The fixture instance returns deterministic results for the challenge, routing, and intelligence demos, along with fixture-backed computer-use and audio behavior. That's useful when someone needs to understand the shape of the application without spending API calls or depending on model variance. The live instance is where routing and intelligence actually reach StepFun models: stepfun_generate produces a real response, stepfun_reason sends the problem through the model at the requested reasoning effort, and stepfun_code returns generated code rather than a canned echo.

Two limitations are documented plainly:

  • The current NitroCloud live-test gateway accepts text content but rejects requests containing image_url. Text-only document reading works through the live computer-use path, but image-dependent tools such as screen analysis and UI inspection do not currently work live through that gateway. They work in fixtures.
  • Both deployed instances point to one shared RAG service running in live gateway mode, so even on the instance named stepfun-mcp-fixture, the five research operations call a real StepFun model. Research is the exception to the keys-free fixture setup. A genuinely zero-cost local demo is possible by running the RAG pipeline locally in fixture mode and pointing a local StepFun MCP server at it; the deployed setup simply can't run a second dedicated RAG instance right now.

Step to the Moon: an onboarding challenge, not a benchmark

A long list of MCP tools says what exists, but doesn't make anyone want to try it. For the MCPCon activation, the gateway includes a small challenge called Step to the Moon: five tasks, each aimed at a different skill — routing, reasoning, vision-shaped work, RAG-shaped work, and code.

A user can list the challenges, run them, and open a scorecard showing how many were completed, the average score, elapsed time, and a rank. The results are deterministic fixtures, so running the same task twice returns the same result. That's intentional: the challenge is an onboarding experience for the MCP application, not a benchmark of StepFun's models, and it's explicitly not meant to be quoted as one.

The scorecard is also where a visual widget earns its place. Rather than ending the interaction with another block of JSON, the MCP client can render the result as an actual card — some responses are fine as text, others clearly want a UI, and NitroStack's Widget SDK covers both without a separate frontend built around the server.

Where NitroStack fits

The core gateway is built with the NitroStack SDK on @nitrostack/core. NitroCloud hosts both deployed environments, and NitroStudio gives the team a place to exercise tools during development without needing a final client setup each time. The Widget SDK handles the Step to the Moon scorecard.

NitroStack doesn't replace StepFun's models. StepFun continues to own the intelligence, its research service owns research, its computer-use service owns computer use, and its audio service owns speech. What NitroStack provides is the layer that makes those capabilities understandable and callable as one MCP application, and that stays the same regardless of which client connects to it.

Reusable takeaway

When a company has several AI capabilities, the most useful MCP strategy isn't always to turn every API endpoint into a tool and call it a platform. Sometimes the better product sits one level above them and answers a simpler question: what should an agent be able to ask this company to do? For StepFun, that answer became routing, intelligence, research, computer use, and audio, all through one MCP server, with separate services kept separate underneath, and with the live-versus-fixture limits stated rather than smoothed over.

Teams building a similar gateway over several existing AI services can look at the NitroStack SDK and NitroCloud for defining one agent-facing MCP surface and hosting it across fixture and live environments.