Building Custom MCP Servers with TypeScript: Complete Architecture & Deployment Guide

Table of Contents(18 sections)
The Model Context Protocol (MCP) defines a standardized interface for AI models to interact with external tools, resources, and context providers. This guide details the construction of a custom MCP server using TypeScript and the official @modelcontextprotocol/sdk, focusing on robust architecture, schema validation, transport mechanisms, and cloud deployment.
Core MCP Server Architecture
An MCP server fundamentally processes incoming MCPRequest objects and generates MCPResponse objects. The server's responsibilities include:
- Transport Layer: Handling communication (e.g., HTTP, stdio).
- Request Deserialization & Validation: Parsing and validating
MCPRequestpayloads against defined schemas. - Tool/Resource Orchestration: Executing requested tools or fetching specified resources.
- Response Serialization: Formatting
MCPResponseobjects.
We will implement a server that supports both stdio and Server-Sent Events (SSE) transports, leveraging Zod for schema validation and Google Cloud Run for deployment.
Project Setup
Initialize a new TypeScript project:
mkdir mcp-server-ts
cd mcp-server-ts
npm init -y
npm install typescript @types/node ts-node zod @modelcontextprotocol/sdk express @types/express @modelcontextprotocol/zod
npm install --save-dev @types/ws ws
npx tsc --init
Configure tsconfig.json for strictness and ESNext modules:
// tsconfig.json
{
"compilerOptions": {
"target": "es2021",
"module": "commonjs",
"rootDir": "./src",
"outDir": "./dist",
"esModuleInterop": true,
"forceConsistentCasingInFileNames": true,
"strict": true,
"skipLibCheck": true,
"resolveJsonModule": true
},
"include": ["src/**/*.ts"],
"exclude": ["node_modules"]
}
Schema Definition and Validation with Zod
The @modelcontextprotocol/zod package provides Zod schemas for MCP types, ensuring strict type safety and validation.
// src/schemas.ts
import {
MCPRequestSchema,
MCPResponseSchema,
ToolDefinitionSchema,
ResourceDefinitionSchema,
PromptTemplateDefinitionSchema,
ToolCallSchema,
ToolResultSchema,
ResourceFetchSchema,
ResourceDataSchema,
PromptTemplateRenderSchema,
PromptTemplateOutputSchema,
} from '@modelcontextprotocol/zod';
// Export all necessary schemas for use throughout the server
export {
MCPRequestSchema,
MCPResponseSchema,
ToolDefinitionSchema,
ResourceDefinitionSchema,
PromptTemplateDefinitionSchema,
ToolCallSchema,
ToolResultSchema,
ResourceFetchSchema,
ResourceDataSchema,
PromptTemplateRenderSchema,
PromptTemplateOutputSchema,
};
// Define custom tool/resource schemas if needed, extending base MCP types
// Example: A custom tool that takes a 'query' string
import { z } from 'zod';
export const SearchToolInputSchema = z.object({
query: z.string().describe("The search query to execute."),
});
export const SearchToolOutputSchema = z.object({
results: z.array(z.string()).describe("A list of search results."),
});
export const SearchToolDefinition = ToolDefinitionSchema.extend({
name: z.literal("search"),
description: z.literal("A tool for performing web searches."),
input_schema: SearchToolInputSchema,
output_schema: SearchToolOutputSchema,
});
export type SearchToolInput = z.infer<typeof SearchToolInputSchema>;
export type SearchToolOutput = z.infer<typeof SearchToolOutputSchema>;
Implementing Tools, Resources, and Prompt Templates
MCP servers expose capabilities through tools, resources, and prompt_templates.
// src/handlers.ts
import {
MCPRequest,
MCPResponse,
ToolCall,
ResourceFetch,
PromptTemplateRender,
ToolDefinition,
ResourceDefinition,
PromptTemplateDefinition,
} from '@modelcontextprotocol/sdk';
import {
SearchToolInput,
SearchToolOutput,
SearchToolDefinition,
} from './schemas';
// --- Tool Implementations ---
async function executeSearchTool(input: SearchToolInput): Promise<SearchToolOutput> {
console.log(`Executing search for: ${input.query}`);
// Simulate an external API call
await new Promise(resolve => setTimeout(resolve, 500));
const results = [
`Result 1 for "${input.query}"`,
`Result 2 for "${input.query}"`,
];
return { results };
}
// Map tool names to their execution functions
const toolExecutors: Record<string, (input: any) => Promise<any>> = {
'search': executeSearchTool,
};
// --- Resource Implementations ---
async function fetchDocumentationResource(id: string): Promise<string> {
console.log(`Fetching documentation for ID: ${id}`);
// Simulate fetching from a database or file system
await new Promise(resolve => setTimeout(resolve, 300));
switch (id) {
case 'mcp-overview':
return "The Model Context Protocol (MCP) standardizes AI model interaction with external systems.";
case 'sdk-usage':
return "The @modelcontextprotocol/sdk provides utilities for building MCP clients and servers.";
default:
throw new Error(`Resource with ID '${id}' not found.`);
}
}
// Map resource names to their fetch functions
const resourceFetchers: Record<string, (id: string) => Promise<any>> = {
'documentation': fetchDocumentationResource,
};
// --- Prompt Template Implementations ---
async function renderSummaryTemplate(variables: Record<string, any>): Promise<string> {
console.log(`Rendering summary template with variables: ${JSON.stringify(variables)}`);
const { document, length } = variables;
if (!document) throw new Error("Missing 'document' variable for summary template.");
// Simulate a complex templating engine
await new Promise(resolve => setTimeout(resolve, 100));
return `Here is a ${length || 'brief'} summary of the document: "${document.substring(0, 50)}..."`;
}
// Map prompt template names to their render functions
const promptTemplateRenderers: Record<string, (variables: Record<string, any>) => Promise<string>> = {
'summary': renderSummaryTemplate,
};
// --- MCP Request Handler ---
export async function handleMCPRequest(request: MCPRequest): Promise<MCPResponse> {
const response: MCPResponse = {
request_id: request.request_id,
tool_results: [],
resource_data: [],
prompt_template_outputs: [],
error: undefined,
};
try {
// Handle tool calls
if (request.tool_calls) {
for (const call of request.tool_calls) {
const executor = toolExecutors[call.name];
if (!executor) {
response.tool_results.push({
call_id: call.call_id,
error: `Tool '${call.name}' not found.`,
});
continue;
}
try {
const output = await executor(call.input);
response.tool_results.push({
call_id: call.call_id,
output: output,
});
} catch (toolError: any) {
response.tool_results.push({
call_id: call.call_id,
error: toolError.message || 'Tool execution failed.',
});
}
}
}
// Handle resource fetches
if (request.resource_fetches) {
for (const fetch of request.resource_fetches) {
const fetcher = resourceFetchers[fetch.name];
if (!fetcher) {
response.resource_data.push({
fetch_id: fetch.fetch_id,
error: `Resource '${fetch.name}' not found.`,
});
continue;
}
try {
const data = await fetcher(fetch.id);
response.resource_data.push({
fetch_id: fetch.fetch_id,
data: data,
});
} catch (resourceError: any) {
response.resource_data.push({
fetch_id: fetch.fetch_id,
error: resourceError.message || 'Resource fetch failed.',
});
}
}
}
// Handle prompt template renders
if (request.prompt_template_renders) {
for (const render of request.prompt_template_renders) {
const renderer = promptTemplateRenderers[render.name];
if (!renderer) {
response.prompt_template_outputs.push({
render_id: render.render_id,
error: `Prompt template '${render.name}' not found.`,
});
continue;
}
try {
const output = await renderer(render.variables);
response.prompt_template_outputs.push({
render_id: render.render_id,
output: output,
});
} catch (templateError: any) {
response.prompt_template_outputs.push({
render_id: render.render_id,
error: templateError.message || 'Prompt template rendering failed.',
});
}
}
}
} catch (e: any) {
console.error("Unhandled error in MCP request handler:", e);
response.error = e.message || 'Internal server error.';
}
return response;
}
// --- Server Capabilities (for /mcp/capabilities endpoint) ---
export const serverCapabilities = {
tools: [
SearchToolDefinition.parse({
name: 'search',
description: 'A tool for performing web searches.',
input_schema: { type: 'object', properties: { query: { type: 'string' } }, required: ['query'] },
output_schema: { type: 'object', properties: { results: { type: 'array', items: { type: 'string' } } }, required: ['results'] },
}),
] as ToolDefinition[],
resources: [
ResourceDefinition.parse({
name: 'documentation',
description: 'Provides access to internal documentation articles.',
id_schema: { type: 'string' },
data_schema: { type: 'string' },
}),
] as ResourceDefinition[],
prompt_templates: [
PromptTemplateDefinition.parse({
name: 'summary',
description: 'Generates a summary of a given document.',
variables_schema: {
type: 'object',
properties: {
document: { type: 'string', description: 'The document content to summarize.' },
length: { type: 'string', enum: ['brief', 'detailed'], description: 'Desired length of the summary.' }
},
required: ['document']
},
output_schema: { type: 'string' },
}),
] as PromptTemplateDefinition[],
};
Transport Layer: stdio vs. SSE
MCP supports various transports. We'll implement both stdio (for local development/CLI tools) and SSE (for HTTP-based, streaming interactions).
stdio Transport
The stdio transport reads MCPRequest from stdin and writes MCPResponse to stdout. Each message is prefixed by its length.
// src/stdioServer.ts
import { MCPRequestSchema, MCPResponseSchema } from './schemas';
import { handleMCPRequest } from './handlers';
import {
readMCPMessage,
writeMCPMessage,
MCPMessage,
} from '@modelcontextprotocol/sdk/stdio';
async function startStdioServer() {
console.log("MCP stdio server started. Waiting for input...");
process.stdin.on('data', async (chunk) => {
try {
const messages = readMCPMessage(chunk);
for (const message of messages) {
if (message.type === 'request') {
const parsedRequest = MCPRequestSchema.parse(message.payload);
console.log(`Received MCP Request: ${parsedRequest.request_id}`);
const response = await handleMCPRequest(parsedRequest);
writeMCPMessage(process.stdout, { type: 'response', payload: response });
} else {
console.warn(`Received unexpected MCP message type: ${message.type}`);
}
}
} catch (error: any) {
console.error("Error processing stdio input:", error);
// Attempt to send an error response if possible
if (error.request_id) { // If we can extract request_id from a partially parsed request
writeMCPMessage(process.stdout, {
type: 'response',
payload: {
request_id: error.request_id,
error: `Invalid MCP request: ${error.message}`,
},
});
} else {
// Fallback for unparseable requests
console.error("Failed to parse incoming MCP request. Cannot send specific error response.");
}
}
});
process.stdin.on('end', () => {
console.log("MCP stdio server stdin ended.");
});
process.stdin.on('error', (err) => {
console.error("MCP stdio server stdin error:", err);
});
}
if (require.main === module) {
startStdioServer();
}
Server-Sent Events (SSE) Transport
SSE provides a persistent HTTP connection for streaming responses. This is ideal for web-based clients.
// src/sseServer.ts
import express from 'express';
import { v4 as uuidv4 } from 'uuid';
import { MCPRequestSchema, MCPResponseSchema } from './schemas';
import { handleMCPRequest, serverCapabilities } from './handlers';
import {
MCPRequest,
MCPResponse,
MCPCapabilities,
} from '@modelcontextprotocol/sdk';
const app = express();
app.use(express.json()); // For parsing application/json
const PORT = process.env.PORT || 8080;
// Middleware to set SSE headers
const setSSEHeaders = (res: express.Response) => {
res.setHeader('Content-Type', 'text/event-stream');
res.setHeader('Cache-Control', 'no-cache');
res.setHeader('Connection', 'keep-alive');
res.setHeader('X-Accel-Buffering', 'no'); // Disable Nginx buffering
};
// Endpoint for server capabilities
app.get('/mcp/capabilities', (req, res) => {
res.json(serverCapabilities);
});
// SSE endpoint for MCP requests
app.post('/mcp/stream', async (req, res) => {
setSSEHeaders(res);
const requestId = uuidv4(); // Generate a unique ID for this request stream
try {
// Validate incoming request body
const mcpRequest: MCPRequest = MCPRequestSchema.parse({
request_id: requestId, // Override or ensure request_id is present
...req.body,
});
console.log(`Received SSE MCP Request: ${mcpRequest.request_id}`);
// Send initial "request_received" event
res.write(`event: request_received\n`);
res.write(`data: ${JSON.stringify({ request_id: mcpRequest.request_id })}\n\n`);
// Process the request
const mcpResponse = await handleMCPRequest(mcpRequest);
// Send the final "response" event
res.write(`event: response\n`);
res.write(`data: ${JSON.stringify(mcpResponse)}\n\n`);
} catch (error: any) {
console.error("Error processing SSE MCP request:", error);
const errorResponse: MCPResponse = {
request_id: requestId,
error: `Invalid MCP request or internal server error: ${error.message}`,
};
res.write(`event: error\n`);
res.write(`data: ${JSON.stringify(errorResponse)}\n\n`);
} finally {
res.end(); // Close the connection after sending response or error
}
});
// Health check endpoint
app.get('/healthz', (req, res) => {
res.status(200).send('OK');
});
export function startSseServer() {
app.listen(PORT, () => {
console.log(`MCP SSE server listening on port ${PORT}`);
});
}
if (require.main === module) {
startSseServer();
}
Main Server Entry Point
A single entry point to choose the server type based on environment variables.
// src/index.ts
import { startStdioServer } from './stdioServer';
import { startSseServer } from './sseServer';
const SERVER_TYPE = process.env.MCP_SERVER_TYPE || 'sse'; // Default to SSE
if (SERVER_TYPE === 'stdio') {
startStdioServer();
} else if (SERVER_TYPE === 'sse') {
startSseServer();
} else {
console.error(`Unknown MCP_SERVER_TYPE: ${SERVER_TYPE}. Must be 'stdio' or 'sse'.`);
process.exit(1);
}
Transport Comparison
| Feature | stdio | SSE (HTTP) |
|---|---|---|
| Protocol | Custom length-prefixed binary/text | HTTP/1.1, text-based |
| Connection | Persistent (stdin/stdout) | Persistent (single request/response stream) |
| Streaming | Bidirectional (via separate pipes) | Unidirectional (server to client) |
| Overhead | Minimal | HTTP headers, event framing |
| Complexity | Low (SDK handles framing) | Moderate (HTTP server, headers, event format) |
| Use Case | CLI tools, local agents, container-internal | Web clients, cloud functions, external services |
| Authentication | OS-level permissions | HTTP headers (Bearer tokens, API keys) |
| Error Handling | Application-level messages | HTTP status codes, application-level events |
| Scalability | Single process | Horizontally scalable (load balancers) |
Containerization with Docker
Containerize the application for consistent deployment.
# Dockerfile
# Use a slim Node.js image for smaller size
FROM node:20-slim AS builder
WORKDIR /app
# Copy package.json and package-lock.json first to leverage Docker cache
COPY package*.json ./
RUN npm install --omit=dev
# Copy source code
COPY . .
# Build TypeScript code
RUN npm run build
# --- Production Stage ---
FROM node:20-slim
WORKDIR /app
# Copy only necessary files from the builder stage
COPY --from=builder /app/package*.json ./
COPY --from=builder /app/node_modules ./node_modules
COPY --from=builder /app/dist ./dist
# Expose the port for the SSE server
EXPOSE 8080
# Set environment variable for SSE server type
ENV MCP_SERVER_TYPE=sse
# Command to run the SSE server
CMD ["node", "dist/index.js"]
Build the Docker image:
docker build -t mcp-server-ts .
Run locally (SSE):
docker run -p 8080:8080 mcp-server-ts
Test with curl:
curl -X POST -H "Content-Type: application/json" \
-d '{
"tool_calls": [
{ "call_id": "call-1", "name": "search", "input": { "query": "latest AI research" } }
],
"resource_fetches": [
{ "fetch_id": "fetch-1", "name": "documentation", "id": "mcp-overview" }
],
"prompt_template_renders": [
{ "render_id": "render-1", "name": "summary", "variables": { "document": "This is a very long document about the history of computing.", "length": "brief" } }
]
}' \
http://localhost:8080/mcp/stream
Expected output (streamed events):
event: request_received
data: {"request_id":"<uuid>"}
event: response
data: {"request_id":"<uuid>","tool_results":[{"call_id":"call-1","output":{"results":["Result 1 for \"latest AI research\"","Result 2 for \"latest AI research\""]}}],"resource_data":[{"fetch_id":"fetch-1","data":"The Model Context Protocol (MCP) standardizes AI model interaction with external systems."}],"prompt_template_outputs":[{"render_id":"render-1","output":"Here is a brief summary of the document: \"This is a very long document about the history of compu...\""}]}
Deployment to Google Cloud Run
Google Cloud Run is an ideal serverless platform for containerized applications, offering automatic scaling and pay-per-use billing.
Prerequisites
- Google Cloud Project configured.
gcloudCLI installed and authenticated.- Cloud Run API enabled.
Deployment Steps
-
Build and Push Docker Image to Google Container Registry (GCR) or Artifact Registry:
bash# For GCR (older, but widely used) docker tag mcp-server-ts gcr.io/<YOUR_PROJECT_ID>/mcp-server-ts:latest docker push gcr.io/<YOUR_PROJECT_ID>/mcp-server-ts:latest # For Artifact Registry (recommended) # Enable Artifact Registry API: gcloud services enable artifactregistry.googleapis.com # Create a repository: gcloud artifacts repositories create mcp-repo --repository-format=docker --location=us-central1 --description="MCP Docker repository" docker tag mcp-server-ts us-central1-docker.pkg.dev/<YOUR_PROJECT_ID>/mcp-repo/mcp-server-ts:latest docker push us-central1-docker.pkg.dev/<YOUR_PROJECT_ID>/mcp-repo/mcp-server-ts:latest -
Deploy to Cloud Run:
bashgcloud run deploy mcp-server-ts \ --image gcr.io/<YOUR_PROJECT_ID>/mcp-server-ts:latest \ --platform managed \ --region us-central1 \ --allow-unauthenticated \ --port 8080 \ --min-instances 0 \ --max-instances 10 \ --memory 512Mi \ --cpu 1 \ --set-env-vars MCP_SERVER_TYPE=sse \ --project <YOUR_PROJECT_ID>--allow-unauthenticated: For public access. For production, consider--no-allow-unauthenticatedand use IAM.--port 8080: Matches theEXPOSEinDockerfile. Cloud Run automatically routes traffic to this port.--set-env-vars MCP_SERVER_TYPE=sse: Ensures the SSE server starts.
-
IAM Authentication (Recommended for Production):
If
--no-allow-unauthenticatedis used, clients must authenticate. For service-to-service communication within GCP, use service accounts.-
Client Service Account: Create a service account for the client (e.g., an AI model service).
-
Grant Invoker Role: Grant the
roles/run.invokerrole to the client service account on your Cloud Run service.bashgcloud run services add-iam-policy-binding mcp-server-ts \ --member="serviceAccount:<CLIENT_SERVICE_ACCOUNT_EMAIL>" \ --role="roles/run.invoker" \ --region us-central1 \ --platform managed -
Client-side Authentication: When making requests from a GCP service, the
gcloudclient libraries orcurlwithgcloud auth print-identity-tokencan automatically handle authentication.bash# Example curl with authenticated token SERVICE_URL=$(gcloud run services describe mcp-server-ts --platform managed --region us-central1 --format 'value(status.url)') TOKEN=$(gcloud auth print-identity-token) curl -X POST -H "Content-Type: application/json" \ -H "Authorization: Bearer ${TOKEN}" \ -d '{ "tool_calls": [ { "call_id": "call-1", "name": "search", "input": { "query": "cloud run deployment" } } ] }' \ "${SERVICE_URL}/mcp/stream"
-
Production Gotchas & Troubleshooting
-
403 Forbiddenon Cloud Run:- Cause:
--no-allow-unauthenticatedwas used during deployment, but the client is not providing a validAuthorizationheader with a Cloud Run invoker token. - Fix:
- Ensure the calling service account has
roles/run.invokeron the Cloud Run service. - Verify the client is correctly generating and attaching an ID token (e.g., using
gcloud auth print-identity-tokenor Google Auth Libraries). - If public access is intended, redeploy with
--allow-unauthenticated.
- Ensure the calling service account has
- Cause:
-
500 Internal Server Error/Container instance crashed:- Cause: Application crash during startup or request processing. Common issues include incorrect
PORTenvironment variable, unhandled exceptions, or out-of-memory errors. - Fix:
- Check Cloud Run logs (
gcloud run services logs read mcp-server-ts --limit 100). Look forError: listen EADDRINUSE(port conflict, unlikely in Cloud Run),UnhandledPromiseRejectionWarning, orMemory limit exceeded. - Ensure your application listens on
process.env.PORT(Cloud Run injects this). OursseServer.tscorrectly usesprocess.env.PORT || 8080. - Increase memory (
--memory) or CPU (--cpu) if logs indicate resource exhaustion. - Add more robust
try...catchblocks around asynchronous operations.
- Check Cloud Run logs (
- Cause: Application crash during startup or request processing. Common issues include incorrect
-
400 Bad Requestfrom/mcp/stream:- Cause: The incoming JSON payload does not conform to
MCPRequestSchema. This often happens due to missing required fields or incorrect data types. - Fix:
- Review the client's request body against
MCPRequestSchema(and any custom schemas). - Add more detailed logging in the server's
catchblock forMCPRequestSchema.parseto output Zod validation errors. - Example:
typescript
import { ZodError } from 'zod'; // ... inside try/catch for MCPRequestSchema.parse } catch (error: any) { if (error instanceof ZodError) { console.error("Zod validation error:", JSON.stringify(error.errors, null, 2)); // ... send detailed error response } else { console.error("Non-Zod error:", error); } }
- Review the client's request body against
- Cause: The incoming JSON payload does not conform to
-
Slow Responses / Timeouts:
- Cause: Long-running tool executions, resource fetches, or prompt template renders. Cloud Run has a default request timeout (e.g., 5 minutes).
- Fix:
- Optimize handler logic.
- Implement asynchronous patterns for very long tasks (e.g., offload to Cloud Tasks or Pub/Sub, and have the MCP server poll for results or receive webhooks).
- Increase Cloud Run request timeout (
--timeout). - Ensure external API calls have appropriate timeouts.
-
SSE Connection Dropping Prematurely:
- Cause: Proxies (like Nginx or Cloud Load Balancers) might buffer responses, preventing immediate streaming. Cloud Run itself handles SSE well, but intermediate proxies might interfere.
- Fix:
- Ensure
X-Accel-Buffering: noheader is set (oursetSSEHeadersfunction does this). - Verify no other proxy in front of Cloud Run is buffering.
- Keep-alive mechanisms: While SSE is inherently keep-alive, if the server is idle for too long, some network components might close the connection. Consider sending periodic "heartbeat" comments (
: comment\n\n) if the server might be idle for extended periods between actual data events.
- Ensure
Frequently Asked Questions
Q1: How do I add a new custom tool or resource to my MCP server?
A1:
- Define Schema: Create a Zod schema for the tool's input and output (or resource's ID and data) in
src/schemas.ts. ExtendToolDefinitionSchemaorResourceDefinitionSchema. - Implement Logic: Write an asynchronous function in
src/handlers.tsthat takes the parsed input and returns the output. - Register Executor: Add your new function to the
toolExecutorsorresourceFetchersmap insrc/handlers.ts. - Update Capabilities: Add your
ToolDefinitionorResourceDefinitionto theserverCapabilitiesarray insrc/handlers.tsso clients can discover it. - Rebuild and Deploy: Rebuild your Docker image and redeploy to Cloud Run.
Q2: Can I use WebSockets instead of SSE for bidirectional streaming?
A2: While MCP itself is transport-agnostic, the @modelcontextprotocol/sdk currently provides specific helpers for stdio and SSE. WebSockets would require a custom implementation of the transport layer. You would need to:
- Set up a WebSocket server (e.g., using
wslibrary with Express). - Implement message framing (e.g., JSON messages with
typeandpayloadfields) over the WebSocket. - Adapt
handleMCPRequestto process incoming WebSocket messages and send responses back over the same connection. This is feasible but requires more manual work than the provided SSE/stdio examples.
Q3: How do I manage secrets (e.g., API keys for external tools) in a Cloud Run MCP server?
A3:
- Cloud Secret Manager: Store secrets in Google Cloud Secret Manager.
- Grant Access: Grant the Cloud Run service account (
<YOUR_PROJECT_ID>@appspot.gserviceaccount.comby default) theSecret Manager Secret Accessorrole on the specific secrets. - Mount as Environment Variables: Configure Cloud Run to mount secrets as environment variables during deployment:
Your application can then accessbash
gcloud run deploy mcp-server-ts ... \ --set-secrets=EXTERNAL_API_KEY=EXTERNAL_API_KEY:latest \ ...process.env.EXTERNAL_API_KEY. - Direct Access (less common): Alternatively, your application can directly call the Secret Manager API using the Google Cloud client libraries, but environment variables are generally simpler for configuration.
Q4: What if my tool execution takes longer than the Cloud Run request timeout?
A4: For long-running operations (e.g., complex ML model inference, large data processing), direct synchronous execution within the MCP request is not suitable. Consider these patterns:
- Asynchronous Task Queue: The MCP server initiates a long-running task (e.g., by publishing a message to Google Cloud Pub/Sub or creating a Cloud Tasks entry). It then immediately returns an
MCPResponseindicating the task has been accepted, possibly with atask_id. - Polling: The client can then periodically poll a separate endpoint on your MCP server (or another service) with the
task_idto check for completion and retrieve results. - Webhooks: The long-running task, once complete, can send a webhook notification to a dedicated endpoint on your MCP server (or another service) to push results back to the client if the client maintains a persistent connection or has a way to receive asynchronous updates. This decouples the request-response cycle from the actual execution time.
Q5: How can I test my stdio server locally?
A5: You can pipe JSON data to its stdin.
- Compile the server:
npm run build - Prepare a request file:
json
// request.json { "request_id": "test-stdio-1", "tool_calls": [ { "call_id": "call-stdio-1", "name": "search", "input": { "query": "stdio transport" } } ] } - Send the request using
stdio-clientfrom@modelcontextprotocol/sdk:Alternatively, you can manually construct the length-prefixed message:bash# Install stdio-client globally or locally npm install -g @modelcontextprotocol/sdk # Run the server in one terminal node dist/index.js # In another terminal, send the request stdio-client --request-file request.json --server-command "node dist/index.js"bash# In one terminal: node dist/index.js # In another terminal: # Get the length of the JSON string # echo '{"request_id":"test-stdio-1","tool_calls":[{"call_id":"call-stdio-1","name":"search","input":{"query":"stdio transport"}}]}' | wc -c # (Let's say it's 123 bytes) # Then send: echo -n -e '\x00\x00\x00\x7B{"request_id":"test-stdio-1","tool_calls":[{"call_id":"call-stdio-1","name":"search","input":{"query":"stdio transport"}}]}' | node dist/index.js # Note: The \x00\x00\x00\x7B represents the 4
Free In-Browser Developer Tools
Clean AI CLI logs, build cron expressions, decode JWTs, and calculate chmod permissions offline.
Related Articles

Building a Custom MCP Client: Connecting Any LLM to Multiple Model Context Protocol Servers
Comprehensive guide covering building a custom mcp client: connecting any llm to multiple model context protocol servers with production-grade architecture and code examples.
Read more
Building Your First MCP Server from Scratch: The Complete Python & Claude Guide
Step-by-step guide to building production Model Context Protocol (MCP) servers with Python, FastMCP, typed tools, resources, and Claude Desktop integration.
Read more
DeepSeek-R1 & Distilled Reasoning Models: Local vLLM Deployment, Quantization & Architecture
Comprehensive guide covering deepseek-r1 & distilled reasoning models: local vllm deployment, quantization & architecture with production-grade architecture and code examples.
Read more