Manufact

How to Test MCP Auth and UI in Claude and ChatGPT

Luigi Pederzani
Luigi PederzaniCo-founder
How to Test MCP Auth and UI in Claude and ChatGPT

A remote MCP server can pass a protocol test and still behave poorly when a person uses it in an AI client. The protocol test checks the server. A real client test checks whether the model chooses the right tool, supplies useful arguments, handles authorization, and displays any UI correctly. To ship confidently, run both.

Start with a protocol baseline

Use the official MCP Inspector to connect to a local or remote server. Check initialization, list the available tools, and call each tool with valid and invalid inputs. Record the expected result, including error responses. This catches problems that do not require a model to reproduce.

npx @modelcontextprotocol/inspector

For repeatable checks, use the Inspector CLI or a protocol test in CI. Keep these checks deterministic: assert the tool name, input schema, response structure, and authorization response. The Inspector CLI documentation gives the current command syntax.

Then test the actual client experience

Connect the same server to each client you support, such as Claude and ChatGPT. Use the same prompts, account permissions, and data fixtures for each run. Do not assume a result in one client predicts a result in another: models, host interfaces, and connection flows can differ. Treat those differences as things to measure, not facts to insert into a test plan in advance.

For each scenario, record the prompt, expected tool, expected arguments, actual tool call, final answer, and any UI or authorization problem. A useful starter matrix is:

ScenarioExpected resultClaude observationChatGPT observation
Direct read requestCalls the intended read tool with the right identifierRun and recordRun and record
Ambiguous requestAsks for missing information or chooses the appropriate toolRun and recordRun and record
Write actionRespects your product's authorization and confirmation rulesRun and recordRun and record
Unauthorized accountReturns a clear, safe errorRun and recordRun and record
Tool failureExplains the failure without inventing a successRun and recordRun and record

Test authorization with separate accounts

Use at least two tenants and users with different permissions. Verify that a token for tenant A cannot access tenant B, that a missing or expired token is rejected, and that a token with insufficient scope cannot execute a protected action. Test these on the server as well as through the client. A friendly client error is useful, but the server must enforce the boundary. OpenAI's authentication guidance also calls for verification of token, scopes, and audience on every invocation.

Test tool descriptions and server instructions

A valid schema does not guarantee that a model will choose a tool. Give each tool a specific name and description: say what it returns, when to use it, and which inputs are required. Prefer a small set of distinct tools over multiple tools that sound interchangeable. Run realistic prompts, including requests the server should decline or clarify.

If the server sends initialization instructions, put the most important routing guidance first. OpenAI's MCP server guide recommends placing key details in the first 512 characters. After changing a tool schema or instructions, reconnect each test client and verify what it actually sees; refresh behavior depends on the client and its configuration.

Test MCP Apps UI where it is used

If your tool returns an MCP Apps interface, test it in each host that your users will use. Check narrow and wide layouts, light and dark themes, loading and error states, keyboard interaction, and what happens when the host lacks an optional capability. A protocol level UI check is valuable, but a real host screenshot catches clipping and broken interactions.

Make the result repeatable

Keep protocol checks in CI. Run the client matrix before launch and after changes to tools, auth, or UI. Save screenshots and request traces with sensitive values removed, then compare failures against the prior run. The result is an evidence based release decision rather than an assumption that a passing curl command means the user experience works.

Manufact's mcp-use Tunnel can expose a local server through a stable public address for client testing. You can also test a preview deployment before connecting the production endpoint.

Share