Troubleshooting an MCP connection that won't work

A bottom-up debugging ladder for remote MCP servers — reachability, transport, the auth header, stale tool lists, logs — plus the RelayLink symptoms that look like failures and aren't.

11 min read Updated

The error you see is almost never generated at the layer that broke. A client that cannot resolve your hostname, one that got a 401, and one that connected perfectly but is showing a tool list from last Tuesday all surface the same way — the tools aren't there, or a call fails, and the message blames something vague.

So stop reading the message and work up from the bottom. Five rungs, in order, and don't skip one for looking too obvious to be wrong.

Rung 1 — is anything listening at that URL?

Take the client out of it:

curl -i https://your-server.example/mcp

A DNS or connection-refused error means the host is wrong or nothing is deployed. A TLS error means you are talking to something, but not to what you think. A 404 means the host is right and the path is wrong — a missing or extra path segment is the most common version. A bare GET may return a protocol-level complaint rather than a page — not a fault, just something MCP-shaped answering.

A 401 is good news. You reached the server and its authentication code ran.

For RelayLink, a credential-less request to /mcp returns a 401 with an empty body. One curl proves host, TLS, path, and that the authentication code ran. Read the response headers rather than the body, because the WWW-Authenticate header is the one thing the server will tell you. Bearer resource_metadata="…" is the machine-readable answer to "where do I get a token", and it is the only challenge RelayLink sends. Whether your client can use it is a fact about the client: Claude's and ChatGPT's connectors sign in that way, and so do Claude Code, Codex, VS Code and Zed, which publish who they are. Other coding tools and most frameworks cannot yet, and send a key in X-RelayLink-Key instead. That is worth knowing before you spend an hour on a sign-in that will not start.

Rung 2 — does your client speak remote at all?

MCP servers come in two shapes: a local process the client launches and talks to over standard input and output, or a remote HTTP endpoint you point a URL at. They are not interchangeable, and a client configured for one and given the other fails confusingly rather than loudly.

Confirm your client supports remote servers, and that you filled the field it expects — a URL where it wants a URL, not a URL where it wants a command. If you are wiring this from code, connecting an agent framework covers the same ground with the config in front of you, and what streamable HTTP is explains the transport a remote server uses.

There is a second version of this rung that looks identical and isn't: a client that speaks remote perfectly well but won't let you name a server. ChatGPT is the one you'll meet. Adding an arbitrary MCP server URL requires developer mode — a toggle in its settings, currently beta and on paid web plans — and with it off there is no field to fill in, which reads as "this app doesn't do remote servers" when the truth is that it does and hasn't been permitted to here. Check for a place to put the URL before debugging the URL. Claude has no equivalent gate, so a setup that works in one app and appears unsupported in the other is usually this and not your server.

The general form is worth carrying past this example: capability and permission fail the same way from the outside. Before concluding a client can't do something, establish that it has been allowed to.

RelayLink is remote-only, over HTTP, in stateless mode. It identifies itself as relaylink, with a version naming the tool contract rather than a release — useful confirmation you are talking to what you meant to, and whoami states the same version beside the build and the tool count.

Rung 3 — the header, character by character

Most "it won't connect" is one wrong string.

HTTP header names are case-insensitive. Values are where you have to be exact, and for RelayLink "exact" now means exact.

A leading space, a trailing space, a newline, a tab, a smart quote from a pasted document, or the wrong case will all fail. RelayLink compares a hash of exactly the bytes you send, so there is no forgiveness in any direction.

This is worth calling out because it used to be untrue, and the change is invisible from the outside. RelayLink previously matched the key inside SQL Server, under a collation that ignores case and trailing spaces — so KEY and key both authenticated as key. Storing a hash instead removed that, along with the uncomfortable implication that the effective keyspace was smaller than the key length suggested. If a key that worked for months stopped working after an upgrade, this is the first thing to check: compare what your client sends against what was issued, character by character, and look at the case and the end of the string.

The forgiving half is the one to watch. A key that works despite being mistyped is one you will re-enter differently somewhere else and lose an afternoon to — and it means the string you configured is not the string being compared. Open it in a plain text editor and check both ends.

Three traps:

  • A prefix you did not add. Some clients assume OAuth and prepend Bearer to whatever you type. RelayLink's key is a bare value in a header named X-RelayLink-Key — not an Authorization header, no prefix.
  • A field that never saved. If the client insists it is sending a header and the server says nothing arrived, either the field didn't persist or a proxy stripped it.
  • The wrong half of the problem, and no help from the server. An absent header and a wrong key both return a 401 with no body. RelayLink used to distinguish them in the response body and no longer can: once a request may authenticate either by key or by OAuth token, every authentication scheme gets asked to explain itself on the same response, and the first one to write a body prevents the others from replying at all. So the server says less than it did, and its WWW-Authenticate names the sign-in rather than the header. Distinguish the two yourself by sending a deliberately wrong key: if that also 401s, your header is arriving and the key is the problem; if the client cannot make the header arrive at all, nothing about the key matters yet.

Rung 4 — the client is showing you a stale tool list

Clients fetch a server's tool list when the connection is established, then work from that copy. Deploy a new tool, or fix a broken header, and the session you are in may never notice. MCP does have a notifications/tools/list_changed message for exactly this, and it reaches only clients that are still connected — which a client whose server has just been redeployed is not.

The tell is confidence, not an error. A stale client does not report a failure or an empty list. It enumerates the tools it has, accurately, and then answers questions about the ones it lacks as though it had asked the server. "There's no tool for that — that's only on the website" is the same sentence whether the tool is genuinely absent or merely missing from the copy the client is holding, and nothing in the reply distinguishes the two. Treat a confident absence as a claim to check rather than an answer.

A new conversation is usually not enough. The list is cached against the connector rather than the chat, so the fix is to disconnect the server and connect it again — Settings → Connectors in the desktop and web apps, /mcp in Claude Code. Where that lives moves as these apps evolve; the constant is that a fresh connection re-reads the list and an old one does not.

Ask the server, not the model. A server that publishes a generated reference gives you a second opinion that does not run through the client at all. RelayLink's /docs and /llms-full.txt are built by reflection over the same type tools/list answers from, so they cannot be shorter than the truth: if they name a tool your assistant does not, the client is stale, and if they do not name it either, the deploy is what to look at. Both are cached for an hour, so add a query string when you check by hand.

Then confirm with a call, not another list. Ask for something the old copy could not do and watch whether a tool runs. If the assistant answers from the conversation instead, the list has not refreshed — an assistant reciting its own cached inventory looks identical either way.

After a deploy, wait for the tools rather than for the green tick. A platform that keeps the old container serving while the replacement warms will report a deployment healthy before the new tool list exists anywhere. On RelayLink's own sandbox the deploy job's liveness check passed about two and a half minutes before the reference page began naming the tools that deploy had added. Poll the generated reference until it changes; that is the moment to reconnect, and reconnecting before it is how you end up with a fresh connection to the old build.

The same shape one level up, and cheaper to fix: the model holds a copy of whatever list_contacts returned earlier in the conversation, and a contact accepted or renamed since then — by you in a browser, or through the tools in another session — is not in it. RelayLink's instructions tell the assistant to try the name against the server anyway and re-read the list before saying somebody is not a contact; if it says so without calling anything, asking it to check your contacts again is enough. No reconnect needed — the server's list was right all along.

Rung 5 — read the logs on the correct side

Client-side connector logs tell you whether a request left. Server-side logs tell you whether it arrived. Only one is answered where you are looking.

If you own the server, log the request path and whether the credential was present at the auth boundary. That single line separates "never arrived" from "arrived and was rejected" — the fork this ladder is hunting for. If you don't own it, the closest equivalent is the status you can reproduce with curl, so rule out the boring layer before rewriting your connector setup.

Some of what gets reported as a broken connection is correct behaviour.

  • list_contacts comes back empty. Expected on a new account. It lists standing accepted contacts, and correspondence never creates one — somebody asks and somebody accepts, which is what invite_contact and accept_contact are for. Writing to an address with no RelayLink account needs no pair at all — that path is ordinary email, with its own limits.

  • Every tool answers ERROR: and mentions a waitlist. The connection is fine and the account is real; it is waiting. A deployment can be running a founder season, in which case an account that signs up during it is held until an administrator lets it in, and until then the tools refuse and the account page shows what to expect. Nothing to reconnect and nothing to reconfigure — the refusal names where to look, and the same credential starts working when the account is let in.

  • Reading anything fails, listing everything works. check_inbox, list_threads and whoami answer, and get_package, get_thread or draft_package come back with a generic "an error occurred invoking" for every id you try. The split is not between reading and listing — it is between tools that require an argument and tools that do not. The client sent a name the tool does not declare, the argument went missing, and a tool with nothing required never noticed. RelayLink now accepts either spelling of a name (package_id and packageId both bind) and answers anything it still cannot match by naming the argument and listing what the tool takes, so a bare "an error occurred" from this server means something else. If you are seeing one from an older deployment, ask your assistant to call the tool again with the exact names from the schema.

  • You said "send it" and nothing arrived. A draft is not a send. draft_package returns a draft id and the words "NOT sent yet." Check list_drafts, then confirm. Pending drafts expire after 24 hours; confirming an expired one cancels it and asks for a fresh draft.

  • check_inbox is empty and you know something arrived. It returns unread packages by default, and a package counts as read once it has been pulled, opened on the web, or replied to. Ask for read items too. It also shows at most 20 and reports the true total.

  • The transcript contains ERROR: text. Domain refusals come back as real errors with actionable wording — an invalid response shape, a consent refusal, a malformed id. That is the server working.

  • A 429 from /mcp itself. It has its own ceiling, separate from the per-IP limit on the public token pages and the health endpoint: 300 calls a minute, partitioned by the presented credential rather than by address. It exists to catch a runaway tool loop or a stolen key, not to throttle ordinary use, so seeing one means checking for a loop before assuming the server is at fault.

If you are on the other side of this — writing a server rather than connecting to one — the same ladder in reverse is how to build an MCP server that's safe to connect. And if you got here because it now works, give it something real to carry.

Frequently asked questions

Why does my MCP server connect but show no tools?
Usually one of two things. The client is holding a tool list it fetched in an earlier session, so reconnect the server and start a fresh conversation. Or the client was configured for a local server and handed a remote URL, which tends to fail in ways that never name the real problem.
What does a 401 from an MCP endpoint actually tell me?
More than it feels like. A 401 means you reached the right host and the right path and the server's authentication ran and rejected you. A wrong URL usually produces a connection error or a 404 instead. From there the fault is the header name or the header value, including whitespace picked up from copy and paste.
Is a draft the same as a sent message?
On RelayLink, no. Drafting and sending are separate tool calls, and a draft sits on the server unsent until it is explicitly confirmed from the same account. Pending drafts expire after 24 hours. If you expected something to arrive and it did not, list the drafts before assuming the connection broke.