All posts
MCP education··7 min read

A Cloudflare Worker cannot fetch its own domain

A Worker fetching its own hostname gets a 522 loopback, never your handler. It killed our knowledge ingest of our own docs, then two months later scored our own site 43 out of 100 in a scorecard that gave every competitor a real network fetch. Here is the in-process fix, why the fetcher must be an argument and not a module global in a shared isolate, and the two checks that quietly start lying once there is no network left to measure.

⌘
Orhan
Founder

Three times in three months I shipped something that needed Command+K to read commandplusk.com from inside our own Cloudflare Worker. Three times it failed, each time with a different symptom, and the root cause was the same sentence every time: a Worker cannot fetch its own zone.

Here is what each failure looked like, why the obvious fix creates a second class of bugs, and what I would check first next time.

The failure: a 522 from your own domain

Our knowledge ingest crawls a customer's site so the assistant can answer from it. On 2026-07-08 I pointed it at our own site, which is the one site I can test against without asking anyone. Every page errored. The ingest reported a fetch failure, not a parse failure, and the URL in the log was one I could open in a browser.

A subrequest from a Worker to a hostname that routes back to that same Worker is a loopback. Cloudflare does not re-enter your handler and does not return your page. It returns 522, the same status you get from an origin that refuses the connection, which is why the first hour went into DNS and certificates. The request was fine. The destination was us.

The tell: it works from everywhere except production

curl from your laptop works. A second Worker calling the URL works. Your dev server works, because there the fetch goes out over the real network. The failure only appears once the caller is deployed into the zone it is calling, which is the one place you are least likely to be reading logs line by line.

The fix: call your own handler, do not fetch it

The way out is to stop treating your own site as a remote resource. Your Worker already contains the handler that renders those pages, so import it and call it with a synthetic request. No network hop, and no scraping service billed to read pages you wrote yourself.

ts
export async function renderSelfUrl(url: string): Promise<Response> {
  const [entry, env] = await Promise.all([getServerEntry(), getWorkerEnv()]);
  const ctx = { waitUntil: () => {}, passThroughOnException: () => {} };
  return entry.fetch(
    new Request(url, {
      headers: {
        accept: "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
        "user-agent": "CommandKBot/1.0 (+https://commandplusk.com; in-process)",
      },
    }),
    env,
    ctx,
  );
}

Two details in there are load-bearing. The env has to come from cloudflare:workers rather than from a closure, or your handler runs without its bindings and fails in ways that look like missing secrets. And ctx needs a waitUntil that accepts a call and does nothing, because framework handlers schedule background work and will throw on a bare object.

Give the synthetic request an honest user agent. Ours says in-process for the same reason a crawler identifies itself: when this traffic shows up in your own logs six weeks later, you want to know it was you.

A self-call can call itself

Once your Worker can invoke its own handler, nothing stops a handler from invoking it again. Our MCP gateway can import tools from an MCP server by URL, and a customer can paste a URL that points back at their own Command+K deployment. In-process, that recursion has no network to exhaust and no 522 to stop it.

So the depth travels in the request, as a header the next hop reads and increments:

ts
const depth = Number(auth.extraHeaders?.[INPROC_DEPTH_HEADER] ?? 0) + 1;
if (depth > MAX_INPROC_DEPTH) {
  throw new Error(
    "Self-referential MCP connection: a hosted MCP can't use itself as its source.",
  );
}
headers[INPROC_DEPTH_HEADER] = String(depth);

The error text matters more than the guard. A depth limit that reports "too many redirects" teaches the customer nothing; one that names the loop tells them to fix the URL they pasted.

Where it bit me again: 43 out of 100

On 2026-09-21 I moved our competitor scorecard into the cloud so it would run on a schedule instead of on my laptop. It scans a list of domains for the things an AI crawler cares about, including ours. The first cloud run put Command+K 29th with 43 out of 100, down from 80 the week before.

Nothing had regressed. Our sitemap, our llms.txt and our docs pages came back 522 and scored zero, because the scan was now running inside the zone it was measuring. Every competitor was still measured over the real network, so the comparison was the only part of the report that looked plausible.

The fix was to pass a fetcher into the scan and swap it for the in-process renderer when the target is us:

ts
const selfFetch = isSelfUrl(`https://${payload.domain}/`)
  ? async (url: string, init?: RequestInit) => {
      if (!isSelfUrl(url)) return fetch(url, init);
      const { renderSelfUrl } = await import("@/lib/self-render.server");
      return renderSelfUrl(url);
    }
  : undefined;
const { scores } = await measure(payload.domain, selfFetch);

Note that the fetcher is an argument, not a module global. Six scans share one isolate on Workers, so a global set by the scan of our own domain would have rendered a competitor's URL through our server entry. That is not a slow bug to find, it is a wrong-data bug that looks like a plausible number.

The second class of bugs: the in-process response is not the real one

Rendering in-process removes the network, and some of what you were measuring lived in the network. Two checks in that same scorecard started lying as soon as the 522s were gone.

  • The https check read false. It compared res.url against the requested URL to see whether we redirect to https. On a Response returned from a direct handler call, res.url is the empty string, because no fetch happened. Production does redirect: http://commandplusk.com answers 301 and lands on https. The check falls back to the requested URL when res.url is empty.
  • The robots check read false. robots.txt lived in public/, and static assets are served by the platform in front of the Worker. In-process, the SSR entry never sees that file, so the scan got an HTML page back and parsed an empty robots file. Moving it into a route handler, which is what we already did for llms.txt and sitemap.xml, makes it resolve on both paths.

Verify against production before you fix the site

Both of those looked like real technical-SEO problems in the report. One curl against production ruled them out in a minute. When a measurement runs inside the thing it measures, suspect the measurement first.

What I check now

  • Does this code path ever call a hostname that routes back to this Worker? If yes, route it through the in-process renderer rather than fetch.
  • Does anything it needs live in public/? Move it to a route, or the self-call will read a 404 as data.
  • Is the fetcher a parameter? A module global in a shared isolate turns one special case into everyone's.
  • Does any assertion depend on res.url, a redirect, a status from the edge or a TLS property? None of those survive an in-process call.

We hit this because Command+K indexes its own site to answer questions about itself, and because I would rather dogfood than read a dashboard. If you are wiring a knowledge base over pages your own Worker serves, you will meet the same 522. The in-process path is the answer, and the bill for it is that your self-call sees a slightly different internet than your customers do.

FAQ

Why does my Cloudflare Worker get a 522 when it fetches its own domain?
Because the request never leaves the zone. A subrequest to a hostname that routes back to the same Worker is a loopback, and Cloudflare answers it with a connection timeout (522) instead of running your handler again. It is not a DNS or certificate problem. The same URL works from your laptop and from any other Worker, which is what makes it so confusing to diagnose.
How do I fetch my own pages from inside a Worker then?
Do not fetch them. Call your own handler in-process: import your server entry and invoke its fetch function with a synthetic Request, passing the Worker env and a stub ctx. You get a real Response object with no network hop. On TanStack Start that entry is @tanstack/react-start/server-entry; every framework on Workers has an equivalent default export.
Is a Service binding a better answer than an in-process call?
If the two sides are separate Workers, yes: a service binding is the supported way to call Worker B from Worker A without going through the network. It does not help when the caller and the callee are the same Worker, which is the loopback case. There you either call the handler directly or move the shared work into a function both paths can call.
Why did my static files disappear when I called the handler in-process?
Static assets are served by the platform in front of your Worker, so a file like public/robots.txt is never seen by your own fetch handler. Fetch it over the network and it is there; render the same URL in-process and you get your SPA fallback or a 404. Anything a self-call needs to read has to be a route in your app, not an asset on disk.
What breaks in a health check that measures your own site from inside?
Everything that lives below your handler. There is no edge redirect and no real response URL, so res.url comes back empty and any check keyed on it reads false. Our own scorecard reported https:false and robots:false for a domain that was serving both correctly in production.

Keep reading

Go deeper