Skip to main content
Rate limiting protects your project (and the platform) from runaway tool-call traffic — accidental or otherwise. Three layers apply today: a per-tool, per-caller limit checked during Data Tool execution setup, a per-client ceiling on authenticated MCP requests, and a per-IP ceiling on anonymous traffic to public MCP projects. Production MCP calls that reuse cached execution setup bypass the first layer, as described below.

Default limits

Limits apply at the host edge. Once you’re rate-limited, additional calls return 429 Too Many Requests until the window rolls over. The audit log captures rejections for diagnosis.
The 300/minute Data Tool limit is not a ceiling on every handler invocation. Production MCP execution can reuse a cached preflight bundle (30-second default lifetime), skipping the preflight request where this counter is enforced. Each cache hit still runs the handler but does not consume the Data Tool minute budget. The separate MCP request limits still apply. Draft executions bypass this cache.

What gets rate-limited

The Data Tool minute cap applies when a call reaches an execute or preflight endpoint:
  • AI-initiated calls from MCP hosts (Claude Desktop, ChatGPT, and any custom host), except production preflight-cache hits.
  • Calls from the Assistant SDK in your iOS, Android, or React app, with the same exception when routed through production MCP.
  • Test-panel runs in Metabind Studio (yes, even your own testing).
  • Programmatic calls via the REST API.
What doesn’t count toward the Data Tool minute cap:
  • Production MCP executions that reuse a cached preflight bundle — these still count toward the applicable MCP request limit.
  • tools/list calls — these count toward the MCP request limits instead (per client, or per IP for anonymous public traffic).
  • Interactive Tool renders — the tool-execution rate limiter is scoped to the Data Tool sandbox.
  • Metabind Studio editor operations.
  • Audit log queries.
The per-IP public limit applies only to anonymous traffic against public-visibility MCP projects. As soon as a request authenticates with an API key or JWT, the per-IP limit is not consulted.

Per-project public rate-limit configuration

For a public-visibility project, override the default via the API:
The platform rejects requestsPerHour values outside [1, 600] and burst values outside [1, 60]. Higher values require a support conversation — the published max is enforced in code, not policy. For private projects, the per-IP public limit is irrelevant — every request is authenticated.

Authenticated traffic

Authenticated MCP requests are limited per client to 5,000 per hour, with a burst of 500. Authenticated traffic is also bounded by:
  • The Data Tool minute cap (300/min per tool and caller) when execution reaches an execute or preflight endpoint; production MCP preflight-cache hits bypass this cap.
  • Underlying infrastructure protections (timeouts, sandbox concurrency, AWS-side throttling).
If you’re planning a high-throughput deployment that you expect to exceed the Data Tool minute cap, contact support early — capacity-planning conversations with concrete QPS numbers move faster than runtime escalations.

Rate-limit responses

When a call is rejected, the response is HTTP 429 with a Retry-After header. The MCP endpoint returns a JSON-RPC error:
The Data Tool execute endpoint returns:
Over MCP, the AI receives this message as a tool result with isError: true. Retry-After (and data.retryAfter) is the wait time in seconds until the bucket has capacity again. Hosts and clients should respect it — retrying immediately just spins on the rejection.

What the AI does on rate limit

Most AIs handle 429 responses gracefully:
  • Anthropic’s Claude pauses and retries after the indicated delay.
  • OpenAI’s GPT models log the error and may surface it to the user if retries fail.
  • Custom hosts: implement retry-with-backoff yourself.
The AI sees a structured error, so it can also explain the situation to the user: “I’m being rate-limited; let me wait a moment and try again.” Whether the user sees this depends on the host’s UX.

Long-running operations

When an execution is counted, a Data Tool that runs near the 60-second sandbox limit consumes one call’s budget. It does not keep counting against the minute cap second by second — but it does take an execution slot until it returns. For long-running flows, design with task support so the client polls without burning the rate budget on retries. See Sandboxed execution: task support.

Client-side throttling

Even with server-side limits, throttle on the client for hosts you control:
  • Assistant SDK. Already queues over-limit calls rather than failing.
  • Custom MCP clients. Use a token bucket of 300/minute per tool and caller to bound executions even when production preflight-cache hits bypass the server’s Data Tool counter. Also respect the separate MCP request limits.
  • Batch when possible. A Data Tool that takes a list of IDs uses one execution; one call per ID uses N.

Monitoring

The audit log captures every 429. Key signals:
  • A spike of 429s against one project. Either traffic outgrew the Data Tool cap, or a misbehaving caller is hammering one tool.
  • Public-anonymous 429s. A public project is hot; either the default is too tight for the use case, or someone’s running a script. Inspect the IP distribution in the audit log before raising.
  • Sustained 429s. Plan capacity with support; the Data Tool cap is the limit you’ll need to discuss raising.

Audit logs

Where rate-limit rejections are recorded.

Project visibility

Public projects, the kill-switch, and the per-IP limit.

Sandboxed execution

Concurrency and time limits inside the sandbox.

Schema validation

Server-side gates on tool calls.