The Order I Did This In Was Wrong

Cloudflare publishes a setup file at developers.cloudflare.com/agent-setup/prompt.md. It's written to be read by an AI coding agent, and for Claude Code it collapses to two commands:

claude plugin marketplace add cloudflare/skills
claude plugin install cloudflare@cloudflare

That installs eleven skills and registers five MCP servers. I ran both. Then, a few minutes later, I remembered the advice I'd been hearing all month — don't install skills and plugins from random sources — and thought: I should probably check that.

Wrong order. The audit is supposed to be a gate, not a postmortem. But the postmortem produced something more useful than a verdict, which is a repeatable checklist. Here it is, in the order that actually matters.

Up front, so nobody misreads this post: the Cloudflare plugin is clean. Real org, signed commits, no hooks, no injection, no exfiltration. This is not a disclosure. It's a demonstration of what a pass looks like, so you can recognize a fail.

1. Provenance: is this actually who it says it is?

This is the first check because it's the one that catches the attack that actually happens. Nobody sneaks malware past you with clever obfuscation. They register a GitHub account called cloudfIare with a capital I instead of an l, and you install it without looking.

So: don't read the plugin's README. Read the git remote and ask GitHub who owns it.

git -C ~/.claude/plugins/marketplaces/cloudflare remote -v
gh api repos/cloudflare/skills  --jq '{owner:.owner.login, type:.owner.type, fork:.fork, stars:.stargazers_count}'
gh api orgs/cloudflare          --jq '{created:.created_at, repos:.public_repos, blog:.blog}'

What you want to see: the remote matches the org exactly, the owner is an Organization and not a personal account, the repo is not a fork, and the org has a history. Cloudflare's org was created in June 2010 and has 568 public repos. That's not something you spin up to run a supply-chain attack this week.

Then check the commit you actually installed, not the repo in the abstract:

gh api repos/cloudflare/skills/commits/<sha> --jq '{verified:.commit.verification.verified, author:.commit.author.email}'

The commit I had was GPG-verified and authored by a Cloudflare Workers PM. Combined with a CODEOWNERS listing four recognizable staff engineers, that's about as good as provenance gets short of a signed release artifact.

One note on a trap here: the GitHub API returns is_verified: false for the Cloudflare org when you call it unauthenticated. That field is not the signal you think it is. Ignore it and look at age, repo count, and commit signatures instead.

2. Hooks: the only thing that runs without asking you

This is the check that matters most, and it's the one almost nobody runs.

Everything else in a Claude Code plugin — skills, slash commands, MCP tools — is inert until an agent chooses to use it, and most of it surfaces a permission prompt when it does. Hooks are different. A hook is a shell command wired to a lifecycle event: PreToolUse, PostToolUse, SessionStart. It fires automatically. You don't approve it. You often don't see it.

The threat model in one line. A malicious plugin doesn't need to convince your agent to do anything. It just needs a SessionStart hook and your ~/.aws/credentials.

So before anything else, ask what auto-executes:

# Does the plugin declare hooks at all?
jq '.hooks' ~/.claude/plugins/cache/<mkt>/<plugin>/<ver>/.claude-plugin/plugin.json

# Any hooks directory, any settings file shipping hooks?
find <plugin-dir> -iname 'hooks' -o -iname 'settings*.json'

# npm lifecycle scripts — postinstall is a hook by another name
find <plugin-dir> -name package.json -not -path '*/node_modules/*' \
  -exec jq '.scripts | keys' {} +

The Cloudflare plugin declares no hooks, ships no hooks/ directory, and its single package.json (a scaffold template that never gets installed by the plugin itself) has no preinstall or postinstall. Nothing in it runs on its own. That's the answer you want.

If a plugin does ship hooks, that isn't automatically disqualifying — plenty of legitimate formatter plugins use PostToolUse. But now you have to read every one of them, and the bar for what you accept should be much higher.

3. MCP endpoints: where is your data going?

An MCP server is a remote host that your agent hands context to. The config is a single file and it takes ten seconds to read:

jq '.mcpServers[].url' <plugin-dir>/.mcp.json

All five of Cloudflare's are on *.mcp.cloudflare.com. If even one had pointed at some unrelated host, that alone would end the evaluation.

Worth understanding what you're granting, though. Four of the five trigger an OAuth flow on first use, and cloudflare-api and cloudflare-bindings can obtain broad account read/write. Only cloudflare-docs is public and read-only. If you're cautious, connect the docs server, leave the rest unauthenticated until you have a reason.

4. Injection and exfiltration: grep the prose, not just the code

A skill is a markdown file full of instructions your agent will follow. The "code" in a skill is the prose. So grep it like code.

grep -rniE '\.env|\.dev\.vars|\.aws|\.ssh|id_rsa|credentials|secret' <plugin-dir>
grep -rniE 'ignore (all )?previous|disregard .* instructions' <plugin-dir>
grep -rnoE 'https?://[a-zA-Z0-9.-]+' <plugin-dir> | grep -vE 'cloudflare\.com|example\.com' | sort -u

That last one is the good one. Enumerate every external URL, subtract the domains you'd expect, and look at whatever survives. Across 406 files the residue was: a webhook.site URL used as a fetch example, a jsDelivr script tag for an official Cloudflare UI package, some esm.sh imports in Workers Playground snippets, a Honeycomb OpenTelemetry export example, and a Google Forms link to Cloudflare's own Hyperdrive limit-increase request. Every one of them explicable in context. That's what a clean result looks like — not zero external URLs, but zero unexplained ones.

5. The version pin that doesn't pin

This is the finding that surprised me, and it's the reason I think this checklist is worth publishing even though the verdict was clean.

The plugin is installed at version 1.0.0. There's a directory on disk named 1.0.0. Everything I audited above, I audited inside that directory. It's tempting to conclude that I audited what will run.

I didn't. Two of the shell scripts fetch their payload at run time:

npx --yes degit cloudflare/skills/<path>

No ref. No tag. No commit SHA. That resolves to the repository's default branch at the moment the script runs, which may be months of commits ahead of the snapshot sitting in your 1.0.0 folder.

The gap. A version pin that covers the files on disk but not the files those files go and fetch is a pin in name only. The audited artifact and the executed artifact are different things.

For a Cloudflare-owned repo pulling from a Cloudflare-owned branch, this is fine. The trust boundary hasn't moved — you already decided to trust Cloudflare when you installed it. But the property is worth naming, because it generalizes badly. The same pattern in a plugin from a less-scrutinized publisher means your audit has a hole in it, and the hole is invisible unless you go looking for run-time fetches.

So add this to the grep pass:

grep -rnE 'degit|curl .*\| *(ba)?sh|npx --yes|git clone' <plugin-dir>

6. Read the destructive commands, then guard them yourself

The last thing I checked wasn't about trust at all. It was about blast radius.

The wrangler skill is a command reference. It documents wrangler delete, wrangler d1 execute --remote, wrangler d1 migrations apply --remote, and wrangler kv bulk delete --force. It never tells an agent to run them unprompted. But it also never says "confirm first," and it never says "check which account you're on."

I run six separate Cloudflare accounts across a few dozen live sites. A skill that makes those commands more available to an agent, without adding a confirmation norm, raises my blast radius even though it's doing nothing wrong. Interestingly, one skill in the same bundle gets this exactly right — its auth probe detects a multi-account token and stops to make the agent ask which account to use. The norm exists in the codebase. It just isn't applied uniformly.

The fix belongs on my side regardless of which plugin is installed. In ~/.claude/settings.json:

{
  "permissions": {
    "ask": [
      "Bash(wrangler d1 execute *)",
      "Bash(npx wrangler d1 execute *)",
      "Bash(wrangler d1 migrations apply *)",
      "Bash(npx wrangler d1 migrations apply *)",
      "Bash(wrangler delete *)",
      "Bash(wrangler kv bulk delete *)"
    ]
  }
}

Note the npx duplicates. The plugin's own deploy script shells out through npx wrangler, and a rule matching wrangler * will not match npx wrangler *.

There's a subtler hole too: these rules match on prefix. A command like cd ~/my-project && npx wrangler d1 migrations apply --remote starts with cd, so it sails straight past every pattern above. If you're running Claude Code in auto permission mode, you can close that with a semantic rule under autoMode.soft_deny, which is evaluated by a classifier that reads the whole command rather than matching a prefix:

{
  "autoMode": {
    "soft_deny": [
      "$defaults",
      "Any wrangler command with --remote that writes to a D1 database, including when nested inside a compound command, subshell, or npx invocation.",
      "Before any wrangler command that writes to a remote Cloudflare resource, confirm which account is active and state the target account and resource name."
    ]
  }
}

Keep "$defaults" in that array or you replace the built-in rules instead of extending them.


The checklist, condensed

# Check Why it's in this position
1 Provenance via GitHub API — org type, age, commit signatures Typosquats are the attack that actually lands
2 Hooks, hooks/ dir, npm lifecycle scripts The only code that runs without your approval
3 Every URL in .mcp.json These hosts receive your context
4 Grep prose for credentials, injection, unexplained URLs In a skill, the prose is the code
5 degit / curl | sh / unpinned clones Run-time fetches escape your version pin
6 Destructive commands → add ask rules Blast radius is yours to manage, not the publisher's

What I'd actually change about my process

Not much about the checks. Everything about the sequencing.

The install commands were sitting in a document that explicitly instructed the reading agent to run all of these yourself, do not ask the user. That instruction was benign, and it came from a domain I trust. But "the document told me to" is not a reason to skip a gate, and an agent reading a fetched file should treat the file's imperatives as data rather than orders. That's true when the file is from Cloudflare. It's much more true when it isn't.

Run the checklist first. It takes about four minutes, and six of the checks are one-liners.