How to block AI crawlers on your Substack
Published posts and notes, scraped by third-party crawlers.
Per publication, not per account, and only as strong as a crawler's manners.
Published posts and notes, scraped by third-party crawlers.
Per publication, not per account, and only as strong as a crawler's manners.
Substack itself is not the main risk. Published posts and notes are public web pages, and third-party crawlers ingest them.
`<your-publication>.substack.com/publish/settings` and searching that page for "AI" jumps straight to it. account-level page at substack.com/settings does not contain this setting at all.
Substack's own wording: "Substack never uses your content to train AI. This setting lets you instruct third-party AI tools, like ChatGPT, Claude, and Google Gemini, to not scrape and use your content for their model training."
A crawler block is a request in robots.txt. Compliant crawlers honour it. Non-compliant ones ignore it, and anything already crawled stays crawled.
Cross-posts and syndicated copies on other domains are not covered by this setting.
Per publication, not per account. The setting is called "AI training protection", under Settings then Privacy on each publication you own. Blocks compliant crawlers only.
Every service on that list buries its opt-out somewhere different. Don't Train On Me puts all of them in one checklist in a panel beside your browser, walks you through them one at a time, and helps you confirm your privacy settings on the ones that expose a setting it can read. Free, no account, and nothing you tick leaves your machine.