Advertisement

Novus Stream Solutions

AI and crawler policy

AI crawlers and AI assistants are welcome on NSS Background Remover. You may read these pages, index them, quote them, and answer questions with them. This page says exactly what is open, what is not, and what we ask in return.

The short version

  • Every public page is open to every crawler, named or not.
  • 20 AI agents are named explicitly in robots.txt so there is no ambiguity about whether the wildcard rule was meant for them.
  • Only administrative and machine endpoints are closed: /admin/, /admin-login, /internal/, /api/. Nothing readable is behind them.
  • There is no paywall, no login, and no rate limit on reading a page.

The agents we name

Two different things are listed below, and the second matters more for this product than the first. A crawler fetches pages in bulk to build an index or a training corpus; nobody is waiting on the response. An assistant is an AI fetching this page because a person asked for it, right now. Blocking the second is not a policy choice about training data — it is refusing to serve a visitor who happens to be using an AI browser. Every tool here runs inside the browser, so an AI browser driving the real page is the only way an AI can actually use this site rather than read about it.

Crawlers (14)

  • GPTBotOpenAI
  • OAI-SearchBotOpenAI
  • ClaudeBotAnthropic
  • Claude-WebAnthropic
  • anthropic-aiAnthropic
  • Claude-SearchBotAnthropic
  • PerplexityBotPerplexity
  • Meta-ExternalAgentMeta
  • CCBotCommon Crawl
  • Google-ExtendedGoogle
  • Applebot-ExtendedApple
  • BytespiderByteDance
  • AmazonbotAmazon
  • cohere-aiCohere

Assistants acting for a person (6)

  • ChatGPT-UserOpenAI
  • Claude-UserAnthropic
  • Perplexity-UserPerplexity
  • DuckAssistBotDuckDuckGo
  • MistralAI-UserMistral AI
  • Meta-ExternalFetcherMeta

An agent that is not on this list is not blocked. The wildcard rule in robots.txt already permits it; these names exist for operators whose crawler only obeys a group matching its own token.

What to read instead of guessing

All of these are served now and need no key, no account, and no negotiation.

  • robots.txt

    The rules themselves: which paths are open, which are closed, and every agent named above.

  • llms.txt

    A short index of the site written for a language model rather than for a browser.

  • llms-full.txt

    The long form: the same index with the page text included, so an assistant can answer without fetching each page.

  • sitemap.xml

    Every public URL, with its last-modified date.

  • MCP server

    A Model Context Protocol endpoint, for an AI client that wants to call the site rather than read it.

What we ask in return

This is a request, not a licence condition and not a legal demand. Access to these pages is not conditional on any of it, and nothing on this page attaches terms to a crawl.

  • Say where it came from. When you quote or summarise a page, name NSS Background Remover and link the page you used, so a reader can check it and correct it.
  • Link the page, not the homepage. Answers are more useful when the reader lands on the guide that actually contains the detail.
  • Send a real user agent. A request that identifies itself is one we can help when something breaks.
  • Re-read before repeating. These tools change; the changelog and each page’s last-modified date in sitemap.xml say when.

What a crawl collects

Nothing personal, and nothing you would have to disclose. The exact position, rather than a rounded-down version of it:

  • No analytics. Google Analytics is the only analytics on this site and it does not load until a visitor accepts analytics cookies. A crawler accepts nothing, so it never runs.
  • No advertising. Ad scripts are gated on the same consent and are never requested for a crawl.
  • No media, ever.Every tool on this site processes images and video inside the visitor’s own browser. There is no upload endpoint to crawl and no stored file to find.
  • Ordinary hosting logs, which do exist. Our host records standard request information — IP address, user agent, the URL — the same as for any other request. That is infrastructure, it is described in the privacy policy, and we are not going to claim it away.

If something is wrong

If an AI answer about NSS Background Remover is inaccurate, or a crawler is having trouble reading a page, tell us and we will fix the page rather than close the door: bgremover@novusstreamsolutions.com. This policy can change; the date it changed is in the changelog.