Our crawler
TODO runs a small, polite crawler that checks official sources for news. This page explains what it does and how to stop it.
User agent
Our crawler identifies itself with this user agent string:
NewsroomBot/1.0 (+https://gentwatch.news/bot)
What it fetches
- Feeds and official APIs, such as RSS and Atom feeds of vendor blogs and changelogs, public research paper feeds, and the GitHub REST API. These are checked at most every two hours, most less often.
- Article pages, rarely. For a small allowlist of primary sources, it may fetch the page of a specific announcement to read its text for fact-checking. Full text is used only while writing and checking; it is not stored or republished.
- Never images. We never download or reuse images from other sites.
How it behaves
- It obeys robots.txt, including
Crawl-delay, and caches robots.txt between visits. - It honours text and data mining reservations (the TDMRep protocol and
noaisignals) by not reading the full text of pages that opt out. - It makes one request at a time per site and uses conditional requests (ETag / Last-Modified), so unchanged feeds cost you almost nothing.
- It never bypasses paywalls, logins or CAPTCHAs.
How to opt out
To block our crawler, add this to your robots.txt:
User-agent: NewsroomBot
Disallow: /
Changes are picked up within a day. You can also email TODO(human) and we will remove your site from our source list.