About AluttaBot
What Alutta’s web crawler is, what it reads, and how to control it on your website.
Who runs it
AluttaBot is run by Alutta, a platform that helps students apply to universities abroad. It reads university websites for one feature, Supervisor Finder, which helps students applying for research degrees find researchers whose work matches their own.
Researchers can read what Supervisor Finder shows about them, and ask to be left out, on our page for researchers.
How to recognise it
Every request AluttaBot makes carries this user agent:
AluttaBot/1.0 (+https://alutta.com/bot)We do not publish a list of IP addresses. If something calling itself AluttaBot behaves differently from what this page describes, it is not us, or something is wrong. Please tell us either way.
What it reads
Only public pages on university websites: staff directories, department pages and researchers’ profile pages. It starts from the university’s own web address and follows links only within that university’s domain.
It reads web pages and nothing else. It never logs in, submits a form, runs a page’s scripts or downloads files.
How hard it pushes
No more than one page a second from any website, and no more than 60 pages in a single visit to a university. If your robots.txt sets a Crawl-delay for AluttaBot, it waits at least that long between pages, up to a minute.
A university is visited when a student starts looking for researchers there. What AluttaBot reads is reused for other students rather than fetched again, so repeat visits are infrequent.
Limiting or blocking it
AluttaBot follows robots.txt. To keep it off a whole website, add:
User-agent: AluttaBot
Disallow: /To keep it out of one part of a site, disallow only that path, for example:
User-agent: AluttaBot
Disallow: /staff/private/To slow it down rather than keep it out, set a delay in seconds between pages:
User-agent: AluttaBot
Crawl-delay: 10It checks robots.txt again at least once a day, so a change applies within 24 hours. To stop it visiting your site altogether without editing robots.txt, use the form below.
What we keep
From each page we keep only what Supervisor Finder shows: a researcher’s name, title and department, a link to their profile page, and a work email address where the page lists one, together with the address of the page it came from.
We do not keep copies of your pages, and we never guess or construct an email address.
Stop it visiting your site
For your university’s web team. We email a link to an address at your website, and AluttaBot stops once you confirm. Nothing changes until then.
Researchers can still appear in Supervisor Finder from their public publication record. Each can ask to be left out on our page for researchers.
Contact
Write to hello@alutta.com with “AluttaBot” in the subject. Tell us the website, roughly when, and what you saw. We read every message.