MemoryCVBot
The job-posting crawler operated by MemoryCV.
Purpose
MemoryCVBot builds an index of employer job postings. It is shown to signed-in MemoryCV users. It may also be licensed to third-party developers, but only for sources with a separate resale decision reviewed by counsel; EU sources are excluded from resale for now.
Identity
- Product token:
MemoryCVBot -
User-Agent:
MemoryCVBot/1.0 (+https://crawler.memorycv.com; crawler@memorycv.com) From: crawler@memorycv.comon every request- Contact: crawler@memorycv.com
- Egress IPs: not fixed yet.
What we fetch
-
Public job-board APIs of registered employers. First planned:
boards-api.greenhouse.io,api.lever.co, Lever's documented EU hostapi.eu.lever.co, andapi.ashbyhq.com. The list grows as adapters ship. - To find boards: a bounded number of pages on approved employer websites, their careers pages, and documented sitemaps.
Limits and robots.txt
-
At most 1 request per second per host, or your
Crawl-delay, whichever is slower. - robots.txt follows RFC 9309: a 4xx response means allowed; a 5xx or unreachable response means we fetch nothing from that host until it is readable again. We cache it for at most 24 hours.
To block us from your whole host, add:
User-agent: MemoryCVBot
Disallow: /
Personal data
Recruiter and other people's names, emails and phone numbers are stripped from developer data.
Opt-out and takedown
- robots.txt stops fetching within 24 hours. It does not remove data already published, and removing the rule makes your host eligible again.
- Email crawler@memorycv.com for a lasting opt-out or removal. We record a durable opt-out for your domain, board or host and keep your request as evidence. An employer opt-out covers all of that employer's job boards, on any platform, including ones we find later. Scans in progress stop as incomplete, no new scan starts, and we acknowledge by email. Nothing re-adds you; only your written request lifts the opt-out.
- Our target for withdrawal from both the MemoryCV and developer feeds is 2 business days.