Guides
Collecting Data You Will Not Be Allowed to Use
Terms of Service Decide What You Can Collect
Plenty of sites can be scraped. The question worth asking is not whether it is possible but whether it is permitted. Most sites carry a clause in their terms restricting or forbidding automated collection, and many publish a robots.txt marking which paths are open to crawlers. Reading the terms and the robots.txt of every target site comes before commissioning anyone. Skip it and you can finish the entire collection and still have nothing you are able to use.
robots.txt Is Not Law and Still Gets Cited
robots.txt is not a binding legal instrument. It is a request from the site operator to automated clients, and a crawler that ignores it will still come back with data. The cost arrives later. Having ignored a published robots.txt is the kind of fact that can be read as intent once a dispute starts. Stay off the paths a site has closed, and where the data is genuinely necessary, write to the operator and ask in writing.
Personal Data Changes the Standard
A post can be public and still carry a name, a contact detail, a profile photo. Once that is in the set, data protection law applies to it, however public the post looked. “It was already public” is not a defense to build a project on. Collecting personal data, storing it, and using it each need their own lawful basis. If personal data is mixed into the target, that part gets reviewed separately or it does not get collected.
Copyright Bites at Reuse, Not Collection
News articles, photographs, full review text: fetching them is the quiet half of the job. Publishing them on your own product is where the exposure sits. Summarizing with a link back to the source and reproducing the original in full sit in completely different places legally. Purpose shifts the answer as well. A set that stays inside the company for analysis and a set that gets published on your own product cannot be judged against the same standard.
Answer These Three Before You Commission It
Have the answers ready and the work moves faster.
- Have you read the target site’s terms and its robots.txt
- Does the collection include personal data, and if so, how will it be handled
- Will the result stay internal, or be shown to the public
With those settled, whoever takes the job can draw a tight scope, and the dataset you end up with is one you are allowed to open.
Start With the Target List
Weighing legal limits while working out what to collect is a lot to carry the first time through. At Weple, that boundary work is step one of a data collection project, settled before anyone writes a fetcher. Send the sites you want covered and what the data is for, and the first thing back is a straight answer on how much of it can actually be done.
In a similar spot
Send us where things stand and we reply with the scope and the price within 24 hours.