
Experienced programmer required to develop code to scrape a website
5170
$$
- Posted:
- Proposals: 1
- Remote
- #137236
- Archived
Description
Experience Level: Intermediate
I have been given your contact details by my colleague David Roberts Jones after you responded to the request for job 134859, which David posted on my behalf.
Basically I'm looking to scrape the UK Charity Commission website to download the 162,000 public domain records. I'm a charity fundraiser and I want to make this information more freely available, but the Charity Commission have just lost 40% of their annual budget and, whilst they don't mind me doing this (their exact words were 'we have no objection'), they are not going to be very helpful.
I'm looking to collect all 162,000 records, but am also looking for a piece of code so that I can collect all the new records that get added each month (around 300 new charities registering).
The charities are identified through their unique identifier, the Charity Commission Number, which for my own charity is 1011102. The number is sequential, so 1146997 will bring up the charity that registered earlier today.
The most straightforward place to start is the website:
http://www.charity-commission.gov.uk/
and if you put in 1146997 into the search box you get 'River Legacy', I'm looking to collect all the 'charity framework' information and all the 'charity contact' information for the monthly update, and all the information for the full down load.
What do you think? It has been tried before, by http://opencharities.org/ but they didn't get it to work very well.
I'm looking to turn this data into a database, so xml or csv would suit well.
Basically I'm looking to scrape the UK Charity Commission website to download the 162,000 public domain records. I'm a charity fundraiser and I want to make this information more freely available, but the Charity Commission have just lost 40% of their annual budget and, whilst they don't mind me doing this (their exact words were 'we have no objection'), they are not going to be very helpful.
I'm looking to collect all 162,000 records, but am also looking for a piece of code so that I can collect all the new records that get added each month (around 300 new charities registering).
The charities are identified through their unique identifier, the Charity Commission Number, which for my own charity is 1011102. The number is sequential, so 1146997 will bring up the charity that registered earlier today.
The most straightforward place to start is the website:
http://www.charity-commission.gov.uk/
and if you put in 1146997 into the search box you get 'River Legacy', I'm looking to collect all the 'charity framework' information and all the 'charity contact' information for the monthly update, and all the information for the full down load.
What do you think? It has been tried before, by http://opencharities.org/ but they didn't get it to work very well.
I'm looking to turn this data into a database, so xml or csv would suit well.
Stuart S.
100% (2)Projects Completed
2
Freelancers worked with
2
Projects awarded
100%
Last project
18 Jul 2013
United Kingdom
New Proposal
Login to your account and send a proposal now to get this project.
Log inClarification Board Ask a Question
-
There are no clarification messages.
We collect cookies to enable the proper functioning and security of our website, and to enhance your experience. By clicking on 'Accept All Cookies', you consent to the use of these cookies. You can change your 'Cookies Settings' at any time. For more information, please read ourCookie Policy
Cookie Settings
Accept All Cookies