Fix CHS (Chesterfield) scraper - #484
Open
symroe wants to merge 1 commit into
Open
Conversation
… bypass The scraper was hitting http://chesterfield.moderngov.co.uk which returned 403. Chesterfield's ModernGov instance is now behind Cloudflare and requires both HTTPS and a browser-level JS challenge to access the ASMX web service endpoint. - Migrate base_url from http:// to https:// - Add http_lib = "playwright" to execute Cloudflare JS challenge via headless Chromium
Member
Author
Re-scrape after 8295740Initial fix: HTTPS migration + Playwright for Cloudflare bypass. Cloudflare blocks all non-Chromium clients from this environment (both wreq and curl receive a 403 "Attention Required" challenge page), so scrape counts can only be confirmed after a Lambda run where headless Chromium will execute the JS challenge.
Generated by Claude Code |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What broke
The scraper was requesting
http://chesterfield.moderngov.co.uk/mgWebService.asmx/GetCouncillorsByWardand receiving a 403. Chesterfield's ModernGov instance is now behind Cloudflare, which blocks the HTTP endpoint entirely and serves a 403 Cloudflare challenge page even on the HTTPS endpoint when accessed with a non-browser TLS fingerprint (confirmed:wreqon HTTPS returns a 4 551-byte Cloudflare HTML block page with"Attention Required! | Cloudflare"title).What was fixed
metadata.json: migratedbase_urlfromhttp://tohttps://councillors.py: addedhttp_lib = "playwright"so headless Chromium executes the Cloudflare JS challenge before fetching the ModGov ASMX XML endpointScrape results
Cloudflare blocks all non-Chromium clients from this environment, so exact counts require a Lambda run. This is the same fix pattern applied to BNS (#472), BRW (#479), CHI (#473), and other ModernGov instances recently upgraded to Cloudflare protection.
Generated by Claude Code