add PubMed guideline scraper - #15
Open
zndr27 wants to merge 2 commits into
Open
Conversation
zndr27
marked this pull request as ready for review
July 30, 2026 00:24
Scrapes the intersection of PubMed's Guideline publication type with the PMC Open Access subset, restricted to English: about 3,000 guidelines. That is the only slice where the full text is both retrievable and openly licensed; the other 40k Guideline records expose an abstract only, under publisher copyright. Two scoping choices, both measured rather than assumed: `Guideline[pt]` rather than `Practice Guideline[pt]`. The former is a strict superset and the 352 records it adds are clinical, not administrative (the 2025 Korean CPR guidelines and similar), so the narrower tag would drop 11% of the corpus for nothing. `English[la]`, which drops 185 records. Every other source in this package is already English-only as a side effect of its entry URL: CPS is a bilingual site scraped through its /en/ routes, WHO publishes in six languages and is scraped through its English listing. PubMed's API returns every language, so the filter has to be explicit to match. The excluded records are largely French CMAJ translations of guidelines already in the corpus. Discovery and extraction use NCBI's E-utilities rather than the rendered pages. esearch paginates by numeric offset, so the page number maps straight onto retstart and an empty page past the end terminates the run. esummary resolves a whole batch of PMIDs to PMCIDs and citation metadata in one request. efetch returns JATS XML carrying body, section structure and license. JATS is close enough to HTML that renaming tags and reusing html_to_markdown is cheaper and less error-prone than a second serializer: table-wrap already contains genuine XHTML tables, and inline markup maps one to one. Licensing is recorded per document rather than claimed for the source, because the Open Access Subset is not uniformly Creative Commons licensed. Censused over all 2,999 English records present on 2026-07-29: CC BY 44.5% publisher terms, no CC license 11.7% CC BY-NC 22.1% of which: Elsevier COVID grant, no CC BY-NC-ND 18.4% <license> element at all (112), CC BY-NC-SA 2.1% PMC OA "unrestricted re-use" CC0 1.3% By what that permits: 45.8% unrestricted for derivative works, 24.1% non-commercial only, 18.4% asserting NoDerivatives, 11.7% needing a case-by-case reading. Presence in the subset is not itself a grant to redistribute: 112 records carry only a copyright line such as "(c) Springer-Verlag Tokyo 2007", and Elsevier's pandemic-era deposits grant free access while still reserving all rights. The license name is parsed from the Creative Commons URL rather than the license-type attribute, which the corpus spells 18 different ways. The copyright statement is captured separately because it is a sibling of <license>, not a child, and it holds the reservation of rights. Figures and supplementary files are recorded in metadata rather than linked. Unlike the HTML sources in MedARC-AI#9 and MedARC-AI#10, which absolutize a real <img src>, JATS carries only a bare filename; the served URL inserts a CDN shard and content hash that appear nowhere in the API response, so a constructed link 404s. Supplementary blocks are pointers too: across 40 sampled guidelines every one referenced an external .docx or .tif rather than inline content, totalling 0.18% of body text. Recording name, label and caption keeps the evidence tables findable without re-scraping. E-utilities calls retry with backoff on 429 and 5xx. One document makes up to two calls back to back and the rate limit is per source address, so NCBI does answer with 429 in practice; without a retry that propagates past the skip handler and kills the whole run. external_id prefers the PMID from the record itself, so an article reached from a PMC URL gets the same identifier as one reached from the listing. Records PMC holds without a deposited body, 0.6% of the corpus, are logged and skipped rather than aborting the run.
Retracted articles keep their `Guideline` publication type and stay in the PMC Open Access subset, so neither the type filter nor the open-access filter excludes them. PMID 37026270, a retracted 2023 rosacea practice pattern, is in the result set today. This is the one place the scraper decides rather than records: licence terms are a policy question with legitimate answers either way, but a retracted guideline is precisely the document a fact verifier must not retrieve as evidence. Prior art: MedPMC (arXiv:2607.07673) filters retraction status the same way when curating PMC at scale.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
Adds a PubMed scraper to the datasets scraping pipeline. Same shape as the NICE scraper: discovery and extraction are separate, and everything outputs a normalized
ScrapedDocument.Scope is the PMC Open Access Subset, the only slice where the full text is both retrievable and openly licensed:
3,030 records as of 2026-08-19. Retracted articles keep their
Guidelinetype and stay in the subset, so nothing else excludes them; they are dropped at search time because a retracted guideline is precisely the document a fact verifier must not retrieve as evidence.How discovery works
esearchpages by numeric offset, 200 per page, and returns the total for the progress bar.esummaryresolves each batch to PMCIDs and citation metadata in one call. Offset paging maps straight ontolist_page(client, page), and an empty page past the end terminates the run.How extraction works
efetchreturns JATS XML per document. Tags are rewritten to HTML and passed through the sharedhtml_to_markdownhelper rather than writing a second serializer. Headings come from<sec>nesting depth; reference lists and<xref>pointers are dropped as citation apparatus.Figures and supplementary files are recorded in
metadatarather than linked: JATS carries a bare filename and the served URL adds a CDN shard and content hash absent from the API response, so a constructed link 404s. Captions stay in the text.Delay is 0.4s, inside NCBI's documented 3 requests/second. Calls retry with backoff on 429, which a real run does draw. Records PMC holds without a body are logged and skipped.
Licensing
The Open Access Subset is not uniformly Creative Commons licensed, so this records rather than decides: each document carries
license,license_url,license_type,license_statementandcopyright_statement. The full census is in the module docstring. A corpus-wide filter probably wants a project-wide answer.CLI
--source allincludes PubMed.--urltakes a PubMed or PMC URL;external_idalways prefers the record's own PMID so both routes give the same ID.Tests
58 tests, offline via
httpx.MockTransport. Rebased onto #21, so registration is one import and one entry inSCRAPERS. Live run of two guidelines gave documents of 86,354 and 25,574 characters.