No single-file HTML output with inlined assets#720
Conversation
Add --single-file (-%Z): after the mirror completes, rewrite every saved page in place with its stylesheets, scripts, images and fonts embedded as data: URIs, while links between pages stay relative. The mirror remains a browsable tree and each page also stands alone. This cannot reuse the -%M path: MHT streams a MIME part per file as it is saved, but a data: URI needs the asset's bytes when the page is written, and pages are normally saved before their assets are fetched. The new pass runs over the finished tree at the tail of httpmirror(), after the update purge. Audio, video, page-to-page links and anything over --single-file-max-size (10 MB default) keep an ordinary link. References carrying a scheme are skipped, which covers data: and makes a second --update run a no-op. Resolution is clamped to the mirror root: the HTML is hostile input. htsopt.h gains two tail-appended fields; VERSION_INFO is untouched. Closes #713 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
Master's #709 took test number 82, so the single-file crawl test moves to 86. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Master took 82 through 85 while this branch was open. The engine self-test also moves to the end of the script: it and the crawl assertions cover different ground, and failing first hid the crawl half. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
sf_relative_from() indexed one byte past from_dir's terminator whenever the page's own directory was a prefix of the asset path, which is the ordinary mirror layout: an out-of-bounds read that also miscounted the ../ prefix and emitted a broken link. Two review agents hit the same ASan trace; the branch had no test because every nested asset in the fixture was under the cap. Also from review: an escaped quote no longer ends a CSS string early (which exposed its contents to the url()/@import scanner), @import url(...) now inlines like the quoted form, a raw-text element ends only on a real end tag rather than any prefix of one, an over-wide tag is copied through by the same quote-aware scan instead of a second quote-blind one, <!--> is an empty comment, whitespace before a tag's > survives, and a page cannot inline more than SINGLEFILE_MAX_PAGE_SIZE, which bounds the multiplicative @import fan-out. A failed encode no longer leaves a payload-less data: prefix behind, and base_dir is matched against the root on a component boundary. sf_readfile now wraps a new readfile2_utf8() rather than being a fifth copy of the readfile family. Docs place --single-file beside -%M instead of implying it supersedes it: MHT stays the better container, this wins on opening anywhere. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
copy_htsopt's two new fields, including the >0 guard that must not let an unset source clear the target's default; -%Z0; and a rejected --single-file-max-size argument falling back to the built-in cap. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
The Windows job builds httrack.exe only, and reaches this file through its *_local-*.test glob, so requiring htsserver failed the whole test there even though every crawl assertion had passed. Skip that half instead, and run the engine self-test before it so the skip cannot swallow it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
Master took test number 86 for the proxytrack cache-longfields test; both that and 87_local-single-file.test now sit in TESTS. The only conflict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
#722 gave slprintfbuff the attribute, so the over-wide-tag fixture's call became the one warning this branch adds over master's baseline. Handle it the way that PR's own self-test code does; the buffer cannot truncate here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <xroche@gmail.com>
|
Two notes for anyone reading the commit list rather than the diff. The branch's first commit still says Separately, #726 tracks something this PR noticed but deliberately did not fix: |
…les (#739) The feature PR added the four LANG_SINGLEFILE* entries to lang.def with English and Francais only; every other locale fell back to English in the WebHTTrack form. Each file is written in its own declared charset. Signed-off-by: Xavier Roche <roche@httrack.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Test 87 was taken by master (proxytrack-nodate) and 90 by the web GUI checkbox test, so the single-file test moves to 91. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
…overage The rewrite spools to a .sfnew temp and renames it over the page, which bypasses filecreate() and with it the HTS_ACCESS_FILE chmod the engine puts on every other mirrored file. Under a restrictive umask the pages came out 0600 while their assets stayed 0644, so a mirror served by a webserver or shared with a group lost read access on exactly the pages. chmod the spool before the rename. Two guards had no coverage: deleting the scheme/data: check in sf_resolve left both the self-test and the crawl test passing, because the fixtures resolved to paths that were absent either way, and the per-page inline budget was never exercised. The fixtures now plant a file where each guard's removal would land the walk, and a self-importing stylesheet measures the budget against a large-budget control. Charging that budget after the nested rewrite instead of before let an @import chain spend what its ancestors had already claimed and drove it negative; charge it up front and refund on failure. The attribute table missed the lazy-loading attributes hts_detect[] already downloads, so a modern page inlined almost nothing: add data-src, data-srcset, lowsrc, object@data and embed@src, and record why the rest stay links. The new web GUI checkbox had no hidden companion input, so it could be ticked but never cleared, which is what #725 fixed for every other box. Master's test 90 catches it once the branch merges. Adds a libFuzzer harness over the rewriter, since it re-serializes hostile HTML, and renames the crawl test to 91 now that master holds 87 and 90. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Master took test 91, so the single-file test moves to 92. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Recorded in htssinglefile.h already; the user-facing guide is where someone raising --single-file-max-size will look. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Conflict in tests/Makefile.am only: master took 89, so the change-report test moves to 93 (PR #720 holds 92). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Its GUI half needs htsserver, which the Windows job does not build, so the test now exits 77 there instead of reporting a pass for assertions it never ran. The skip list is pinned, so it has to be declared. Both Windows jobs reported fail=0; only the list check was red. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Master took test 92 and added an expected Windows skip; the single-file test moves to 93 and both skip lists are combined. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Master landed the change report (#721) and tests 88-93 while this branch was open. Every conflict was both branches appending to the same tail, so each is resolved as a union: htsopt.h keeps the change-report fields at their published offsets and appends the single-file pair after them, and the lang files, GUI templates and selftests carry both features' entries. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
Master took 88 through 93 and #720 claims 94. Signed-off-by: Xavier Roche <xroche@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com>
#718) * Read sitemap files so URLs nothing links to are found HTTrack finds URLs only by parsing links, so anything a site publishes solely in its sitemap stayed invisible: robots.txt was already parsed, but its Sitemap: lines were ignored and nothing else in the tree touched sitemaps. Adds opt-in --sitemap (-%m), which probes the start host's robots.txt and falls back to /sitemap.xml, and --sitemap-url (-%mu) for an explicit document. Handles <urlset> and nested <sitemapindex>, plain or gzipped. Discovered URLs enter with the full depth budget but still go through the wizard, so filters and scope rules decide; a sitemap is not a filter bypass. The parser reads attacker-controlled XML off the network, so it is capped on URL count, index nesting, decompressed size and decompression ratio, and child sitemaps must stay on the host that named them. Closes #712 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: distinguish the sitemapindex log line from a urlset one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: fix an out-of-bounds read, tighten the parser and the tests lienrelatif() walked back from the last character of its current-path argument without checking the path was non-empty, reading one byte before the stack buffer. htsAddLink is the first caller to pass an empty savename, which sitemap documents have because they are ingested rather than mirrored, so ASan caught it on the new crawl test. The parser drops a value whose numeric character reference decodes outside printable ASCII, rather than leaving the reference verbatim and seeding a URL the site never published, and classifies a document by its real root element, so a comment naming the other one no longer flips urlset and sitemapindex. The robots.txt line reader is bounded by the body size instead of relying on a NUL terminator. The self-test moves to 01_zlib-sitemap.test: MSan runs 01_engine-* only, because an uninstrumented libz floods it with false positives. Tests gain the assertions the earlier ones were missing: which of the robots.txt route and the /sitemap.xml fallback was taken, that the sitemap documents stay out of the mirror, that the off-host child sitemap is refused, the sitemapindex nesting cap, the per-document URL cap at its production value, and copy_htsopt coverage for the two new fields. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: state what the decompression cap actually binds on deflate tops out near 1032:1, so hts_codec_maxout never binds before the 64 MiB cap; the old comment implied a ratio guard that cannot fire. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: a child sitemap is a fetch, so filters and robots.txt must gate it An adversarial review found that a <sitemapindex> <loc>, and a robots.txt Sitemap: line, went straight to hts_record_link: the request went out even when a -* rule or robots.txt Disallow covered it. Only the <urlset> half ran through the wizard. Gate the document itself on the filters and on robots.txt, which is all that can apply: the wizard proper wants a referring link, and its up/down travel rules would judge a child sitemap against the parent sitemap's own directory. The robots.txt probe is exempt, being the request that fetches the rules. A 301 also used to end ingestion silently, since the engine re-queues the target as a fresh link that carried no sitemap marking. That hit any site redirecting http to https. The marking now follows the redirect. The "N URL(s) added" counter reported what the scanner emitted rather than what was taken, which hid both of the above; it now reads "N of M". The fallback to /sitemap.xml keys on the same corrected count, so a robots.txt whose only Sitemap: line is off-host or filtered still falls back. Root classification skips a UTF-8 BOM and an XML namespace prefix, and the doc list is cleared when a mirror starts rather than only when it ends. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: gate each fetch by who asked for it, and anchor travel on the start URL A nine-agent review found a scope escape: the sitemap document was its own `premier`, so the wizard measured travel from wherever the site chose to put its sitemap. A root /sitemap.xml therefore widened a /deep/dir/ crawl to the whole host. The ingester now points the wizard at the crawl's own start link and lets each seeded URL become its own anchor, which is what a command-line seed gets. Robots handling was both mistimed and undifferentiated. The Sitemap: lines are now collected by robots_parse, on the same body in the same fetch, and acted on after the parsed rules are installed rather than before; and the decision comes from a new hts_robots_forbids extracted out of the wizard, so the sitemap path inherits the -s1 filters-win override instead of a stricter hand-rolled check. The four fetches are no longer treated alike: a sitemap the user names is user intent, one the site declares invites the fetch, only the guessed /sitemap.xml obeys a Disallow, and the URLs listed inside stay fully gated. Also: hts_unescapeEntities replaces the private entity decoder, whose guard tests and fuzz corpus it silently forfeited; hts_codec_head replaces hts_zhead, which is only defined under HTS_USEZLIB; the composed URL buffer now fits two maximal components plus a scheme, which a 2046-byte --sitemap-url reached; the bounded search is promoted to htstools as hts_memstr; and the live state moves from httrackp into htsoptstate, leaving two installed fields rather than three. Tests gain the scope escape, the three robots cases, a cap-boundary control, the handler invocation count and a compression-bomb decode. Every one was checked against a deliberately broken build. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: add the fuzz harness the parser was missing, and drop truncated Sitemap: lines The file header called the scanner fuzzable while fuzz/ registered ten harnesses and none for it. fuzz-sitemap feeds it raw XML, gzip-framed bodies and truncated streams off a heap copy of exactly the input size, so an overread is an ASan report rather than a quiet pass, with a four-file seed corpus. 60000 runs clean under ASan+UBSan. robots_parse now drops a Sitemap: line that filled its scratch buffer instead of handing on the half URL it was truncated to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * fuzz: keep only the four sitemap seed inputs A libFuzzer run writes its finds into the first corpus directory, and 191 of them were committed with the harness. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * selftest: bound the sitemap document builders' snprintf accumulation snprintf returns the length it wanted to write, so accumulating it blind lets the next offset and size argument walk past the buffer. Guard each step the way the argv builder above already does, and give the per-URL loop a real remaining-space bound instead of a fixed 33. Signed-off-by: Xavier Roche <xroche@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: date the new files 2026 The headers were copied from an existing file and kept its 1998 year. Signed-off-by: Xavier Roche <xroche@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: keep the ingestion state out of htsoptstate htsoptstate is embedded by value as httrackp.state, so a field at its tail shifts every httrackp member declared after it: an offsetof probe put warc_file at 141752 on master and 141760 on the branch. Move the pointer to httrackp's own tail, where every existing offset holds and copy_htsopt still ignores it. Also renumber the crawl test to 89, master having taken 87 and 90, and give the new option8 checkbox the hidden companion that 90_webhttrack-checkbox-clear requires, plus its row in that test's table. Signed-off-by: Xavier Roche <xroche@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * tests: renumber the sitemap crawl test to 95 Master took 88 through 93 and #720 claims 94. Signed-off-by: Xavier Roche <xroche@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: realign the htsopt.h comments after the single-file merge Master's longer LLint declarator moved the block's comment column. Signed-off-by: Xavier Roche <xroche@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Xavier Roche <roche@httrack.com> * sitemap: translate the new GUI strings into the remaining 28 locales (#738) The feature PR added the four LANG_SITEMAP* entries to lang.def with English and Francais only; every other locale fell back to English in the WebHTTrack form. Each file is written in its own declared charset. Signed-off-by: Xavier Roche <roche@httrack.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Signed-off-by: Xavier Roche <roche@httrack.com> Signed-off-by: Xavier Roche <xroche@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
HTTrack's only one-file output was
-%M, amultipart/relatedMHT archive. Firefox and Safari never read those, so you cannot hand someone an.mhtand count on it opening.--single-file(-%Z) produces ordinary HTML instead: once the mirror completes, every saved page is rewritten in place with its stylesheets, scripts, images and fonts embedded asdata:URIs. Links between pages stay relative, so the mirror is still a browsable tree and any single page can also be handed to someone on its own. The assets do end up on disk twice.-%Mstays, and the two are good at different things: MIME carries text parts without the base64 tax, keeps each resource's original URL and stores a shared asset once, while single-file HTML wins on opening by double-click anywhere. The help text and the guide say which to reach for; neither presents one as the successor.It cannot reuse the MHT path.
postprocess_file()streams a MIME part per file as it is saved, but adata:URI needs the asset's bytes at the moment the page is written, and pages are normally saved before their assets are fetched. So this is a pass over the finished tree, hooked at the tail ofhttpmirror()after the update purge. By then the engine has made every reference relative, so resolving one is path arithmetic against the page's own directory, clamped to the mirror root: the HTML is hostile input and a crafted../../..must not reach a file outside the project. A reference carrying a scheme is skipped, which coversdata:and makes a second--updaterun a no-op rather than a double-encode. Audio and video keep their links, as do page-to-page links and anything over--single-file-max-size(10 MB by default, reported in the log).htsopt.hgains two tail-appended fields; I have not touchedVERSION_INFO.An inlined stylesheet becomes a
data:URL, and per the URL Standard adata:URL has an opaque path. So a relative reference left inside one does not resolve: the parser returns failure, whatever path we write there. That only bites an asset the stylesheet referenced and the cap kept as a link, and raising--single-file-max-sizepast it inlines the asset too.htssinglefile.hand the docs record the limitation.-#test=singlefilebuilds a fixture mirror and asserts what is inlined, what keeps its link, the cap, the per-page inline budget, the mirror-root clamp and idempotence.tests/93drives a real crawl, decodes the emitted payloads with Python and compares them byte for byte against the served files, checks the rewritten page keeps the mode the engine gave it, and checks the webhttrack form still emits the right command line. The negative-case fixtures plant a real file wherever deleting a guard would send the resolver, so no assertion can pass just because the target was missing; I deleted each guard in turn and watched the suite go red. A libFuzzer harness covers the rewriter, since it re-serializes attacker-controlled HTML.The option is wired into the webhttrack Build tab, all 30 locales, the help text, a regenerated man page and the command-line guide.
Closes #713