-
Notifications
You must be signed in to change notification settings - Fork 121
Intro
Danny Lin edited this page Jan 20, 2024
·
28 revisions
WebScrapBook is a browser extension that captures the web page faithfully with various archive formats and customizable configurations, for future retrieval, organization, annotation, and editing. This project inherits from legacy Firefox add-on ScrapBook X.
- Capture faithfully: A web page shown in the browser can be captured without losing any subtle detail. Metadata such as source URL and timestamp are also recorded.
- Customizable capture: WebScrapBook can save selected area in a page, save source page (before processed by scripts), or save page as a bookmark. How to capture images, audio, video, fonts, frames, styles, scripts, etc. are also customizable. A web page can be saved as a folder, a ZIP-based archive file (HTZ or MAFF), or a single HTML file.
- Page editing: A web page can be highlighted, annotated, or edited before or after a capture.
- Organizable collections: Captured pages can be organized in the browser sidebar using one or more scrapbooks, and each scrapbooks holds a hierarchical tree structure to organize data items. Notes using HTML or markdown format can also be created and managed. (*)
- Fulltext searching: Each scrapbook can be further indexed for a rich-feature search (using title, fulltext, comment, source URL, create time, modify time, etc.). (*)
- Remote access: Captured data can be hosted with a central backend server and be read or edited from other devices. Alternatively, a scrapbook can generate a static site index and be distributed as a static web site. (*)
- Mobile support: WebScrapBook supports mobile browsers such as Firefox for Android and Kiwi browser. You can capture and edit the web page from a mobile phone or tablet.
- Legacy ScrapBook support: Scrapbooks created from legacy ScrapBook or ScrapBook X can be converted into WebScrapBook-compliant format for usage. (*)
- All or partial functionality of a starred feature above requires a running collaborating backend server, which can be easily set up using PyWebScrapBook.
- An HTZ or MAFF archive file can be viewed using the built-in archive page viewer, using PyWebScrapBook or other assistant tools, or by opening the index page after unzipping.
WebScrapBook is available for Chromium-based browsers (Google Chrome, Edge, Opera, Vivaldi, Brave, etc.), and Firefox-based browsers (Firefox for Desktop or Android, Tor Browser, etc.). Just go to the app store of the corresponding browser and install this extension.
Known mobile browsers that support installation of the extension:
- Android: Firefox for Android, Kiwi Browser
- iOS: none
You can also install this extension from source code, as long as the browser supports.
- Download the latest source code from this repository and unpack.
- Run
build/pack.cmd
(on Windows) orbuild/pack.sh
(on POSIX/Linux). - There will be
dist/WebScrapBook.zip
(for Chromium) anddist/WebScrapBook.xpi
(for Firefox) generated for installation.
- Make sure the browser supports installation from a package.
- Known supported: Chrome, Edge, Brave, Kiwi
- Go the the extension management page.
- Check
Developer mode
. - Load the
.zip
package file generated in the above section in the extension management page (through a button likeLoad Package
, or dragging and dropping the file into the management page).
- Make sure the browser supports installation of an unsigned add-on. (Consult this document for more details.)
- Known NOT supported: Firefox Release, Firefox Beta
- Known supported: Firefox ESR, Firefox Nightly, Firefox Developer Edition, Waterfox, Tor Browser
- Set config
xpinstall.signatures.required
tofalse
. (Typeabout:config
in the URL address bar to enter the config page.) - Load the
.xpi
package file generated in the above section in a tab (throughInstall Add-on From File
command in the add-on management page, orFile > Open File
from the main menu, or dragging and dropping the file to the tab list).