Skip to content

Tika based link (URL) extractor for httpreserve

Notifications You must be signed in to change notification settings

httpreserve/tikalinkextract

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

38 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

tikalinkextract

Tika client for httpreserve

Demo

asciicast

Use with Wget

Extract the links from your files using seeds option

./tikalinkextract -seeds -file archives-nz-demo/ > transferlinks.txt

Use the seeds to generate a warc file

wget --page-requisites --span-hosts --convert-links  --execute robots=off --adjust-extension --no-directories --directory-prefix=output --warc-cdx --warc-file=accession.warc --wait=0.1 --user-agent=httpreserve-wget/0.0.1 -i transferlinks.txt 

Known Issues

  • HTTP links that are formatted in such a way to be split across lines, thus include a newline \n character.

Resources that might help

License

Tika is licensed as follows: https://www.apache.org/licenses/

This tool is licensed GNU General Public License Version 3. Full Text