PDFTOHTML(1) General Commands Manual PDFTOHTML(1)
pdftohtml - program to convert PDF files into HTML, XML and PNG images
pdftohtml [options] <PDF-file> [<HTML-file> <XML-file>]
This manual page documents briefly the pdftohtml command. This manual
page was written for the Debian GNU/Linux distribution because the
original program does not have a manual page.
pdftohtml is a program that converts PDF documents into HTML. It gener-
ates its output in the current working directory.
A summary of options are included below.
Show summary of options.
first page to print
last page to print
-q do not print any messages or errors
-v print copyright and version info
-p exchange .pdf links with .html
-c generate complex output
-s generate single HTML that includes all pages
-i ignore images
generate no frames. Not supported in complex output mode.
use standard output
zoom the PDF document (default 1.5)
-xml output for XML post-processing
output text encoding name
owner password (for encrypted files)
user password (for encrypted files)
force hidden text extraction
-fmt image file format for Splash output (png or jpg). If complex is
selected, but -fmt is not specified, -fmt png will be assumed
do not merge paragraphs
-nodrm override document DRM settings
adjust the word break threshold percent. Default is 10. Word
break occurs when distance between two adjacent characters is
greater than this percent of character height.
outputs the font name without any substitutions.
Pdftohtml was developed by Gueorgui Ovtcharov and Rainer Dorsch. It is
based and benefits a lot from Derek Noonburg's xpdf package.
This manual page was written by Soren Boll Overgaard <email@example.com>,
for the Debian GNU/Linux system (but may be used by others).
pdfdetach(1), pdffonts(1), pdfimages(1), pdfinfo(1), pdftocairo(1),
pdftoppm(1), pdftops(1), pdftotext(1)pdfseparate(1), pdfsig(1), pdfu-