HTML → FB2
Converting a novel from HTML to FB2: what each format is, what gets lost, the pitfalls for long fiction, and the right tools.
What is HTML?
HTML / XHTML is the structural language of the web and is literally the content format inside an EPUB; .mhtml (MHTML) packs a page together with all its resources (images, CSS) into one file, which is what "save as a single page" produces. XML is the general markup language that formats such as FB2 and OPF are themselves written in. The whole family is tagged text where the tags carry the structure, so how much structure survives a conversion depends on how well the tool understands them.
.html.htm.xhtml.xml.mhtmlWhat is FB2?
FB2 (FictionBook 2) is an open XML-based e-book format that is very popular in Russian-speaking regions. Body text, metadata and illustrations all live in one XML file, with images usually embedded as Base64, so a single .fb2 is a complete book that is easy to distribute. It has clean structure and carries chapters and author data natively; the drawbacks are larger size (Base64 inflates images by roughly a third) and less reader support outside Russian-language markets.
.fb2What to watch out for when converting HTML to FB2
Extracting from the source
All structure lives in the tags, but real pages mix in navigation, ads and footers that a naive conversion carries along. Extract the content container first (a readability pass), and note that <img> usually points outside the file — images must be downloaded and embedded before producing a single-file format.
Writing the target format
FB2 requires structured XML body text, Base64-embedded images and metadata inside <description>. It supports almost no CSS, so it suits text-plus-illustration books rather than layout-heavy titles.
Where the two formats interact
- Chapter headings in the source are ordinary lines while the target demands a real table of contents. Converters try to match patterns such as "Chapter N"; when the match is poor the TOC is garbled or the book has a single chapter. Check that TOC entries match the actual chapter count.
What gets lost
- Both sides are structurally similar; only fine typographic detail is lost
What you gain
- A real table of contents with chapter jumps
Post-conversion checklist
- Convert the first three chapters as a trial and check for lost text, interleaved lines or misplaced blocks before doing the whole book.
- Check chapter splitting: the TOC entry count should match the real chapter count.
- Jump to five random spots and search a character name to confirm the text is fully searchable.
- Confirm illustration count and placement — check that images are not all dumped at the end of a chapter.
- Fill in title, author and language — a wrong language tag breaks typesetting such as Chinese line breaking.
Converting novels: what is different
Fiction is the most common and the most peculiar kind of e-book: hundreds or thousands of chapters, almost no figures, often nothing but a TXT source, and Chinese sources that frequently carry anti-piracy garbling. These are the points that bite hardest.
- A long novel can run to hundreds or thousands of chapters, so the table of contents decides whether the result is usable. Normalise chapter headings to one recognisable pattern first — say "Chapter N Title" alone on its own line — or the target will miss chapters, merge them, or collapse the whole book into one.
- Volume markers ("Volume One", "Volume Two") usually degrade to ordinary headings, losing the volume/chapter hierarchy. If that matters, promote volume headings one level in an editor afterwards so they are not siblings of chapter titles.
Tools for converting HTML to FB2
Any of the tools below can handle this direction; pick one by your operating system and how many files you have. Read the previous section first — doing the steps out of order (converting a scanned file before running OCR, for instance) usually means starting over.
Calibre
FreeDesktop appThe de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.
CloudConvert
FreeOnline serviceA general online conversion service with no install and broad format coverage, fine for the occasional file. Files are uploaded to a third party, so avoid it for copyrighted or private content, and the free quota is limited.
SumatraPDF
FreeReaderA lightweight reader for PDF, EPUB, MOBI, CBZ and more that can save PDF content as text. It is a fast way to check whether text extracts cleanly — if the selection comes out garbled, OCR is needed before converting.
Common questions
What is lost when converting HTML to FB2?
Both sides are structurally similar; only fine typographic detail is lost.
Which tool should I use to convert HTML to FB2?
Recommended: Calibre, CloudConvert, SumatraPDF. Calibre fits this pair best. The de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.
What should I watch out for when converting HTML to FB2?
Chapter headings in the source are ordinary lines while the target demands a real table of contents. Converters try to match patterns such as "Chapter N"; when the match is poor the TOC is garbled or the book has a single chapter. Check that TOC entries match the actual chapter count.
Converted it — now what do you read it with?
Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.
Download Tide ReaderTXT and EPUB; local-first reading, with optional cloud sync and WebDAV.
