Skip to main content
Conversion
Format conversion

CBZ/CBR → Markdown/DOCX

Converting a novel from CBZ/CBR to Markdown/DOCX: what each format is, what gets lost, the pitfalls for long fiction, and the right tools.

Tide Reader Editorial6 min read

What is CBZ/CBR?

A comic book archive is not a real format but a folder of images packed in order: .cbz is ZIP, .cbr is RAR, .cbt is TAR and .cb7 is 7z, renamed so readers recognise it. There is no text layer and no typography — each page is a single full-page image, and reading order comes from filename sorting (which is why unpadded names like 1.jpg, 2.jpg … 10.jpg put page 10 right after page 1).

Extensions.cbr.cbz.cbt.cb7
Typical useComics, art books and books scanned to images; fundamentally an image bundle rather than a text book.

What is Markdown/DOCX?

Markdown (.md) and Word documents (.docx) are grouped here because they play the same role in conversion: lightweight structured authoring sources. Markdown is plain text plus markers (# headings, ** bold) — clean but limited. A .docx is a ZIP of Office Open XML with far richer layout (pagination, headers, footers, comments, tracked changes). Both are good conversion sources rather than final reading formats.

Extensions.md.docx
Typical useAuthoring and layout source files; an intermediate format before converting to EPUB or TXT.

What to watch out for when converting CBZ/CBR to Markdown/DOCX

Extracting from the source

A comic archive is a renamed archive: .cbz=ZIP, .cbr=RAR, .cbt=TAR, .cb7=7z. It has no text layer, and page order comes from filename sorting — if names are unpadded (1.jpg, 2.jpg … 10.jpg) page 10 lands right after page 2, so rename with zero padding first.

Writing the target format

Producing .docx suits further editing, but reading is not its goal: pagination shifts with Word versions and fonts, and phone rendering is fragile. If the destination is an e-reader, EPUB is the right endpoint.

Where the two formats interact

  • Layout must be rebuilt: the paragraph, column and figure positions in the source were arranged for a fixed page and mean nothing once the target reflows. The converter can only re-stitch a text stream, so paragraphs may merge wrongly, columns may interleave and figures may land mid-sentence. Read the first few chapters afterwards, watching dialogue breaks and figure placement.
  • The source has no text layer — every page is an image. No text-targeted conversion can conjure text: you either get blank output or the image inserted as one oversized object. Getting searchable, reflowable text requires OCR first, and OCR on Chinese fiction introduces character errors that need proofreading.
  • The source is one full-page image per page while the target reflows: each image is inserted in sequence, and with no text layer you end up with an image stream. The result is usually huge and unsearchable — a worse reading experience than the original.
  • The source has no chapter concept (comics and PDFs are page-based) while the target expects navigation. The converter can only split mechanically by page or fixed length, producing a semantically empty TOC. Ignore it if you do not need navigation, but do not expect real chapters to appear.

What gets lost

  • Both sides are structurally similar; only fine typographic detail is lost

What you gain

  • Reflow: text re-lays-out for any screen and font size
  • An editable, annotatable source document

Post-conversion checklist

  1. Convert the first three chapters as a trial and check for lost text, interleaved lines or misplaced blocks before doing the whole book.
  2. Check chapter splitting: the TOC entry count should match the real chapter count.
  3. Confirm illustration count and placement — check that images are not all dumped at the end of a chapter.
  4. Fill in title, author and language — a wrong language tag breaks typesetting such as Chinese line breaking.
  5. Verify page order: filenames must be zero-padded (001.jpg) or page 10 follows page 2.

Converting novels: what is different

Fiction is the most common and the most peculiar kind of e-book: hundreds or thousands of chapters, almost no figures, often nothing but a TXT source, and Chinese sources that frequently carry anti-piracy garbling. These are the points that bite hardest.

  • A long novel can run to hundreds or thousands of chapters, so the table of contents decides whether the result is usable. Normalise chapter headings to one recognisable pattern first — say "Chapter N Title" alone on its own line — or the target will miss chapters, merge them, or collapse the whole book into one.
  • Volume markers ("Volume One", "Volume Two") usually degrade to ordinary headings, losing the volume/chapter hierarchy. If that matters, promote volume headings one level in an editor afterwards so they are not siblings of chapter titles.

Tools for converting CBZ/CBR to Markdown/DOCX

Any of the tools below can handle this direction; pick one by your operating system and how many files you have. Read the previous section first — doing the steps out of order (converting a scanned file before running OCR, for instance) usually means starting over.

Kindle Comic Converter

FreeDesktop app

A converter built for comic archives: CBZ/CBR/CB7 in, output tuned for Kindle, Kobo and similar devices. It handles image scaling, greyscale, double-page spreads and zero-padding of filenames, all of which are tedious by hand.

Platforms: Windows / macOS / Linuxgithub.com/ciromattia/kcc ↗

Calibre

FreeDesktop app

The de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.

Platforms: Windows / macOS / Linuxcalibre-ebook.com ↗

SumatraPDF

FreeReader

A lightweight reader for PDF, EPUB, MOBI, CBZ and more that can save PDF content as text. It is a fast way to check whether text extracts cleanly — if the selection comes out garbled, OCR is needed before converting.

Platforms: Windowswww.sumatrapdfreader.org ↗

Common questions

What is lost when converting CBZ/CBR to Markdown/DOCX?

Both sides are structurally similar; only fine typographic detail is lost.

Which tool should I use to convert CBZ/CBR to Markdown/DOCX?

Recommended: Kindle Comic Converter, Calibre, SumatraPDF. Kindle Comic Converter fits this pair best. A converter built for comic archives: CBZ/CBR/CB7 in, output tuned for Kindle, Kobo and similar devices. It handles image scaling, greyscale, double-page spreads and zero-padding of filenames, all of which are tedious by hand.

What should I watch out for when converting CBZ/CBR to Markdown/DOCX?

Layout must be rebuilt: the paragraph, column and figure positions in the source were arranged for a fixed page and mean nothing once the target reflows. The converter can only re-stitch a text stream, so paragraphs may merge wrongly, columns may interleave and figures may land mid-sentence. Read the first few chapters afterwards, watching dialogue breaks and figure placement. The source has no text layer — every page is an image. No text-targeted conversion can conjure text: you either get blank output or the image inserted as one oversized object. Getting searchable, reflowable text requires OCR first, and OCR on Chinese fiction introduces character errors that need proofreading. The source is one full-page image per page while the target reflows: each image is inserted in sequence, and with no text layer you end up with an image stream. The result is usually huge and unsearchable — a worse reading experience than the original. The source has no chapter concept (comics and PDFs are page-based) while the target expects navigation. The converter can only split mechanically by page or fixed length, producing a semantically empty TOC. Ignore it if you do not need navigation, but do not expect real chapters to appear.

Converted it — now what do you read it with?

Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.

Download Tide Reader

TXT and EPUB; local-first reading, with optional cloud sync and WebDAV.