Putting AI-Translated Text Back into Manga Bubbles Wasn’t So Simple

Yuxino,•Making Things
中EN

I experimented with recognizing Japanese text in manga, translating it into Chinese, and putting it back in the original regions. That keeps the translation on the page. Changing the language does not automatically make it comfortable to read, though.

The words are part of the image, so they cannot be copied like text on a web page. First I need to find them and turn those pixels into text for translation. That is OCR.

Find a region, then read its text

Japanese manga often has vertical text and handwriting mixed with characters and linework. The current Japanese pipeline uses two tools in sequence. comic-text-detector (opens in a new tab) finds text regions. The program crops each region and passes it to the ONNX version of manga-ocr-2025 (opens in a new tab) to read the characters. Detection and recognition are consecutive steps.

The crop determines what the recognizer can see. Half a sentence in the crop means half a sentence to recognize. Nearby drawing lines may also be mistaken for characters. When the translation looks wrong, checking that crop is more useful than immediately changing the translation prompt.

Putting the Chinese back in the original position

The detected region is a rectangular text box, not the complete outline of a speech bubble. Its position and size are kept as proportions of the original image. A box starting a quarter of the way across the page still starts there when the image is resized. The translation can then follow the bubble.

Each piece of dialogue also gets a number before going to DeepSeek for Chinese translation. The response includes those numbers, letting the program match each translation to its original box without relying on response order.

The number links the same bubble across both pages. This illustrates the method, not an OCR test.

Recognition runs locally. Translation needs a connection and my own API key, sending recognized text rather than the page image. That also removes some context: the translator cannot see expressions or actions. When dialogue omits the subject, the words alone may not tell it who is being discussed.

If the Chinese seems wrong, I can compare the OCR text with the Japanese image first. Misrecognized text is a problem before translation; if that text is correct, I can examine the translation instead.

Two characters split across columns

These crops come from the same page in an earlier Mac test, rather than showing how every page turns out. The thin red outlines mark text regions. Click either image to enlarge it.

Original Japanese
Chinese translation

The right-hand box contains Chinese, but the two characters in 孩子 (“child”) have split across columns. The type is heavy and the ellipsis looks like scattered dots. It is readable, but awkward.

The current renderer tries a font size based on the region and estimates how many columns it needs. If their combined width will not fit, it reduces the font size and tries again. Vertical columns run from right to left, the block is centered horizontally in the box, and some punctuation is rotated before drawing.

That calculation checks whether characters fit, without choosing column breaks based on words. Both characters in 孩子 stay inside the box, but separating them still makes reading awkward. That is something this approach does not handle well yet.

Covering the original text has a similar limitation. The renderer samples an approximate background color, paints over the Japanese, and draws Chinese on top. Plain white bubbles are easier. A patch of color cannot restore texture or character linework that it covers.

Comments