Translate webtoons and comics as you scroll
A chapter is split into dozens of separate images, and picking them one by one is not realistic. Live mode reads whatever is on screen, lays the translation over the original lettering, and speaks each bubble as it comes into view.
One bubble, one sentence
Text recognition returns one box per line, and translating those lines separately is what ruins most attempts: a single line of dialogue comes back as four disconnected fragments. Caliptic merges the lines of a bubble into whole sentences before translating, using the spacing between lines rather than the height of the letters — the letter height varies with descenders even at the same size, and that is precisely what made earlier attempts cut a bubble in half. Centered lettering is detected too, so the translation sits centered in the bubble the way the original did.
Scroll; it keeps up
You do not select anything. Turn on live mode and read as you normally would: as new panels come into view they are recognized, translated and covered with the translation, and each block is read aloud when it becomes visible. The voice can be turned off entirely — the translations still appear, which is what you want in a shared room. Reading order follows the vertical flow that webtoons use.
Where it falls short
It is optical recognition, so stylised sound effects and hand-lettered text are hit and miss. Right-to-left panel ordering, as used in Japanese manga, is not handled yet; vertical webtoons are. And it will not match a human scanlation — it is for the series nobody is translating, and for reading ahead. Image translation needs the free macOS app alongside the extension; text-only pages work without it, on any platform.
Does the translation cover the original text or sit beside it?
It covers it. A sidebar makes you look away from the panel and lose the reading flow, which defeats the point in a comic. The translation is drawn as an opaque block over the original lettering, sized to it; if the translation runs longer, the block grows downward rather than clipping the text.
Is anything uploaded when it reads the page?
No. The captured region is cropped in memory, recognized on your own machine and discarded. It is never written to disk by the extension and never sent anywhere. Translation uses the browser's on-device model, and the voice is a local system voice.