Subtitles for whatever my computer is playing

Yuxino,•Making Things
中EN

mimi is a live subtitle tool I built. It transcribes dialogue from videos, games, and livestreams playing on the computer and translates it into the selected language. The original app keeps running, with a movable subtitle window over it.

Expressions and tone help when there are no subtitles, but they only tell me so much about what someone said. A drama may have a subtitle file elsewhere; games and livestreams are harder to cover that way. mimi takes its name from the Japanese word for ears, みみ.

It captures system audio output and sends it to the configured cloud service for recognition and translation. There is no microphone listening to the speakers, and no separate integration for each player. The captions appear in their own window and are not written into the video file. Immersive Mode hides the frame, leaving only the captions.

The clip below uses English dialogue from Sintel to show Chinese captions and Immersive Mode. It uses mimi’s real frontend, replaying one cloud service’s subtitle responses at their original receive times. The recording harness substitutes for native audio capture and shortcuts. It shows how captions appear, but does not measure audio capture or live latency in an installed app.

Sintel © Blender Foundation (opens in a new tab), CC BY 3.0 (opens in a new tab).

Why words can change after they appear

One of the most annoying things about live captions is having the words change halfway through reading them. But wait until someone has said a whole lot, and the dialogue is over before the captions appear. The scene has already moved on. Captions that don’t match the picture are painful to watch. I’m still working on this in mimi.

With an existing subtitle file, a player knows the whole line in advance and shows it at the right video time. mimi must hear the audio first. A partial recognition result is provisional; more audio can revise words, punctuation, or boundaries. A final result ends revision for that audio segment. It guarantees neither correct words nor a complete thought. Google’s streaming recognition documentation (opens in a new tab) treats the likelihood of an interim result changing as a separate measure of stability.

Translation adds another problem. “I think this idea…” could end with “is great” or “won’t work.” Even a correctly recognized opening leaves the opinion unknown. Languages also arrange words differently, so adding context can require rewriting a translation’s earlier words. Research on online translation (opens in a new tab) shows this happening without changing the original prefix. The paper discusses hiding an unstable ending or keeping a new translation closer to the previous version. The former adds delay; the latter can preserve an unsuitable translation.

An example of missing context, not a recognition or latency measurement.

mimi replaces these drafts within the current caption block instead of adding every revision as a new sentence. Confirmed transcript and translation pairs move into the captions above. That keeps the opening from being printed over and over, but the provisional translation I’m reading can still change.

When to start translating

The services need different handling. In mimi’s low-latency mode, the service returns both the transcript and translation. They can arrive separately, so mimi pairs them using the service’s utterance identifiers. A late translation for the previous sentence must not be attached to the next transcript.

Another route uses Audio 3.0 for recognition, then sends the text to Qwen-MT for translation. mimi waits briefly for a draft to stop changing and tries translating the completed portion up to a sentence-ending punctuation mark. Continuous speech has a maximum-wait trigger too: once the draft is long enough, mimi can request a preview of the current text without waiting for a full stop. Punctuation and these timers decide when to try a translation. They do not turn the draft into a confirmed result.

When the recognizer confirms a segment, mimi translates the confirmed text in order. Only the latest preview request is kept; older requests are cancelled, and late results cannot overwrite a newer preview. Confirmed translations take priority over previews of the next sentence. This preserves order, but can make the next preview wait longer. If translations build up too far, this route reports an error and reconnects rather than accumulating dialogue that falls further and further behind.

The OpenAI route works differently again: the original and translation arrive as separate streams of appended text. mimi uses returned timing information and sentence endings on both sides to assemble corresponding portions into confirmed captions. One original sentence may become two translated sentences, so pairing the first sentence on each side is not enough. Without enough information to match the portions, the text has to stay in the current caption and wait.

Shorter segments give captions a chance to appear earlier, but the meaning may be incomplete and the translation may need revising. Longer segments give the translator more context, but the character may have finished the action before the caption appears.

By the time the translation appears, the scene has changed. Sometimes even the person on screen has changed. I’m still trying to remember what was being said in the last scene. lol.

How to display the text as it arrives

Once the translation arrives, I still need to be able to read it. If the interface changes with every revision from the service, I get halfway through a line and have to go back to the start. A burst of words, or a line pushed away just as I begin reading it, is difficult too. This live-subtitling study (opens in a new tab) compares word-by-word, block, and scrolling displays. The pace of text generation and the pace of reading need separate consideration.

mimi currently holds successive revisions briefly and updates the caption after the text stops changing for a moment. If new words keep arriving, it periodically takes the latest version rather than waiting for the stream to stop altogether. That can interrupt reading less often, but every extra wait also makes the caption later. I’m still adjusting this.

Assuming these updates arrive close together, the lower row holds them briefly and displays the latest. This is an illustration, not a recorded mimi caption stream.
Current display timing details

Draft originals appear after roughly 180 milliseconds without a change, and translations after roughly 400 milliseconds. Under continuous updates, the interface takes the latest version after at most roughly 750 milliseconds and 1.5 seconds respectively. Confirmed results and clearing update immediately. These timers start when the interface receives text; audio capture, network, recognition, and translation time come before that.

For a long current sentence, the window keeps only the last few lines visible and lets earlier text roll off the top. When text arrives or a draft becomes a confirmed caption, the layout also tries to avoid a sudden change in height or replaying the entrance animation. Recent work addresses scrolling during continuous updates: the next movement should continue from the text’s current position, without interrupting an unfinished glide and making it jump again.

Smoother scrolling still leaves the translation free to change. I want captions to appear early without interrupting me halfway through reading them. That part is still unfinished.

The transcript and translation can also arrive separately. In translation-only mode, mimi waits for the translation rather than temporarily substituting a Japanese line because the original arrived first. To see what has been recognized, I can select original or bilingual mode.

Why it currently uses cloud services

When I chose not to make mimi run a large local model, my main concern was accuracy. If it mishears the dialogue, a fluent translation won’t rescue it. It also needs to respond quickly, so the quality of a completed paragraph is only part of the problem.

Installation matters too. Users aren’t necessarily programmers, and even programmers can face hurdles downloading models, setting up a runtime, and dealing with versions and hardware compatibility. People don’t all have powerful computers. I can’t plan everyone’s experience around my own machine. Model size, memory use, and whether recognition and translation can keep up while running together all affect whether someone can use it.

But those were my considerations. Saying “local models aren’t accurate” was too sweeping. There are plenty of models I haven’t even tested. That was a pretty reckless claim.

Some people will accept lower accuracy in exchange for a local setup without API charges. If all audio and text really stay on the machine, they also have more control over privacy. A suitable model on suitable hardware can be faster too. Local does not automatically mean slow, and no API charges does not mean no memory use, electricity, or maintenance. People accept different costs. I was being a bit narrow-minded about that.

For now, mimi still leaves model execution to cloud providers and handles audio, captions, and display itself. That involves network and processing waits, credentials, and service charges.

It hears the sound, but cannot see the picture

System audio may contain music, effects, and dialogue from another page. Closing unrelated playback is more direct than expecting the recognizer to know which voice I want. If someone points at the screen and says “this one,” mimi cannot tell what they mean either.

An example of visual context missing from audio, not an app recording or recognition result.

An existing, checked subtitle track can be used directly. When one is unavailable, mimi can provide the dialogue. An odd line still needs checking against the original audio and scene.

Setup requires service credentials; languages, modes, and charges depend on the provider. System audio goes to that provider for processing. mimi does not use the microphone or record the screen. Saving captions and recording system audio are off by default. I can enable them in settings to keep local records and export TXT or WAV files. Caption timestamps mark when text was confirmed, rather than the video’s playback position, so the exports are not an aligned subtitle track for the original video.

There are now packages for Apple silicon and Intel Macs running macOS 13 or later, plus Windows x64 and Linux x86_64 previews. Audio capture on a physical Intel Mac still needs verification; Linux window behavior also depends on the desktop environment. Downloads and platform details are in the mimi repository (opens in a new tab).

Comments