Putting video frames together for AI

Yuxino,•Making Things
中EN

KOMA is a video analysis tool I built. Part of its story comes from a company AI hackathon: we tried making something with AI and entered it in the competition. It was just the two of us, really—partners, I suppose, haha. I learned a lot.

At first, the two of us picked a topic fairly casually. Neither of us was sure whether video analysis would actually be useful. It sounded fun, so we gave it a go. And we got to watch videos at work as part of the experiment. Excellent.

Put the frames in one picture

At the time, having a vision-language model analyze a video directly meant a long wait. These models can process both images and text. Extracting frames and analyzing them one at a time was slow too, and cost money.

One approach we used during the competition was to combine sampled frames into a grid and send that picture to the model. Several separate images became one image it could examine together.

That also made debugging easier. We could put the model’s answer next to the frames it had seen and inspect them together. Repeated frames or a gap in the sequence were much easier to spot than when flipping through individual pictures.

An illustration of the method, not footage from the competition. The six-panel layout does not represent a fixed setting we used.

What if the order number flashes past?

Which frames to choose, and how large to keep them, depends on what we want to find.

Suppose a customer sends a recording of themselves using a system and asks us to look into a problem. They open an order’s details, the order number appears briefly, and then they move to another page. Along with seeing which button they clicked, we might need to identify the order.

If our samples skip that moment, the combined picture won’t contain the number. It was in the video, but the model never saw it. Asking again won’t help; we first need to return to the footage and find the moment when the details were open.

Taking more samples makes that less likely, but leaves us with more pictures. Squeeze them all onto one fixed-size canvas, and each panel gets smaller. The model might recognize the order page while being unable to read the number. Keeping that frame as a larger, separate image is more useful than adding still more frames to the grid.

Or we may have the right frame, with clear text, and the model still gets a digit wrong. When an answer is wrong, we need to check the picture: was the number missing, too small to read, or visible but misread? Changing the question alone won’t necessarily fix any of those problems.

Sampling at regular intervals spreads pictures across a video without guaranteeing that we catch the order details. Choosing frames based on visual changes might miss a number changing in one small area too. One approach worth trying is to locate the rough moment, then sample that section more closely. That adds processing and waiting.

A dog in human clothes confused the model

Models differed a lot too. While checking their analyses, I watched plenty of absurd, funny videos at work. Some had me laughing really hard.

I remember one where an owner had dressed a dog in human clothes, and the dog was running along the ground. Some models described it as a person pretending to be a dog.

The video was already funny, and the model added another misunderstanding. That answer alone doesn’t tell us exactly where it went wrong. What it does show is a plausible description that got the subject of the video wrong.

Today, KOMA (opens in a new tab) is a self-hostable website that organizes videos using sampled images and transcribed speech. It accepts uploaded files and downloadable MP4 URLs. That describes the current implementation, not a claim that all these features were finished during the competition.

Showing it to people revealed a real need

The competition was pretty intense. I remember a few hundred teams signing up. Some dropped out before finishing, though most completed their projects; I can’t say how much effort everyone put in. We finished with a pretty good placing. A ranking doesn’t tell the whole story, but I was happy about it for a while, and we also had a budget for food and drinks.

The competition introduced me to more colleagues from other departments. Video analysis had started as something I was casually playing with. It turned out that some colleagues actually needed it. Talking to them opened up possibilities for collaboration, and even for turning the work into something that could count toward KPIs.

Sometimes I wonder whether the whole world is winging it a little. Maybe I can be braver, take ownership, and try building something others haven’t tried. Something good might come out of it.

Friends, be a little braver and a little more outgoing. I used to think entering these competitions and going to these events was bullshit. I wasn’t going to win, so what was the point? But if all I do is think about it and never try, well, none of it will have anything to do with me.

After twenty or thirty years of being an introvert, here I am telling people to be more outgoing. I need to drop that old attitude. Who knows? Trying might at least get me a few free meals and drinks.

Comments