Two Ways I Tried Making AI Video on AutoDL
I had used the MiniMax H3 video API to make a Gwen transformation video, gradually changing a game character from one skin to another. My notes put that round at RMB 35, including tests and the finished video. Trying more ideas meant spending more.
For the August 9 experiment, I added RMB 10 to AutoDL, rented a machine, installed the environment, downloaded the model, and set up a ComfyUI workflow.
Getting it running first
ComfyUI connects generation steps as nodes. The canvas shows where images enter, which model is used, and where the result comes out. A workflow file (opens in a new tab) saves those connections and settings. Borrowing one doesn’t install its models and plugins too.
As a first-time user, I spent a while on the setup. If a borrowed workflow can’t find its model, relies on a plugin that isn’t installed, or exceeds the machine’s GPU memory, its nodes need checking. Before I could generate a clip, I’d already waited for downloads and worked on the environment.
The few clips I got looked much like those I’d made through the H3 API. That was an impression from those clips; I hadn’t held the images and settings constant for a comparison. Running it myself didn’t remove the retries, either. Change a few words, wait for another clip, dislike it, and try again.
RMB 10 was a top-up, not the cost of a video. While the rented machine is running, downloads, setup, and discarded clips all take time. Those belong in the calculation too.
Being able to change the workflow was useful. For an occasional short video, the preparation took a while.
August 22: trying hosted workflows
I later tried H3 workflows on AutoDL.Art. This time I chose a ready-made workflow and submitted a job. The platform supplied the models and environment, without my renting and setting up a machine first.
The historical page advertised RMB 0.01 per second for 480P and 768P. That is a price label from August 22, not a current quote. The call log has separate columns for queue time, inference time and charge.
| Created on August 22 | Workflow | Charge | Queue time | Inference time |
|---|---|---|---|---|
| 01:06 | Multiple images and audio | RMB 0.05 | 36s | 1m 53s |
| 01:08 | Multiple image references | RMB 0.10 | 37s | 3m 9s |
| 13:41 | Multiple images and audio | RMB 0.15 | 1m 2s | 8m 3s |
These minutes describe how long the jobs spent running inference, not the length of their output videos. The workflow name included “15 seconds,” but the screenshot does not show request parameters or output metadata. Neither that name nor the charge proves the actual clip length.
The hosted workflow removed the environment setup. Queueing and inference still took time. These three jobs used different workflows and inputs, so their timings do not establish a speed comparison.
Comments