Two fighters. Two teams. Everyone writes the moves. A live video model turns those ideas into a shared spectacle.
Watch what happened when a crowd’s attack ideas became real generated fights. The good exchanges, the awkward pauses, and the experiments trying to close the gap.
Generation prompts Read & copy +MEOWTOR / CHOMP Real Director output · baseline prompt study480p 10.125 sec
The APIs behind the fight
APIs let the game ask other services to do specific jobs. These are the main services used in the gameplay experiments:
Google Gemini — plan the action. It selects player moves, turns them into scene instructions, and checks images of the finish. It also makes fighter artwork. The game’s own rules decide damage and the winner.
fal’s H3 Max Director — make the fight video. It turns the scene instructions into moving pictures and action sounds. The game sends new instructions as the fight continues.
ElevenLabs — add voice and sound. Its speech API makes announcer audio. Its sound-effects API was used for crowd and transition sounds. These are separate from the sounds that Director makes with the video.
Cloudflare Stream — send the live picture to viewers. It carries the broadcast from the host to spectators. Cloudflare R2 stores saved media, D1 stores game records, and Workers runs the game’s server.
Each service can finish its job at a different moment. That is why these experiments test both the quality of the fight and the wait before players can see it.
What happens between one attack and the next?
What we tried
Experiments run from earliest to latest. Colors and labels show what each experiment tests. Each exchange is a short set of attacks and counters; a generation prompt is the instruction sent to the video AI.
What we learned
SEPTEMBER 8, 2026 / DIRECTOR 1.1 TRIALS
More control did not always make a better fight.
Eleven bounded trials tested end images, longer sessions, supplied audio and higher resolutions. The recommendation from this round: keep SlopFight’s current generation defaults.
These were real video-service trials using a controlled script, not a full Studio app match. They compare features of the same current endpoint; they do not establish an old-model versus new-model improvement. No production settings or player records were changed by these tests.
1. An ending image can erase the move that gets you there.
Three pairs repeated the same LASERPAW/VOLTJAW finisher with seeds 111, 222 and 333. Each pair used the same starting art, prompt and settings; one added a target ending image.
All three text-only controls retained a yarn sphere and rainbow trail, although literally riding the sphere remained ambiguous. All three end-image trials skipped that setup. Both approaches showed the shark falling and staying down. The target improved a particular grooming pose, not whether the intended winner won.
Seed
Text only: first frame
End image: first frame
111
13.166 s
17.697 s
222
11.973 s
22.124 s
333
9.454 s
17.346 s
Mean
11.531 s
19.056 s
What we learned: the end-image runs started 7.525 seconds slower on average. This is three repeats of one matchup, with a fixed run order and provider variability—not a general latency benchmark. Seed 222’s target run reached the 40-second time limit. Keep end images off by default; test them selectively when the staging benefit is worth the delay.
2. A longer connection still needs recovery.
One session lasted 120.978 seconds, showing 112.030 seconds of video. It still had a 0.416-second gap between decoded frames. A separate short control had a 5.200-second frame gap. Reported session allowances changed during testing, so a fixed maximum timer would be a poor substitute for reading the live allowance.
The longer session also ran a scripted practice match. Moves appeared roughly 6–11 seconds after submission, even when the service accepted the prompt earlier. Character proportions and staging drifted, and unwanted text-like markings appeared.
What we learned: retain session recovery. These timings exclude the game’s judge and choreographer, spectators, balances, commentary mixing and finish verification. A complete Studio match and maximum-session replacement test remain unverified.
3. Supplying sound worked. Running out produced silence.
The recording contained the supplied 110 Hz sound, then the queued 880 Hz sound, then the 330 Hz replacement. Replacing audio did not immediately interrupt material already dispatched. Exhaustion produced very quiet or silent samples; an invalid audio source was rejected without ending the video session.
Default audio and explicit 96, 128 and 192 kbps settings all produced recordings. The test used synthetic tones to make source identity measurable. These were not production effects, and no supplied impact was shown to line up with a game-authorized hit.
What we learned: keep commentary on its separate track. Signal measurements and model-assisted audio reviews do not establish which setting sounds best; direct human listening is still needed. The saved files were re-encoded, so their bitrate does not measure the incoming audio bitrate.
4. More pixels worked, but did not fix the fight.
Requested
Recorded dimensions
First frame
First 3 generation times
480p
832 × 480
8.948 s
1.811 / 1.748 / 1.764 s
768p
1344 × 768
13.005 s
5.921 / 4.985 / 4.975 s
1080p
1920 × 1080
13.598 s
6.995 / 6.835 / 6.846 s
Each short live sample reported zero dropped frames, but higher resolutions took longer to generate. A separate ten-second offline re-recording produced 2.66 / 2.78 / 5.07 MB at 480p / 768p / 1080p. The 1080p encoder used a higher bitrate, so this is not an isolated resolution comparison. A late 1080p close-up also hid the defeated fighter.
What we learned: retain 480p as the default. Trial 768p selectively; higher detail alone did not improve move fidelity or finish framing. These tests did not validate the live spectator path.
The next hypothesis: finish with both faces visible.
A local prompt change made at asks for both full faces to remain recognizable through the final second and actual last frame. It preserves masks, nonhuman anatomy and the resolved result: a defeated fighter stays down, and the action cannot be shortened or reset just to reveal a face.
Not yet established: whether this reduces distortion in the next scene. The eleven recordings below used the earlier prompts. Thirty prompt checks passed after this change, but no new video trial or production deployment was performed for it.
Decision at the end of this round: keep 480p, text-directed finishers, separate commentary, audio gating, finish verification and session recovery. Add capability/audio telemetry after integration review. The original report records 52 targeted tests and three offline session-cleanup checks; those are not full-game acceptance tests. Estimated Director usage cost was $14.4–$14.9, excluding model-assisted review charges—not an invoice or confirmed charge.
Watch the eleven trials
Actual recordings from the existing test run, presented as MP4 viewing transcodes with audio. No trials were rerun for this post. The long recording contains synthetic audio tests and a scripted practice match, rather than a complete multiplayer game.
Finisher · seed 111 · text only
Approx. 15:35:15–15:35:47 PDT · 32.200 s measured start to stop
Finisher · seed 111 · end image
Approx. 15:36:06–15:36:42 PDT · 36.731 s measured start to stop
Finisher · seed 222 · text only
Approx. 15:37:02–15:37:33 PDT · 30.981 s measured start to stop
Finisher · seed 222 · end image
Approx. 15:38:12–15:38:52 PDT · 40.029 s measured start to stop
Finisher · seed 333 · text only
Approx. 15:39:17–15:39:45 PDT · 28.487 s measured start to stop
Finisher · seed 333 · end image
Approx. 15:40:22–15:40:58 PDT · 36.364 s measured start to stop
480p / 128 kbps · audio and scripted practice match
Approx. 15:41:15–15:43:16 PDT · 120.978 s measured start to stop
768p / 128 kbps
Approx. 15:43:44–15:44:16 PDT · 32.048 s measured start to stop
1080p / 128 kbps
Approx. 15:44:48–15:45:21 PDT · 32.616 s measured start to stop
480p / 96 kbps
Approx. 15:45:49–15:46:17 PDT · 28.308 s measured start to stop
480p / 192 kbps
Approx. 15:47:05–15:47:33 PDT · 28.056 s measured start to stop
Experiment timestamps
All times below are approximate PDT (UTC−07:00), September 8, 2026. The original harness recorded elapsed time, not wall-clock event times. Start and first-frame times were reconstructed from the original recording-save timestamp minus measured elapsed durations. Unmeasured recorder finalization, upload and save time shifts these estimates later. The elapsed durations remain the original measurements.