SHIPPED · v0.10.0
Bring your own generation APIs, with a spend cap and multilingual narration
New
Use your own generation APIs in a directed video. Turn on Generation APIs, choose a provider for voice, stills and motion, set a spend cap, and let the harness generate the media alongside its local render. Runway, fal.ai and Gemini are available for pictures and motion; ElevenLabs handles narration.
Choose providers in the prompt. Requests such as “Use Gemini for images and Runway for motion” or “Use fal PixVerse for clips” select the matching lanes while the API switch is on. English and Vietnamese requests work, and explicit instructions such as “do not use fal” turn that lane off.
More hosted models. fal.ai adds FLUX Schnell stills plus Wan 2.5 and PixVerse 4.5 motion. Runway uses Gen-4.5 for text-to-video and offers a text-capable image default. Gemini adds Veo 3.1 Fast alongside Gemini image generation.
Automatic multilingual narration. ElevenLabs now chooses a compatible model and account voice from the language and delivery requested in the prompt. Vietnamese narration uses a supported Vietnamese voice, and a finished video can be re-voiced from the chat without rebuilding its pictures.
Richer original artwork. The drawing system adds a much larger library of people, animals, clothing, rooms, buildings, vehicles, food, devices, documents and everyday props, with better posing, expressions, acting, camera movement and ambient motion.
Reference boards for every style. Style packs now include boards for people, animals, props and sets, giving the director a concrete visual target before it draws or reviews a scene.
Improved
Spend stays under your control. The composer shows estimated use and available provider balance when it can be read. Every hosted call is checked against the run cap, logged with its provider and model, and listed in the delivery notes.
Cloud failures have a useful fallback. A failed, exhausted or quota-limited lane can continue locally or stop at a review gate, according to the run setting. One provider failure no longer loses the rest of the production.
Better looking scenes by default. Layout, colour, scale, subject placement and visual variety are checked more strictly, with dedicated construction tools for characters, sets, props and shot composition.
Generated pictures become real scene assets. Stills and short motion plates are inserted into the reel with their prompt, model and provenance, reused when unchanged, and disclosed in the final delivery.
Fixed
Fixed fal.ai queue results failing for models with nested API paths by using the status and response URLs returned by fal.ai.
Fixed Gemini video quota and operation errors being mistaken for successful generations, and corrected supported video durations and reference-video budget estimates.
Fixed Vietnamese ElevenLabs requests selecting an incompatible model or silently falling back to an English voice.
Fixed hosted media choices being ignored by the video harness after they were selected in the composer or written in the prompt.