AI Router · CLI · MCPCheapest eligible quotes before you create
how-to · activation

Write motion prompts for image-to-video generation

Describe camera and environmental movement without needlessly recreating the approved scene.

Start with 20 free credits
Tool rack

Separate the source image from the motion instruction

Treat an image-to-video request as two inputs with different jobs. The still supplies the visible starting composition; the text supplies the change over time. Google's current image-to-video guidance says the source already provides subject, scene, and style, and recommends concentrating the prompt on camera movement, subject animation, and environmental change. Runway gives the same division of labor more explicitly: the input image acts as the first frame and carries composition, subject matter, lighting, and style, while the text describes motion, camera work, and temporal progression.

Before drafting, inspect the still for motion that it already implies. A runner leaning forward, wind-blown clothing, motion blur, dust, or diagonal speed lines can pull the generated movement in a direction that conflicts with the words. Runway's guide warns that contradictory visual cues can require more iteration and recommends changing the input image when repeated prompt edits do not resolve the conflict. This is why a motion prompt begins with an image check, not with a longer description of everything already visible.

For an OfflineCreator MCP run, `generate` reserves an image-to-video job when an upload is required, and `upload_input` then accepts the source through a local `filePath` or `imageBase64` field before submitting it. The published tool schema accepts PNG, JPEG, WebP, or GIF content types and `16:9`, `9:16`, or `1:1`; it does not publish a duration field. Keep provider-only controls out of the MCP call unless the current tool schema exposes them.

Credit gauge

Confirm the current model and spend before testing a prompt

OfflineCreator Studio's current public model catalog lists Kling 2.6 Pro Motion as an image-to-video model for image animation at 42 credits. Treat that as a current catalog snapshot, not a permanent price, availability promise, or quality ranking. Call `list_models` and `get_credits` immediately before a paid attempt, show the operator the returned model and cost, and set a one-generation approval boundary.

The underlying fal Kling 2.6 Pro image-to-video endpoint documents a different, provider-level input contract: `prompt` and `start_image_url` are required, `duration` can be five or ten seconds with five as the default, and `end_image_url` is optional. Its example response contains a generated MP4 file record. Those details explain the provider envelope but do not add a duration or image-URL argument to OfflineCreator MCP; the public MCP path reserves first, then uploads through `upload_input`.

Prompt anatomy

Build one readable motion sentence

Use a compact order: camera behavior, subject action, environmental action, then endpoint. Name direction and pace instead of stacking cinematic adjectives. A practical draft for a still of a ceramic bird on a windowsill is: “Locked camera. The ceramic bird remains still while the sheer curtain moves gently inward once and soft leaf shadows travel slowly across the wall; motion settles back to near-stillness for the final second.” The still already defines the bird, window, framing, palette, and light, so the prompt does not recreate them.

When the camera should move, state one path and one endpoint: “The camera makes a slow straight push-in toward the subject as steam rises in one thin curl; stop on a stable medium close-up.” Do not combine push-in, orbit, pan, zoom, rack focus, and a transformation in the same short request. Google's guide identifies physical camera movement, subject animation, and environmental animation as distinct motion types; choosing only what the shot needs makes the intended result easier to review.

Write constraints as review targets, not guarantees. “No cut,” “locked camera,” or “minimal subject motion” communicates intent, but the output still needs inspection. Runway specifically suggests locked-off or perfectly still camera language for minimal-motion attempts while also noting that video models are designed to create motion. A constraint can make a failure legible; it cannot prove that a model will obey.

Camera
Choose locked, push in, pull back, pan, truck, or orbitAdd direction, pace, and a stopping composition; use one dominant path.
Subject
Name one observable actionPrefer a blink, turn, step, or controlled object movement over a chain of events.
Environment
Animate one supporting elementWind, fog, steam, water, shadows, or fabric should support rather than compete with the subject.
Endpoint
Describe where motion settlesA stable final beat gives the reviewer a clear target and the editor a usable exit.
Output contact sheet

Use matched documentation examples as evidence, not benchmarks

Current first-party guides publish matched input, prompt, and result presentations. Google's image-to-video best-practices page pairs a source image with a resulting video and the prompt “A sweeping drone-like aerial view starting from ground level and rising to reveal the entire landscape in epic proportions.” Runway's image-to-video guide presents input-image, prompt, and result rows, including a controlled comparison where prevalent motion cues such as blur and dust are reduced while the requested prompt remains about a parked, motionless car and an aggressive camera arc.

These examples support two narrow conclusions: motion language is evaluated together with the first frame, and changing the first frame can be a valid revision when its visual cues fight the prompt. They are not independent benchmarks, customer results, or evidence that a different still, model, ratio, or provider will reproduce the shown movement. The embedded result media also does not substitute for a project-specific generation record.

For this guide's ceramic-bird example, the matched review record is therefore still prospective: save the exact source still, prompt, model record, generation ID, unedited output, and a frame-by-frame decision. Record whether the camera stayed locked, whether only the curtain and shadows moved, and whether the final second settled. Until those artifacts exist, label the text an editorial prompt example rather than a tested output.

Google example
Source image + aerial-rise prompt + displayed resultFirst-party demonstration of a camera-motion instruction paired with image-to-video media; no cross-model performance inference.
Runway comparison
Same motion request, different strength of implied-motion cuesFirst-party matched example showing why the source image belongs in prompt diagnosis.
This page
Ceramic-bird prompt has no generated result in this research passA later review needs the unedited clip and explicit pass or rejection notes before it becomes a matched project example.
Workflow timeline

Revise the input that caused the earliest visible failure

Run the workflow in a fixed order: approve the still and intended movement, check `list_models` and `get_credits`, reserve with `generate`, upload through `upload_input`, wait for completion, download the unedited result, and compare it with the written motion contract. The public MCP implementation returns an upload-required reservation without polling, while `upload_input` uploads and submits the job. The public tools also expose `wait_generation` for polling and `download_output` for a short-lived signed URL after completion.

Classify the earliest visible mismatch before changing anything. If the camera moves despite a locked request, revise the camera sentence. If the ceramic bird moves before the curtain does, clarify the subject constraint. If the source pose or blur implies a conflicting action, revise the still instead of adding more negative phrases. If the desired shot contains several beats, split the concept; Google's current guidance recommends one focused scene for a short video rather than chaining distinct events.

Change one input at a time and retain everything else: source-image revision, prompt revision, model record, ratio, generation ID, output, first failing timestamp, and decision. This creates a diagnosable pair. It does not establish a universal prompting rule, because one accepted or rejected generation cannot measure reliability across models or source images.

Related circuit

Move to the drift-review guide when the output changes identity, geometry, text, materials, or background details that the source was meant to anchor. Use the storyboard-frame guide when the idea needs several distinct moments rather than one continuous motion instruction. Return to the creative-workflow directory when the task changes from writing the prompt to selecting, generating, retrieving, or approving another asset.

Canonical plate

Keep the prompt guide inside its evidence boundary

This page owns the narrow question of how to write motion instructions around an existing first frame. It does not own general image-to-video setup, product-authenticity review, multi-shot storyboarding, model comparison, or a promise that negative instructions preserve identity, text, geometry, lighting, or exact camera paths.

The first-party matched examples show documented inputs beside displayed results, but this research pass did not generate a clip or capture a project-specific source/prompt/output set. The current community retrieval timed out before emitting an evidence report or review queue, so it cannot establish either useful discussion or community silence. Keep claims about reliability, quality, creator preference, editing effort, speed, and customer outcomes out of the page until attributable evidence and task-matched artifacts support them.