·7 min read

Animate your album cover for Spotify Canvas with any AI: prompts for Veo, Kling, Higgsfield & Seedance

You can turn a static cover into a Spotify Canvas with almost any video AI — Veo 3, Kling 2, Higgsfield, Seedance, Runway. Here's how prompts shape the result, and copy-paste prompts you can use today.

Spotify Canvas is just a short 9:16 video, 3–8 seconds long, looping behind your track on mobile. That means you don't strictly need a dedicated tool — any modern image-to-video AI can technically produce one. Veo 3, Kling 2.x, Higgsfield, Seedance, Runway Gen-4, Luma Ray, Pika — they'll all take your cover and animate it.

The catch: results vary wildly. Two things decide what you actually get back — the model you pick, and the prompt you write. This guide walks through both, and gives you copy-paste prompts tuned for each major model so you can try them yourself.

Why the model matters more than you think

Every video model has a personality. Veo 3 (Google) leans cinematic and physically plausible — great for realistic photography covers, less great for illustrated artwork it tends to "correct" into realism. Kling 2.x is currently the strongest at preserving the input image's style, which makes it a safe default for illustrated or painted covers. Higgsfield specializes in camera-movement presets (push-in, orbit, parallax) and is excellent when you want a controlled cinematic move rather than subject animation. Seedance (ByteDance) is fast, cheap, and surprisingly good with stylized 2D, though it can drift on faces. Runway Gen-4 is the all-rounder and handles text and logos better than most.

Pick the model that matches your cover. Photographic, moody portrait? Veo or Runway. Hand-drawn or painted? Kling or Seedance. Want a controlled camera move on a still scene? Higgsfield.

Why your result will look different every time

Image-to-video is non-deterministic. The same prompt on the same model with the same image will produce a different clip every run — sometimes subtly, sometimes dramatically. Hands warp. Faces shift. Backgrounds invent details that weren't there. This isn't a bug, it's how diffusion video works in 2026.

Plan to generate 3–5 takes for any cover and pick the best one. Budget for it: a single Veo 3 generation can cost a few dollars; Kling and Seedance are cheaper but still add up. If a generation looks 90% right but has one broken frame, try again — don't try to fix it in post.

Prompt anatomy that works on every model

Regardless of model, a Canvas prompt has four parts: subject motion (what moves), speed (how fast), camera (does it move too?), and style lock (what must NOT change). The style lock is the part most people skip and it's why their cover ends up looking like a different image.

Write your prompts in English, even if you and your audience aren't English speakers. Every major video model (Veo, Kling, Higgsfield, Seedance, Runway, Luma, Pika) is trained predominantly on English captions and follows English instructions more accurately — non-English prompts get translated internally and lose nuance, especially around camera moves and style locks. Keep the prompt in English; the output is just pixels, so your listeners won't see a difference.

Template: "[Subject] [does motion] [slowly/gently]. [Camera instruction]. Preserve original art style, colors, composition, and character design. Loop. Vertical 9:16."

Copy-paste prompts by model

Below are prompts tuned to each model's quirks. Replace the bracketed parts with what's actually on your cover. Run each one 2–3 times and keep the best take.

Veo 3 (Google, via Gemini / Vertex / Flow)

Veo wants cinematic, descriptive language and responds well to camera direction. Keep motion subtle or it will over-animate.

"Cinematic subtle motion on a still album cover. [Subject, e.g. the character in the foreground] [action, e.g. breathes slowly, hair drifting in a faint breeze]. Background [action, e.g. soft volumetric light shifts, dust particles float]. Camera: extremely slow push-in, almost imperceptible. Maintain the exact art style, color palette, and composition of the source image. No new objects, no style change. 5 seconds, vertical 9:16."

Kling 2.x (Kuaishou)

Kling is the most style-faithful of the major models, which is why it's the default behind a lot of animated illustration work right now. Keep prompts short and concrete — it over-thinks long ones.

"[Subject] [gentle action, e.g. blinks slowly and the smoke behind them drifts upward]. Subtle, dreamy motion. Keep original illustration style, line work, and colors unchanged. Static camera. Loop. 5s, 9:16."

Higgsfield

Higgsfield's strength is its camera-motion presets — pick one rather than describing the camera in words. Pair a preset like "Slow Push In" or "Parallax" with a tiny note about subject motion.

"Preserve the album cover exactly. [Subject] [tiny motion, e.g. the neon sign flickers once, the water in the foreground ripples gently]. Keep all characters, typography, and composition identical to the source. Loopable, 9:16, 5 seconds." Then select a camera preset (Slow Push In or Parallax usually wins for Canvas).

Seedance (ByteDance / Doubao)

Seedance is fast and cheap, great for testing prompt ideas. It handles 2D and anime-style art particularly well. It can drift on faces, so prefer prompts that animate environment over face when your cover has a portrait.

"Animate this cover with calm, looping motion. [Environment action, e.g. clouds drift slowly across the sky, light flickers on the water]. [Subject action, e.g. the figure's coat sways gently]. Do not change the art style, colors, or character. 9:16 vertical, 5 seconds."

Runway Gen-4 / Luma Ray / Pika

All three behave similarly enough that one prompt template covers them. They're strong all-rounders and handle text on covers better than most.

"Subtle living-photo motion on this album cover. [Subject] [small action]. Background [small action]. Preserve every element of the original image — style, color, characters, typography. No camera movement, or extremely slow drift only. vertical 9:16, 1080×1920."

After you generate: format for Spotify

Whatever model you use, Spotify Canvas requires an MP4, exactly 1080×1920, 3–8 seconds, looping, under ~8MB. Most general video models export at 720×1280 or 1080×1920 but not always at the right duration or with a clean loop. You'll likely need to trim, scale, and re-encode in something like Handbrake, CapCut, or ffmpeg before uploading to Spotify for Artists.

Or just skip all of this: animatecover.com

Everything above works, and we genuinely recommend trying it — there's something fun about wrangling a prompt until it gives you the loop you imagined. But if you just want a Canvas that's ready to upload, with no model-picking, no aspect-ratio fixing, no trial-and-error budget that quietly turns into $40, that's exactly what AnimateCover is built for.

AnimateCover is purpose-built for Spotify Canvas: it picks the right model for your cover automatically, locks the output to 1080×1920, ensures a, and exports a Spotify-ready MP4. A typical cover costs a fraction of a single Veo generation, and you can re-roll cheaply until you love the result. Upload your cover, write one short line about how it should move, and you'll have your Canvas in a couple of minutes.