Selected Showcase

LightX2V8-step + Spark-H390% sparse

Prompts come from the VDN-H3 project page.

01 / 10
PromptShow full promptCollapse promptCinematic, close-up low angle shot, the camera shakes slightly. An on-screen young Japanese woman in her early 20s (S1), featuring striking, silky bright-white bob-cut hair with soft straight bangs, flawless pale snow-white porcelain skin, sharp black winged cat-eye eyeliner, and expressive dark hazel eyes, stands in the center foreground. She wears a casual nylon black Japanese streetwear jacket layered over a cotton white cropped top blouse, revealing a lean athletic figure. Behind her, the red steel structure of Tokyo Tower rises against a bright, slightly overexposed sky. The lighting is neutral daylight with deep focus and unedited candid realism. She holds a smartphone with her right arm extended in a selfie pose, gives a warm, natural smile to the camera, and says, <d>[Japanese] やっと東京に着いたよ...!</d> [Shot 2] At 00:01.000, the camera cuts to a medium tracking shot following her. Mid-stride into a moving crowd at Shibuya Crossing, the white-haired girl (S1) turns her head back toward the camera with a playful look, her white bob swaying. Neon signs in vivid colors are sharp and bright behind her under real night exposure, casting dynamic reflections on her face. [Shot 3] At 00:01.900, the shot transitions to a close-up static shot inside a convenience store. The girl (S1) holds a triangular wrapped onigiri up to the camera lens in the center frame, raising her eyebrows with a cute, questioning look. The lighting is a harsh fluorescent interior cool white light, revealing realistic skin texture on her pale face. [Shot 4] At 00:02.800, the camera cuts to a medium static shot. The girl (S1) is seated on a wooden bench in Yoyogi Park, looking away from the camera to the right. Dappled sunlight filters through green leaves in the background, casting soft warm patches on her white hair. A subtle lens smudge is visible on the edge of the frame, adding amateur camera realism. [Shot 5] At 00:03.700, the camera cuts to a close-up static shot. The girl (S1) is illuminated by warm practical lantern light from the left, giving her skin an amber hue. She takes a bite of a round takoyaki, her eyes widen instantly, and she pulls back laughing, waving her right hand near her mouth. She exclaims, <d>[Japanese] あつっ!でも、めっちゃ美味しい!</d> [Shot 6] At 00:05.500, the shot transitions to a medium static shot at night. The girl (S1) stands before a glowing, colorful vending machine. Mixed LED light from the machine reflects on her porcelain skin. She gives a slight, playful smirk directly to the camera. [Shot 7] At 00:06.400, the camera cuts to a medium static shot inside a train. Her beautiful face and white hair are reflected in the glass of the moving train window on the right side of the frame. Outside the window, city lights are blurred into long horizontal streaks. She looks out thoughtfully, her gaze directed away from the camera. [Shot 8] At 00:07.300, the shot transitions to a low angle static shot. The girl (S1) walks forward away from the camera, passing under a massive vermilion wooden torii shrine gate. The top of the gate fills the upper frame, lit by natural afternoon daylight. [Shot 9] At 00:08.200, the camera cuts to a close-up static shot. At a digital light exhibit, the girl (S1) looks up slowly in slow motion. Intricate blue, violet, and gold light patterns are projected directly onto her face and white hair. The focus shifts, making her face sharp while the background lights diffuse into bokeh. [Shot 10] At 00:10.000, the camera tilts up from a close-up on her face. The girl (S1) has a small, impressed smile. The upward tilt reveals a massive giant robot statue standing tall behind her against the sky. The wide-angle framing captures the scene under natural daylight. [Shot 11] At 00:11.000, the shot transitions to a medium tracking shot from behind in a lantern alley. The camera follows the girl (S1) as she walks down the narrow path. Warm glowing paper lanterns line the walls, casting dark shadows on the alley. She glances back over her right shoulder once with a warm smile. [Shot 12] At 00:12.000, the camera cuts to a close-up selfie shot that shakes slightly. The girl (S1) extends her arm holding the smartphone, revealing the stunning Tokyo skyline at night sprawling behind her. The wind gently moves her white hair under natural night exposure. She says, <d>[Japanese] こんな夜景、初めて見るな...</d> [Shot 13] At 00:13.000, the camera cuts to a medium shot that shakes slightly. The girl (S1) crouches down in Nara Deer Park, holding out a crispy round deer cracker in her right hand. A brown deer in front of her bows politely toward her, making her burst out laughing. Her white bob shifts as she giggles under warm natural lighting. overall_soundscape: The soundscape begins with bright urban room tone and audible street ambiance, followed by the chaotic murmur of a bustling crowd and sharp footsteps at the crossing. A sharp crinkle of plastic packaging echoes in the store, giving way to the soft rustle of leaves and a gentle breeze in the park. Loud, distinct laughter and a sharp exhalation dominate over the faint sizzle of hot street food, transitioning seamlessly into the rhythmic clatter of a moving train over tracks. Gentle footsteps crunch on gravel under the shrine gate, later replaced by a pronounced electronic hum at the light exhibit and the soft rustle of clothing as she walks down a quiet alley. A steady gust of wind hits the microphone on the rooftop, ending with a distinct deer snort and clear, bright giggles in the park. non_diegetic_music: A fast-paced, upbeat electronic pop track with heavy bass drops and a driving rhythmic tempo, featuring synthesized beats and a dynamic energy that maintains a steady pace throughout the clip.
integrated_multimodal_description: [Shot 1] Cinematic, close-up low angle shot, the camera shakes slightly. An on-screen young Japanese woman in her early 20s (S1), featuring striking, silky bright-white bob-cut hair with soft straight bangs, flawless pale snow-white porcelain skin, sharp black winged cat-eye eyeliner, and expressive dark hazel eyes, stands in the center foreground. She wears a casual nylon black Japanese streetwear jacket layered over a cotton white cropped top blouse, revealing a lean athletic figure. Behind her, the red steel structure of Tokyo Tower rises against a bright, slightly overexposed sky. The lighting is neutral daylight with deep focus and unedited candid realism. She holds a smartphone with her right arm extended in a selfie pose, gives a warm, natural smile to the camera, and says, <d>[Japanese] やっと東京に着いたよ...!</d> [Shot 2] At 00:01.000, the camera cuts to a medium tracking shot following her. Mid-stride into a moving crowd at Shibuya Crossing, the white-haired girl (S1) turns her head back toward the camera with a playful look, her white bob swaying. Neon signs in vivid colors are sharp and bright behind her under real night exposure, casting dynamic reflections on her face. [Shot 3] At 00:01.900, the shot transitions to a close-up static shot inside a convenience store. The girl (S1) holds a triangular wrapped onigiri up to the camera lens in the center frame, raising her eyebrows with a cute, questioning look. The lighting is a harsh fluorescent interior cool white light, revealing realistic skin texture on her pale face. [Shot 4] At 00:02.800, the camera cuts to a medium static shot. The girl (S1) is seated on a wooden bench in Yoyogi Park, looking away from the camera to the right. Dappled sunlight filters through green leaves in the background, casting soft warm patches on her white hair. A subtle lens smudge is visible on the edge of the frame, adding amateur camera realism. [Shot 5] At 00:03.700, the camera cuts to a close-up static shot. The girl (S1) is illuminated by warm practical lantern light from the left, giving her skin an amber hue. She takes a bite of a round takoyaki, her eyes widen instantly, and she pulls back laughing, waving her right hand near her mouth. She exclaims, <d>[Japanese] あつっ!でも、めっちゃ美味しい!</d> [Shot 6] At 00:05.500, the shot transitions to a medium static shot at night. The girl (S1) stands before a glowing, colorful vending machine. Mixed LED light from the machine reflects on her porcelain skin. She gives a slight, playful smirk directly to the camera. [Shot 7] At 00:06.400, the camera cuts to a medium static shot inside a train. Her beautiful face and white hair are reflected in the glass of the moving train window on the right side of the frame. Outside the window, city lights are blurred into long horizontal streaks. She looks out thoughtfully, her gaze directed away from the camera. [Shot 8] At 00:07.300, the shot transitions to a low angle static shot. The girl (S1) walks forward away from the camera, passing under a massive vermilion wooden torii shrine gate. The top of the gate fills the upper frame, lit by natural afternoon daylight. [Shot 9] At 00:08.200, the camera cuts to a close-up static shot. At a digital light exhibit, the girl (S1) looks up slowly in slow motion. Intricate blue, violet, and gold light patterns are projected directly onto her face and white hair. The focus shifts, making her face sharp while the background lights diffuse into bokeh. [Shot 10] At 00:10.000, the camera tilts up from a close-up on her face. The girl (S1) has a small, impressed smile. The upward tilt reveals a massive giant robot statue standing tall behind her against the sky. The wide-angle framing captures the scene under natural daylight. [Shot 11] At 00:11.000, the shot transitions to a medium tracking shot from behind in a lantern alley. The camera follows the girl (S1) as she walks down the narrow path. Warm glowing paper lanterns line the walls, casting dark shadows on the alley. She glances back over her right shoulder once with a warm smile. [Shot 12] At 00:12.000, the camera cuts to a close-up selfie shot that shakes slightly. The girl (S1) extends her arm holding the smartphone, revealing the stunning Tokyo skyline at night sprawling behind her. The wind gently moves her white hair under natural night exposure. She says, <d>[Japanese] こんな夜景、初めて見るな...</d> [Shot 13] At 00:13.000, the camera cuts to a medium shot that shakes slightly. The girl (S1) crouches down in Nara Deer Park, holding out a crispy round deer cracker in her right hand. A brown deer in front of her bows politely toward her, making her burst out laughing. Her white bob shifts as she giggles under warm natural lighting. overall_soundscape: The soundscape begins with bright urban room tone and audible street ambiance, followed by the chaotic murmur of a bustling crowd and sharp footsteps at the crossing. A sharp crinkle of plastic packaging echoes in the store, giving way to the soft rustle of leaves and a gentle breeze in the park. Loud, distinct laughter and a sharp exhalation dominate over the faint sizzle of hot street food, transitioning seamlessly into the rhythmic clatter of a moving train over tracks. Gentle footsteps crunch on gravel under the shrine gate, later replaced by a pronounced electronic hum at the light exhibit and the soft rustle of clothing as she walks down a quiet alley. A steady gust of wind hits the microphone on the rooftop, ending with a distinct deer snort and clear, bright giggles in the park. non_diegetic_music: A fast-paced, upbeat electronic pop track with heavy bass drops and a driving rhythmic tempo, featuring synthesized beats and a dynamic energy that maintains a steady pace throughout the clip.
PromptShow full promptCollapse prompt3D CG fantasy animation, cinematic photorealism, a wide shot frames ancient jungle ruins at dawn as the camera slowly pushes in toward a moss-covered stone bridge. Warm, volumetric sunbeams pierce through the dense, vibrant green tropical foliage, casting long shadows. Crystal-clear water flows smoothly beneath the bridge, while a thick, luminous mist drifts steadily through the serene environment. [Shot 2] At 00:01.916, the camera cuts to a medium close-up of a young female elf with pointed ears and long, flowing blue hair, wearing an emerald-green linen tunic and leather bracers. Crouched on the mossy bridge, she extends her right hand toward the stream below. A glowing, translucent fox formed entirely of flowing water and light begins to emerge from the surface beside her, sending sparkling water particles into the air. The camera shakes slightly, adding subtle intimacy to the framing. [Shot 3] At 00:03.833, the shot transitions to a closer medium close-up as the camera slowly pushes in. The fully materialized water-fox from Shot 2 playfully circles the elf's extended hand, nuzzling her fingers with its fluid snout. The elf smiles softly, her shoulders relaxing. Magical, luminescent ripples spread outward across the stream's surface below them, reflecting the teal light. [Shot 4] At 00:05.750, the camera cuts to a medium shot, executing a smooth tracking shot to the right as the elf and the water-fox move together across the ancient stone bridge. The fox bounds gracefully between mossy stones, leaving a vibrant trail of glowing, suspended water droplets in its wake. Dense jungle foliage heavily frames the foreground and background of the stone path. [Shot 5] At 00:07.666, the scene shifts to a medium wide shot where the camera performs a dynamic arc shot around the pair. The elf sweeps her arms upward in a graceful gesture. Luminous, fluid streams of teal water rise from the stream, spiraling like living ribbons around her and the magical companion. The water-fox darts joyfully through the swirling aquatic ribbons, its body blending seamlessly with the magical currents. [Shot 6] At 00:09.583, the camera cuts to a medium shot and pans right to capture their fast but smooth motion. The water-fox leaps energetically through floating rings of glowing water, sending sparkling droplets splashing playfully around the elf. She throws her head back in a visual laugh, reaching both hands toward the creature as the suspended magical droplets catch the warm, volumetric sunlight. [Shot 7] At 00:11.500, the camera cuts to a cinematic wide shot and gently pulls out. The elf and the water-fox sit side-by-side at the far edge of the stone bridge, their backs partially turned as they overlook the glowing jungle stream. The fox slowly settles its fluid body onto the stone beside her hip. Warm yellow sunlight breaks through the dense canopy above, illuminating the drifting mist and suspended water particles as the environment fades into a soft, hazy light. overall_soundscape: A continuous, loud flow of rushing water dominates the environment, accompanied by the distinct rustle of leaves in the breeze. As the fox materializes, pronounced, resonant magical chimes and clearly heard liquid splashes ring out. Audible, rhythmic splashing and the fluid sloshing of water follow as the creature bounds across the stones, mingling with distinct, glassy tinkling sounds of water droplets popping in the air and a faint but clearly audible creature trill. non_diegetic_music: An orchestral fantasy score plays at a slow tempo, featuring continuous string melodies, sustained low cellos, and light woodwind flourishes.
integrated_multimodal_description: [Shot 1] 3D CG fantasy animation, cinematic photorealism, a wide shot frames ancient jungle ruins at dawn as the camera slowly pushes in toward a moss-covered stone bridge. Warm, volumetric sunbeams pierce through the dense, vibrant green tropical foliage, casting long shadows. Crystal-clear water flows smoothly beneath the bridge, while a thick, luminous mist drifts steadily through the serene environment. [Shot 2] At 00:01.916, the camera cuts to a medium close-up of a young female elf with pointed ears and long, flowing blue hair, wearing an emerald-green linen tunic and leather bracers. Crouched on the mossy bridge, she extends her right hand toward the stream below. A glowing, translucent fox formed entirely of flowing water and light begins to emerge from the surface beside her, sending sparkling water particles into the air. The camera shakes slightly, adding subtle intimacy to the framing. [Shot 3] At 00:03.833, the shot transitions to a closer medium close-up as the camera slowly pushes in. The fully materialized water-fox from Shot 2 playfully circles the elf's extended hand, nuzzling her fingers with its fluid snout. The elf smiles softly, her shoulders relaxing. Magical, luminescent ripples spread outward across the stream's surface below them, reflecting the teal light. [Shot 4] At 00:05.750, the camera cuts to a medium shot, executing a smooth tracking shot to the right as the elf and the water-fox move together across the ancient stone bridge. The fox bounds gracefully between mossy stones, leaving a vibrant trail of glowing, suspended water droplets in its wake. Dense jungle foliage heavily frames the foreground and background of the stone path. [Shot 5] At 00:07.666, the scene shifts to a medium wide shot where the camera performs a dynamic arc shot around the pair. The elf sweeps her arms upward in a graceful gesture. Luminous, fluid streams of teal water rise from the stream, spiraling like living ribbons around her and the magical companion. The water-fox darts joyfully through the swirling aquatic ribbons, its body blending seamlessly with the magical currents. [Shot 6] At 00:09.583, the camera cuts to a medium shot and pans right to capture their fast but smooth motion. The water-fox leaps energetically through floating rings of glowing water, sending sparkling droplets splashing playfully around the elf. She throws her head back in a visual laugh, reaching both hands toward the creature as the suspended magical droplets catch the warm, volumetric sunlight. [Shot 7] At 00:11.500, the camera cuts to a cinematic wide shot and gently pulls out. The elf and the water-fox sit side-by-side at the far edge of the stone bridge, their backs partially turned as they overlook the glowing jungle stream. The fox slowly settles its fluid body onto the stone beside her hip. Warm yellow sunlight breaks through the dense canopy above, illuminating the drifting mist and suspended water particles as the environment fades into a soft, hazy light. overall_soundscape: A continuous, loud flow of rushing water dominates the environment, accompanied by the distinct rustle of leaves in the breeze. As the fox materializes, pronounced, resonant magical chimes and clearly heard liquid splashes ring out. Audible, rhythmic splashing and the fluid sloshing of water follow as the creature bounds across the stones, mingling with distinct, glassy tinkling sounds of water droplets popping in the air and a faint but clearly audible creature trill. non_diegetic_music: An orchestral fantasy score plays at a slow tempo, featuring continuous string melodies, sustained low cellos, and light woodwind flourishes.
PromptShow full promptCollapse promptWatercolor and pencil concept art evolving into cinematic reality, wide shot, the camera pushes in over a flat hand-drawn architectural masterplan of a riverside village on sketchbook paper. Visible pencil lines, construction marks, and watercolor stains cover the surface. As the camera continuously pushes in and pedestals down toward a central bridge, the ink lines gain depth. Trees extrude from the paper, watercolor canals begin to shimmer, and tiny sketch-like figures start moving. Buildings rise into three-dimensional forms with textured roofs and shifting shadows. The paper surface dissolves as the transformation accelerates. The camera tracks low above the canal while the water flows and trees sway gently. The final pencil traces vanish into soft, warm sunlight, leaving a fully realized, photorealistic living village in the final framing. overall_soundscape: A soft scraping of pencil on paper and a faint paper rustle open the scene, followed by a gentle whoosh as the world extrudes. This transitions into the audible bubbling of flowing water, subtle wind rustling the tree leaves, and the distant, faint chatter and footsteps of the village figures. non_diegetic_music: Solo acoustic guitar and sustained ambient string pads, slow tempo, building steadily into a warm, orchestral swell that resolves on a bright, major chord.
integrated_multimodal_description: [Shot 1] Watercolor and pencil concept art evolving into cinematic reality, wide shot, the camera pushes in over a flat hand-drawn architectural masterplan of a riverside village on sketchbook paper. Visible pencil lines, construction marks, and watercolor stains cover the surface. As the camera continuously pushes in and pedestals down toward a central bridge, the ink lines gain depth. Trees extrude from the paper, watercolor canals begin to shimmer, and tiny sketch-like figures start moving. Buildings rise into three-dimensional forms with textured roofs and shifting shadows. The paper surface dissolves as the transformation accelerates. The camera tracks low above the canal while the water flows and trees sway gently. The final pencil traces vanish into soft, warm sunlight, leaving a fully realized, photorealistic living village in the final framing. overall_soundscape: A soft scraping of pencil on paper and a faint paper rustle open the scene, followed by a gentle whoosh as the world extrudes. This transitions into the audible bubbling of flowing water, subtle wind rustling the tree leaves, and the distant, faint chatter and footsteps of the village figures. non_diegetic_music: Solo acoustic guitar and sustained ambient string pads, slow tempo, building steadily into a warm, orchestral swell that resolves on a bright, major chord.
PromptShow full promptCollapse promptUltra-sharp CGI sculpture of the faceless humanoid (S1) carved from opaque white soap, featuring long soft hair and an oversized coat with smooth rounded forms and a velvety matte surface, under soft studio rim lighting against a pitch black background. The faceless humanoid (S1) brings its hands close and delicately blows a tiny translucent soap bubble into the surrounding dark space. [Cut to 07.1] [Shot 2] Close-up of the tiny soap bubble floating away into the pitch-black shadows, catching a soft glint of white light on its shimmering surface. overall_soundscape: Soft breath sound, gentle air puff, and quiet bubble wobble. non_diegetic_music: N/A
integrated_multimodal_description: [Shot 1] Ultra-sharp CGI sculpture of the faceless humanoid (S1) carved from opaque white soap, featuring long soft hair and an oversized coat with smooth rounded forms and a velvety matte surface, under soft studio rim lighting against a pitch black background. The faceless humanoid (S1) brings its hands close and delicately blows a tiny translucent soap bubble into the surrounding dark space. [Cut to 07.1] [Shot 2] Close-up of the tiny soap bubble floating away into the pitch-black shadows, catching a soft glint of white light on its shimmering surface. overall_soundscape: Soft breath sound, gentle air puff, and quiet bubble wobble. non_diegetic_music: N/A
PromptShow full promptCollapse promptCinematic, WS, a handheld Tracking Shot follows behind a tall, slender elf princess entering a sunlit battlefield strewn with shattered timber and stone rubble under bright, clear daylight. The princess has fair freckled skin, blue-gray eyes, long wavy platinum-blonde hair, and pointed ears; she wears a silver branch circlet and a flowing silver battle gown with a scale-textured mantle, holding a gleaming silver longsword in her right hand. Ahead of her in the bright, high-contrast light, armored green-skinned orcs occupy successive depth pockets across the dry earth. A nearby orc wearing a spiked iron helmet and wielding a jagged axe charges forward. The princess sharply drops into a crisp micro-anticipation pose. A brilliant white-silver flash completely engulfs her, and she instantly vanishes, leaving the air perfectly clear with zero motion blur. The continuous tracking movement drifts steadily forward through the empty space over the debris. A second local white-silver flash erupts to the right of the charging orc, revealing the princess fully formed in mid-air in an attack stance. She delivers a sharp diagonal Zornhau cut to the orc's chest, releasing a sudden, crisp spray of bright green fluid. The defeated orc falls back heavily as the steady tracking continues, smoothly shifting left to keep the unbroken action sharply resolved. Another distinct orc entirely clad in rust-colored chainmail lunges from the side with a spear. The princess vanishes in another dry flash just as the spear tip passes through her previous position. A local flash illuminates the space directly behind the chainmail orc, and the princess reappears, driving a precise straight thrust into his back. Bright green fluid bursts outward in sharp droplets as he drops to his knees. The tracking motion smoothly retreats backward over a discarded wooden shield as a heavy orc captain wearing dark iron shoulder armor and clutching a massive iron cleaver descends from a pile of rubble. The captain swings the cleaver downward in a brutal arc; the princess vanishes in a flash of white-silver light a fraction of a second before the blade strikes the earth. A final white-silver flash cracks the air directly above the captain. The princess materializes above him, bringing her silver longsword down in a vertical Scheitelhau finishing strike. Crisp green fluid erupts from the captain's shoulder armor as he crashes into the dust, while the princess lands gracefully, her platinum-blonde hair and silver mantle settling with perfect clarity against the bright sunlit background. overall_soundscape: The bright daylight environment is filled with a subtle, continuous ambient wind and distant battle rumble. This background is sharply punctuated by the loud, pronounced dry crackle and popping impact of the white-silver flashes. Distinct, sharp metallic swooshes and loud ringing clangs dominate the foreground as the silver longsword cuts the air and meets armor. Each successful sword strike is immediately followed by a pronounced, wet squelch and splashing sound of green fluid erupting, culminating in loud, heavy thuds and clattering metal as the defeated armored orcs crash into the hard dirt. Rapid, heavy footsteps of charging orcs thump continuously against the earth. non_diegetic_music: Fast-paced, relentless cinematic percussion featuring heavy taiko drums and sharp wooden clacks, driving an accelerating tempo that synchronizes tightly with the sword strikes, maintaining pure rhythmic tension without any sweeping melodic lines.
integrated_multimodal_description: [Shot 1] Cinematic, WS, a handheld Tracking Shot follows behind a tall, slender elf princess entering a sunlit battlefield strewn with shattered timber and stone rubble under bright, clear daylight. The princess has fair freckled skin, blue-gray eyes, long wavy platinum-blonde hair, and pointed ears; she wears a silver branch circlet and a flowing silver battle gown with a scale-textured mantle, holding a gleaming silver longsword in her right hand. Ahead of her in the bright, high-contrast light, armored green-skinned orcs occupy successive depth pockets across the dry earth. A nearby orc wearing a spiked iron helmet and wielding a jagged axe charges forward. The princess sharply drops into a crisp micro-anticipation pose. A brilliant white-silver flash completely engulfs her, and she instantly vanishes, leaving the air perfectly clear with zero motion blur. The continuous tracking movement drifts steadily forward through the empty space over the debris. A second local white-silver flash erupts to the right of the charging orc, revealing the princess fully formed in mid-air in an attack stance. She delivers a sharp diagonal Zornhau cut to the orc's chest, releasing a sudden, crisp spray of bright green fluid. The defeated orc falls back heavily as the steady tracking continues, smoothly shifting left to keep the unbroken action sharply resolved. Another distinct orc entirely clad in rust-colored chainmail lunges from the side with a spear. The princess vanishes in another dry flash just as the spear tip passes through her previous position. A local flash illuminates the space directly behind the chainmail orc, and the princess reappears, driving a precise straight thrust into his back. Bright green fluid bursts outward in sharp droplets as he drops to his knees. The tracking motion smoothly retreats backward over a discarded wooden shield as a heavy orc captain wearing dark iron shoulder armor and clutching a massive iron cleaver descends from a pile of rubble. The captain swings the cleaver downward in a brutal arc; the princess vanishes in a flash of white-silver light a fraction of a second before the blade strikes the earth. A final white-silver flash cracks the air directly above the captain. The princess materializes above him, bringing her silver longsword down in a vertical Scheitelhau finishing strike. Crisp green fluid erupts from the captain's shoulder armor as he crashes into the dust, while the princess lands gracefully, her platinum-blonde hair and silver mantle settling with perfect clarity against the bright sunlit background. overall_soundscape: The bright daylight environment is filled with a subtle, continuous ambient wind and distant battle rumble. This background is sharply punctuated by the loud, pronounced dry crackle and popping impact of the white-silver flashes. Distinct, sharp metallic swooshes and loud ringing clangs dominate the foreground as the silver longsword cuts the air and meets armor. Each successful sword strike is immediately followed by a pronounced, wet squelch and splashing sound of green fluid erupting, culminating in loud, heavy thuds and clattering metal as the defeated armored orcs crash into the hard dirt. Rapid, heavy footsteps of charging orcs thump continuously against the earth. non_diegetic_music: Fast-paced, relentless cinematic percussion featuring heavy taiko drums and sharp wooden clacks, driving an accelerating tempo that synchronizes tightly with the sword strikes, maintaining pure rhythmic tension without any sweeping melodic lines.
PromptShow full promptCollapse promptDocumentary videography in the style of Martin Parr and Martha Cooper, handheld and super shaky camcorder video, medium tracking shot at night down the wet street past the abandoned toy store with broken windows and faded paint. The cute Eurasian woman (S4) with black hair, wearing the DIY full-face helmet cobbled together from scrap plastic parts, electronics circuit boards, broken glass, and multicolored wires spilling messily down her shoulders, paired with the oversized white tracksuit, points ahead while riding the zebra. Another scruffy stray dog (S5) suddenly emerges from the shadows near the storefront, barking warningly. The grey Italian greyhound dog with a white chest patch (S1) stops and barks back. The camera pans sharply to follow the action in the heavy downpour. overall_soundscape: Multiple dogs barking overlapping, heavy rain pouring, and thunder rumbling. non_diegetic_music: N/A
integrated_multimodal_description: [Shot 1] Documentary videography in the style of Martin Parr and Martha Cooper, handheld and super shaky camcorder video, medium tracking shot at night down the wet street past the abandoned toy store with broken windows and faded paint. The cute Eurasian woman (S4) with black hair, wearing the DIY full-face helmet cobbled together from scrap plastic parts, electronics circuit boards, broken glass, and multicolored wires spilling messily down her shoulders, paired with the oversized white tracksuit, points ahead while riding the zebra. Another scruffy stray dog (S5) suddenly emerges from the shadows near the storefront, barking warningly. The grey Italian greyhound dog with a white chest patch (S1) stops and barks back. The camera pans sharply to follow the action in the heavy downpour. overall_soundscape: Multiple dogs barking overlapping, heavy rain pouring, and thunder rumbling. non_diegetic_music: N/A
PromptShow full promptCollapse promptMedium-wide handheld documentary shot in the photographic style of Martin Parr and Martha Cooper, featuring high-contrast theatrical staging with a pitch-black backdrop. In the center-left, a woman with dark hair sits upright on a light brown wooden chair, wearing a full-length, long-sleeved light purple gown. Beside her, a man with short dark hair sits directly on the light-colored stage floor in a cross-legged position, wearing a dark navy blazer, trousers, and dark shoes. A soft white spotlight shines vertically down upon them. In the extreme foreground, a low, warm golden-orange light source illuminates sparse yellow wildflowers and dandelions growing along the stage edge, casting an upward warm glow. The woman and man face each other, moving their mouths as they sing in operatic harmony. The handheld camcorder slowly sways from side to side. overall_soundscape: Echoing operatic singing voices, subtle rustle of gown fabric, and faint stage ambience. non_diegetic_music: N/A
integrated_multimodal_description: [Shot 1] Medium-wide handheld documentary shot in the photographic style of Martin Parr and Martha Cooper, featuring high-contrast theatrical staging with a pitch-black backdrop. In the center-left, a woman with dark hair sits upright on a light brown wooden chair, wearing a full-length, long-sleeved light purple gown. Beside her, a man with short dark hair sits directly on the light-colored stage floor in a cross-legged position, wearing a dark navy blazer, trousers, and dark shoes. A soft white spotlight shines vertically down upon them. In the extreme foreground, a low, warm golden-orange light source illuminates sparse yellow wildflowers and dandelions growing along the stage edge, casting an upward warm glow. The woman and man face each other, moving their mouths as they sing in operatic harmony. The handheld camcorder slowly sways from side to side. overall_soundscape: Echoing operatic singing voices, subtle rustle of gown fabric, and faint stage ambience. non_diegetic_music: N/A
PromptShow full promptCollapse promptCinematic, a medium shot from a slightly low angle where the camera executes a Tracking Shot to the left, following alongside the subject. An on-screen young Japanese woman (S1) in her early twenties with a slim build and shoulder-length dark brown hair appears in the center frame. She wears an elegant, feminine summer dress made of flowing white chiffon with a soft pink and yellow floral pattern. She holds a clear plastic cup filled with a cold green iced drink, condensation visible on the outside. She walks casually through a charming Japanese city street in the midground, lined with traditional wooden storefronts and soft green foliage in the background. Warm afternoon sunlight streams from the top right, casting dappled shadows on the pavement and gently illuminating her hair. She walks naturally, turning her head to look around at the shops. A gentle breeze flutters her hair and the light fabric of her dress. [Shot 2] At 00:05.000, the camera cuts to a medium close-up where the framing continues to Shake Slightly to simulate handheld filming. The woman (S1) is now stopped in the center of the frame in front of a small quaint café with a navy blue noren curtain and potted ferns. She brings the cold drink to her lips and takes a small sip. Lowering the cup, she pulls a slim white smartphone from her pocket with her free hand, glances at the illuminated screen, and then pockets it again. She looks around the street with a soft, cute smile, her shoulders relaxed in the warm light. [Shot 3] At 00:09.800, the shot transitions to a close-up where the camera continues to Shake Slightly. The woman (S1) suddenly notices the camera lens. Her eyes light up, and her lips part into a sweet, endearing smile. She slightly raises her iced drink toward the lens in a gentle "cheers" motion. Looking directly into the camera, the young woman (S1) says naturally, <d>[Japanese] 今日は自分の時間。</d> After speaking, she gracefully pivots her body to the right, turning her back to the camera, and begins walking away down the sunlit street as the clip ends. overall_soundscape: A soft wind hiss and distant traffic hum provide an ambient outdoor backdrop. Pronounced rhythmic tapping of light footsteps on pavement is clearly heard, accompanied by the distinct rustle of flowing chiffon fabric. A sharp clink of ice cubes against a plastic cup rings out in the foreground, followed by a soft sipping sound. The subtle click and clatter of a smartphone casing being handled is audible, leading directly into the clear, close-up spoken female voice at the end. non_diegetic_music: Solo acoustic guitar, slow tempo, featuring a bright and airy fingerpicked melody with no percussion.
integrated_multimodal_description: [Shot 1] Cinematic, a medium shot from a slightly low angle where the camera executes a Tracking Shot to the left, following alongside the subject. An on-screen young Japanese woman (S1) in her early twenties with a slim build and shoulder-length dark brown hair appears in the center frame. She wears an elegant, feminine summer dress made of flowing white chiffon with a soft pink and yellow floral pattern. She holds a clear plastic cup filled with a cold green iced drink, condensation visible on the outside. She walks casually through a charming Japanese city street in the midground, lined with traditional wooden storefronts and soft green foliage in the background. Warm afternoon sunlight streams from the top right, casting dappled shadows on the pavement and gently illuminating her hair. She walks naturally, turning her head to look around at the shops. A gentle breeze flutters her hair and the light fabric of her dress. [Shot 2] At 00:05.000, the camera cuts to a medium close-up where the framing continues to Shake Slightly to simulate handheld filming. The woman (S1) is now stopped in the center of the frame in front of a small quaint café with a navy blue noren curtain and potted ferns. She brings the cold drink to her lips and takes a small sip. Lowering the cup, she pulls a slim white smartphone from her pocket with her free hand, glances at the illuminated screen, and then pockets it again. She looks around the street with a soft, cute smile, her shoulders relaxed in the warm light. [Shot 3] At 00:09.800, the shot transitions to a close-up where the camera continues to Shake Slightly. The woman (S1) suddenly notices the camera lens. Her eyes light up, and her lips part into a sweet, endearing smile. She slightly raises her iced drink toward the lens in a gentle "cheers" motion. Looking directly into the camera, the young woman (S1) says naturally, <d>[Japanese] 今日は自分の時間。</d> After speaking, she gracefully pivots her body to the right, turning her back to the camera, and begins walking away down the sunlit street as the clip ends. overall_soundscape: A soft wind hiss and distant traffic hum provide an ambient outdoor backdrop. Pronounced rhythmic tapping of light footsteps on pavement is clearly heard, accompanied by the distinct rustle of flowing chiffon fabric. A sharp clink of ice cubes against a plastic cup rings out in the foreground, followed by a soft sipping sound. The subtle click and clatter of a smartphone casing being handled is audible, leading directly into the clear, close-up spoken female voice at the end. non_diegetic_music: Solo acoustic guitar, slow tempo, featuring a bright and airy fingerpicked melody with no percussion.
PromptShow full promptCollapse promptCinematic, a low-angle tracking shot follows behind a male photographer in his mid-30s sprinting through a bombed European street under overcast daylight. He has a stubble-covered face with a visible eyebrow wound, wearing a dark overcoat, shirt, trousers, dark shoes, and a leather satchel. A single silver-black period press camera hangs from a strap across his chest. The olive-charcoal-sepia environment is filled with stone ruins, rubble, black smoke, crackling fires, and a leaning red-and-cream tram. A shell detonates ahead, lifting heavy masonry and snapping overhead cables. The tram lurches as the photographer dives behind a stone block, shielding his camera. Debris flies close to the lens. He remains completely still as his stunningly wide eyes open through the settling dust. [Shot 2] At 00:01.916, the camera cuts to a tight static shot that shakes slightly. The photographer's dusty fingers check the silver-black camera, revealing a mechanical indicator showing one exposure left. He opens his leather satchel, exposing only empty film sleeves. The focus shifts from the camera's mechanical indicator to his face as he understands his situation; he closes the camera, rises to his feet, and his shoulders slump briefly. [Shot 3] At 00:03.833, the camera cuts to a tracking shot following laterally beside the photographer from Shot 1 as he runs past the leaning red-and-cream tram. Indistinct civilians cross the background beneath thick black smoke and collapsing architecture. Reaching a vantage point, he raises the camera to his eye. Through the viewfinder framing, his finger reaches the shutter button but stops without pressing it. His jaw tenses as he breathes heavily. [Shot 4] At 00:06.229, the shot transitions to a static shot framing a woman trapped beside the tram beneath a light wooden timber. She has auburn hair, gray-green eyes, a bleeding temple wound, a headscarf, a beige coat, a blue-gray dress, a scarf, stockings, shoes, and a silver locket. She coughs forcefully and reaches out her hand. The photographer briefly frames her in the foreground, then his finger leaves the shutter. He lowers the camera with softened eyes and immediately sprints toward her. [Shot 5] At 00:08.625, the camera cuts to an arc shot circling the pair as the photographer clears heavy stones, lifts the timber, and pulls the woman from Shot 4 upright, revealing her realistic physical weight. A stone building facade cracks in the background, erupting in thick volumetric dust. Her arm crosses his shoulders, and they stagger-run forward while the tram's glass windows burst violently behind them. The silver-black camera swings naturally from his strap. Both display terrified, exhausted expressions as they struggle forward. [Shot 6] At 00:11.979, the camera cuts to a tracking shot keeping tight on their side profiles. Another explosive blast knocks them down onto the rubble. The photographer shields the woman, and the swinging camera aggressively hits his chest. He catches it, accidentally pressing the shutter button as they look deeply at each other, realizing they are alive. Bright firelight reflects across the glass camera lens. The motion freezes for a fraction of a second, then the volumetric dust continues raining down over them. [Shot 7] At 00:13.416, the camera cuts to a static shot showing the resulting photograph filling the frame. The image is a monochrome, imperfect, tilted, and highly grainy 35mm print, depicting the photographer supporting the woman amidst thick smoke, her silver locket catching the light. The frame holds completely still on this image. overall_soundscape: A deafening boom initiates the scene, followed by loud crashing masonry, snapping cables, and the loud clatter of groaning metal, which abruptly transitions into a high-pitched whine and a slow, muffled heartbeat. As dusty grit patters audibly, a distinct metallic click of a camera mechanism and rustling leather cut through, gradually giving way to the chaotic foreground crunch of boots on rubble, heavy panting, crackling fires, and distant thudding artillery. The ambient roar dips beneath a sharp audible inhale, creaking leather straps, and frantic coughing, before wood violently splinters and glass shatters with a loud crash during a frantic run; this chaos abruptly drops into silence, punctuated by one enormous, dominant mechanical clack of a shutter, a soft winding click, and finally, faint heartbeats blending with distant crackling flames. non_diegetic_music: Initially completely silent, the score introduces a minimal, slow-tempo sustained low solo cello halfway through the scene, which is eventually joined by sparse, restrained percussion beats, culminating in a single, resonating low cello note at the end.
integrated_multimodal_description: [Shot 1] Cinematic, a low-angle tracking shot follows behind a male photographer in his mid-30s sprinting through a bombed European street under overcast daylight. He has a stubble-covered face with a visible eyebrow wound, wearing a dark overcoat, shirt, trousers, dark shoes, and a leather satchel. A single silver-black period press camera hangs from a strap across his chest. The olive-charcoal-sepia environment is filled with stone ruins, rubble, black smoke, crackling fires, and a leaning red-and-cream tram. A shell detonates ahead, lifting heavy masonry and snapping overhead cables. The tram lurches as the photographer dives behind a stone block, shielding his camera. Debris flies close to the lens. He remains completely still as his stunningly wide eyes open through the settling dust. [Shot 2] At 00:01.916, the camera cuts to a tight static shot that shakes slightly. The photographer's dusty fingers check the silver-black camera, revealing a mechanical indicator showing one exposure left. He opens his leather satchel, exposing only empty film sleeves. The focus shifts from the camera's mechanical indicator to his face as he understands his situation; he closes the camera, rises to his feet, and his shoulders slump briefly. [Shot 3] At 00:03.833, the camera cuts to a tracking shot following laterally beside the photographer from Shot 1 as he runs past the leaning red-and-cream tram. Indistinct civilians cross the background beneath thick black smoke and collapsing architecture. Reaching a vantage point, he raises the camera to his eye. Through the viewfinder framing, his finger reaches the shutter button but stops without pressing it. His jaw tenses as he breathes heavily. [Shot 4] At 00:06.229, the shot transitions to a static shot framing a woman trapped beside the tram beneath a light wooden timber. She has auburn hair, gray-green eyes, a bleeding temple wound, a headscarf, a beige coat, a blue-gray dress, a scarf, stockings, shoes, and a silver locket. She coughs forcefully and reaches out her hand. The photographer briefly frames her in the foreground, then his finger leaves the shutter. He lowers the camera with softened eyes and immediately sprints toward her. [Shot 5] At 00:08.625, the camera cuts to an arc shot circling the pair as the photographer clears heavy stones, lifts the timber, and pulls the woman from Shot 4 upright, revealing her realistic physical weight. A stone building facade cracks in the background, erupting in thick volumetric dust. Her arm crosses his shoulders, and they stagger-run forward while the tram's glass windows burst violently behind them. The silver-black camera swings naturally from his strap. Both display terrified, exhausted expressions as they struggle forward. [Shot 6] At 00:11.979, the camera cuts to a tracking shot keeping tight on their side profiles. Another explosive blast knocks them down onto the rubble. The photographer shields the woman, and the swinging camera aggressively hits his chest. He catches it, accidentally pressing the shutter button as they look deeply at each other, realizing they are alive. Bright firelight reflects across the glass camera lens. The motion freezes for a fraction of a second, then the volumetric dust continues raining down over them. [Shot 7] At 00:13.416, the camera cuts to a static shot showing the resulting photograph filling the frame. The image is a monochrome, imperfect, tilted, and highly grainy 35mm print, depicting the photographer supporting the woman amidst thick smoke, her silver locket catching the light. The frame holds completely still on this image. overall_soundscape: A deafening boom initiates the scene, followed by loud crashing masonry, snapping cables, and the loud clatter of groaning metal, which abruptly transitions into a high-pitched whine and a slow, muffled heartbeat. As dusty grit patters audibly, a distinct metallic click of a camera mechanism and rustling leather cut through, gradually giving way to the chaotic foreground crunch of boots on rubble, heavy panting, crackling fires, and distant thudding artillery. The ambient roar dips beneath a sharp audible inhale, creaking leather straps, and frantic coughing, before wood violently splinters and glass shatters with a loud crash during a frantic run; this chaos abruptly drops into silence, punctuated by one enormous, dominant mechanical clack of a shutter, a soft winding click, and finally, faint heartbeats blending with distant crackling flames. non_diegetic_music: Initially completely silent, the score introduces a minimal, slow-tempo sustained low solo cello halfway through the scene, which is eventually joined by sparse, restrained percussion beats, culminating in a single, resonating low cello note at the end.
PromptShow full promptCollapse prompt2D-animated, extreme close-up, pull out. The scene opens instantly on a flat cel-colored graphic close-up of a young woman's right eye, featuring a large turquoise almond shape, bold black winged eyeliner, and soft pink-purple eyeshadow. She blinks exactly once. The camera sharply pulls out to a tight portrait as the woman turns toward the camera with a confident, sideways smile. Her full face reveals pointed elf-like ears, a small nose, and small stylized lips. Her extremely long golden-blonde hair, featuring dramatic outward-pointing layered spikes and soft peach-pink gradient tips, sweeps across the frame in large, layered graphic shapes. Oversized turquoise hoop earrings hang from her ears. Flat red, turquoise, and warm-yellow starbursts explode outward in the background, illuminated by a flat, even, bright studio-style cel light, as large condensed white text "GOLDEN HEAT" pops directly behind her head. [Shot 2] At 00:01.400, the camera cuts to a full-body tracking shot moving alongside the woman from Shot 1 in side-profile. She walks confidently across a completely abstract, flat red background. Her full outfit is visible: a red choker with a small gold ornament, an asymmetrical red cropped top with white accent panels, an asymmetrical red mini skirt, a circular gold belt buckle, stacked gold bracelets, long translucent golden-yellow fabric panels flowing from the waist, and elegant red strappy high heels. Her long layered hair and the translucent yellow panels trail behind her. The camera then executes a rapid arc shot, spinning around her sharp graphic silhouette to settle on a front-facing medium composition. [Shot 3] At 00:02.800, the shot transitions to a medium static shot. The woman from Shot 1 lifts one hand toward her face, playfully tilts her head, and sharply flicks her wrist outward. Her stacked gold bracelets clink visually, creating three circular graphic echoes that expand rapidly toward the lens, transforming into a wipe transition. Inside these expanding circles, quick inset close-ups flash sequentially: her turquoise eye, her turquoise hoop earring, her circular gold belt buckle, and finally her red strappy heel. [Shot 4] At 00:04.200, the camera cuts to a montage sequence of medium static shots divided into bold comic-style panels. The woman from Shot 1 strikes a confident front pose, then shifts into a sharp side profile, followed by resting one hand on her hip for an over-the-shoulder glance, and finishes with a playful head tilt. The flat backgrounds alternate rapidly between deep crimson, warm cream, turquoise, and golden yellow. Her spiked golden-blonde hair repeatedly breaks across the sharp geometric panel borders. [Shot 5] At 00:05.800, the shot changes to an extreme close-up static shot. The composition focuses entirely on the turquoise hoop earring from Shot 1 swinging on her pointed ear. The circular earring scales up rapidly until it fills the entire screen, becoming a solid turquoise ring transition. Inside the expanding ring, flat gold spark shapes and thick red graphic rays rotate rapidly against a cream background. [Shot 6] At 00:07.100, the camera cuts to a full-body low-angle tilt up. The woman from Shot 1 takes one powerful step forward toward the camera. On the impact of her red strappy heel, the flat floor instantly disappears into a massive, jagged red graphic starburst. The camera rapidly tilts up from her red heel, traveling past the flowing golden fabric panels, the gold belt buckle, the red cropped top, and settling on her face. She finishes the movement with one hand firmly on her hip and a confident, direct gaze. [Shot 7] At 00:08.800, the shot transitions to a medium static shot. The woman from Shot 1 executes a massive hair flip. Her enormous golden-blonde hair sweeps dramatically from the left side of the frame to the right, functioning as a natural animated wipe. The peach-pink hair tips stretch into flat graphic smear frames across the screen, revealing her dark silhouette standing against a vibrant turquoise background entirely filled with flat, rotating retro star shapes. [Shot 8] At 00:10.300, the camera cuts to a rapid series of extreme close-up static shots. Ultra-fast beauty cuts flash on screen: her bold eyeliner and turquoise eye, her small confident smile, her pointed ear with the turquoise hoop earring, her stacked gold bracelets, her circular gold belt buckle, and the crossed red straps of her heel. Each distinct detail appears centered inside a different rapidly changing geometric frame, flashing in perfect sync. [Shot 9] At 00:11.600, the shot changes to a full-body static shot. The woman from Shot 1 strikes a clean fashion pose against an abstract background composition of oversized red circles, golden sunbursts, and turquoise geometric shapes. She rotates slightly from a three-quarter angle to face the camera directly, while her extremely long golden hair fans out dramatically behind her in crisp, angular shapes. Massive bold typography reading "SUNSET ICON" slides vertically downward behind her sharp silhouette. [Shot 10] At 00:13.100, the camera cuts to a tight portrait static shot. The woman from Shot 1 executes a rapid sequence: a quick eye blink, a sharp head tilt, a fast hair sweep, and a final confident facial expression with one eyebrow slightly raised. On the last movement, the frame freezes into a warm-toned hero portrait. She places one hand on her hip. Her massive golden hair forms a huge, jagged graphic silhouette behind her. The bold typography "GOLDEN HEAT" expands horizontally across the abstract background, while small turquoise stars and gold sparkle symbols continuously orbit the entire composition. overall_soundscape: A loud, distinct electronic snap and crisp pop accompany the opening blink and head turn, followed immediately by rhythmic, sharp high-heeled footsteps clicking clearly against a solid surface. As the visual transitions occur, pronounced synthesized whooshes, high-pitched sparkle hits, and heavy metallic clatters from the gold bracelets ring out in the foreground. The sequence is punctuated by a loud, resonant bass thud on the heel impact, crackling energy sizzles during the starbursts, and a final crisp synthetic chime as the typography locks into place. non_diegetic_music: High-energy Japanese dance-pop song, fast tempo, featuring a punchy four-on-the-floor electronic kick drum and bright synthesizer leads. Enthusiastic, upbeat female J-pop vocals drive the melody, seamlessly matching the rapid rhythmic cuts and hyper-kinetic energy of the animation.
integrated_multimodal_description: [Shot 1] 2D-animated, extreme close-up, pull out. The scene opens instantly on a flat cel-colored graphic close-up of a young woman's right eye, featuring a large turquoise almond shape, bold black winged eyeliner, and soft pink-purple eyeshadow. She blinks exactly once. The camera sharply pulls out to a tight portrait as the woman turns toward the camera with a confident, sideways smile. Her full face reveals pointed elf-like ears, a small nose, and small stylized lips. Her extremely long golden-blonde hair, featuring dramatic outward-pointing layered spikes and soft peach-pink gradient tips, sweeps across the frame in large, layered graphic shapes. Oversized turquoise hoop earrings hang from her ears. Flat red, turquoise, and warm-yellow starbursts explode outward in the background, illuminated by a flat, even, bright studio-style cel light, as large condensed white text "GOLDEN HEAT" pops directly behind her head. [Shot 2] At 00:01.400, the camera cuts to a full-body tracking shot moving alongside the woman from Shot 1 in side-profile. She walks confidently across a completely abstract, flat red background. Her full outfit is visible: a red choker with a small gold ornament, an asymmetrical red cropped top with white accent panels, an asymmetrical red mini skirt, a circular gold belt buckle, stacked gold bracelets, long translucent golden-yellow fabric panels flowing from the waist, and elegant red strappy high heels. Her long layered hair and the translucent yellow panels trail behind her. The camera then executes a rapid arc shot, spinning around her sharp graphic silhouette to settle on a front-facing medium composition. [Shot 3] At 00:02.800, the shot transitions to a medium static shot. The woman from Shot 1 lifts one hand toward her face, playfully tilts her head, and sharply flicks her wrist outward. Her stacked gold bracelets clink visually, creating three circular graphic echoes that expand rapidly toward the lens, transforming into a wipe transition. Inside these expanding circles, quick inset close-ups flash sequentially: her turquoise eye, her turquoise hoop earring, her circular gold belt buckle, and finally her red strappy heel. [Shot 4] At 00:04.200, the camera cuts to a montage sequence of medium static shots divided into bold comic-style panels. The woman from Shot 1 strikes a confident front pose, then shifts into a sharp side profile, followed by resting one hand on her hip for an over-the-shoulder glance, and finishes with a playful head tilt. The flat backgrounds alternate rapidly between deep crimson, warm cream, turquoise, and golden yellow. Her spiked golden-blonde hair repeatedly breaks across the sharp geometric panel borders. [Shot 5] At 00:05.800, the shot changes to an extreme close-up static shot. The composition focuses entirely on the turquoise hoop earring from Shot 1 swinging on her pointed ear. The circular earring scales up rapidly until it fills the entire screen, becoming a solid turquoise ring transition. Inside the expanding ring, flat gold spark shapes and thick red graphic rays rotate rapidly against a cream background. [Shot 6] At 00:07.100, the camera cuts to a full-body low-angle tilt up. The woman from Shot 1 takes one powerful step forward toward the camera. On the impact of her red strappy heel, the flat floor instantly disappears into a massive, jagged red graphic starburst. The camera rapidly tilts up from her red heel, traveling past the flowing golden fabric panels, the gold belt buckle, the red cropped top, and settling on her face. She finishes the movement with one hand firmly on her hip and a confident, direct gaze. [Shot 7] At 00:08.800, the shot transitions to a medium static shot. The woman from Shot 1 executes a massive hair flip. Her enormous golden-blonde hair sweeps dramatically from the left side of the frame to the right, functioning as a natural animated wipe. The peach-pink hair tips stretch into flat graphic smear frames across the screen, revealing her dark silhouette standing against a vibrant turquoise background entirely filled with flat, rotating retro star shapes. [Shot 8] At 00:10.300, the camera cuts to a rapid series of extreme close-up static shots. Ultra-fast beauty cuts flash on screen: her bold eyeliner and turquoise eye, her small confident smile, her pointed ear with the turquoise hoop earring, her stacked gold bracelets, her circular gold belt buckle, and the crossed red straps of her heel. Each distinct detail appears centered inside a different rapidly changing geometric frame, flashing in perfect sync. [Shot 9] At 00:11.600, the shot changes to a full-body static shot. The woman from Shot 1 strikes a clean fashion pose against an abstract background composition of oversized red circles, golden sunbursts, and turquoise geometric shapes. She rotates slightly from a three-quarter angle to face the camera directly, while her extremely long golden hair fans out dramatically behind her in crisp, angular shapes. Massive bold typography reading "SUNSET ICON" slides vertically downward behind her sharp silhouette. [Shot 10] At 00:13.100, the camera cuts to a tight portrait static shot. The woman from Shot 1 executes a rapid sequence: a quick eye blink, a sharp head tilt, a fast hair sweep, and a final confident facial expression with one eyebrow slightly raised. On the last movement, the frame freezes into a warm-toned hero portrait. She places one hand on her hip. Her massive golden hair forms a huge, jagged graphic silhouette behind her. The bold typography "GOLDEN HEAT" expands horizontally across the abstract background, while small turquoise stars and gold sparkle symbols continuously orbit the entire composition. overall_soundscape: A loud, distinct electronic snap and crisp pop accompany the opening blink and head turn, followed immediately by rhythmic, sharp high-heeled footsteps clicking clearly against a solid surface. As the visual transitions occur, pronounced synthesized whooshes, high-pitched sparkle hits, and heavy metallic clatters from the gold bracelets ring out in the foreground. The sequence is punctuated by a loud, resonant bass thud on the heel impact, crackling energy sizzles during the starbursts, and a final crisp synthetic chime as the typography locks into place. non_diegetic_music: High-energy Japanese dance-pop song, fast tempo, featuring a punchy four-on-the-floor electronic kick drum and bright synthesizer leads. Enthusiastic, upbeat female J-pop vocals drive the melody, seamlessly matching the rapid rhythmic cuts and hyper-kinetic energy of the animation.

Overview

Adaptive BSA

Block-Sparse Attention (BSA) splits attention into two complementary branches: an exact branch that computes selected blocks exactly, and a compressed branch that covers the remaining content at much lower complexity. However, fixed block partitioning limits the exact branch, and mean estimation — a common construction of the compressed branch — biases its estimates. We introduce Spark-Attn, comprising Spark-Reblock and Spark-Reweight, to improve the two branches. We refer to its MiniMax-H3 implementation as Spark-H3.

01 / METHOD

Spark-Reblock

~2 min read

Block-level scoring works better when tokens within each block are more similar. Common block layouts group consecutive tokens or use fixed spatio-temporal tiles. These layouts provide efficient, regular blocks, but do not account for differences across inputs and layers.

Two aspects of video diffusion help explain why these differences matter.

Noise weakens the spatio-temporal prior. Denoising starts from independent Gaussian noise and gradually recovers structure. Nearby positions in the noise do not have an inherent similarity advantage. A spatio-temporal neighborhood that aligns well with a clean video may therefore be less well suited to the attention-relevant similarity structure during early denoising.

Denoising from noise to data
01Illustration: Yang Song, Generative Modeling by Estimating Gradients of the Data Distribution.

Similarity patterns vary across inputs and layers. Each layer and head learns its own similarity measure, so which tokens are similar depends on both the model’s learned weights and the input. A single fixed block partitioning may not match these varying similarity patterns well.

We propose Spark-Reblock to account for both effects by grouping tokens with similar attention preferences into the same blocks. More precisely, two queries are considered similar if they produce similar attention-score patterns over the key distribution, while two keys are considered similar if they receive similar scores across the query distribution. We quantify these relationships using a Mahalanobis cosine distance induced by the opposite side's second moment:

dQ(qi,qj)=1−cos⁡ ⁣(MK1/2qi, MK1/2qj),MK=E[kk⊤], d_Q(q_i,q_j)=1-\cos\!\big(M_K^{1/2}q_i,\,M_K^{1/2}q_j\big),\qquad M_K=\mathbb{E}[kk^\top],

and symmetrically for keys with MQ=E[qq⊤]M_Q=\mathbb{E}[qq^\top]. These distances make “similar” explicit: query tokens are grouped under dQd_Q, and key tokens under its symmetric counterpart dKd_K. Starting from the full token set, our iterative hierarchical partitioner repeatedly splits every current group according to the corresponding distance while enforcing the assigned child capacities. The process continues level by level until every leaf contains exactly the target block size. Because the hierarchy adapts to the current input, layer, and head, it can follow their different similarity structures; its time complexity is O(Nlog⁡N)\mathcal{O}(N \log N), where NN is the number of tokens.

The following animation illustrates recursive partitioning: split the parent according to assigned child capacities, apply the same operation within each child, and stop at the leaf block size. Token IDs track membership through the tree and the attention matrix shows the resulting permutation.

For ablation, we use a BSA baseline with fixed block partitioning. Each query block's exact branch covers the top 10% of video key-value blocks; its compressed branch mean-pools the remaining blocks. For a query block QAQ_A:

BSA⁡(QA)=Softmax⁡ ⁣(QAK~A⊤d)V~A. \operatorname{BSA}(Q_A)=\operatorname{Softmax}\!\left(\frac{Q_A\widetilde{K}_A^\top}{\sqrt{d}}\right)\widetilde{V}_A.

The effective keys K~A\widetilde{K}_A and values V~A\widetilde{V}_A are assembled by concatenating over all key-value blocks BB:

K~A=Concat⁡B{KB,B∈SA,1∣B∣kˉB⊤,B∉SA,V~A=Concat⁡B{VB,B∈SA,1∣B∣vˉB⊤,B∉SA. \begin{aligned} \widetilde{K}_A &=\operatorname{Concat}_{B} \begin{cases} K_B, & B\in\mathcal{S}_A, \\[4pt] \mathbf{1}_{|B|}\bar{k}_B^\top, & B\notin\mathcal{S}_A, \end{cases} \\[8pt] \widetilde{V}_A &=\operatorname{Concat}_{B} \begin{cases} V_B, & B\in\mathcal{S}_A, \\[4pt] \mathbf{1}_{|B|}\bar{v}_B^\top, & B\notin\mathcal{S}_A. \end{cases} \end{aligned}

Here, SA\mathcal{S}_A is the set of blocks selected for the exact branch, and kˉB\bar{k}_B and vˉB\bar{v}_B are the mean key and value of block BB. The all-ones vector 1∣B∣\mathbf{1}_{|B|} repeats each mean across the block's token positions, so every token contributes with equal weight — a uniform mean estimate. A single row-wise softmax normalizes both branches' contributions together.

The compressed branch matches Sol-Attn's mean-pooled approximation, while routing differs: we use a fixed Top-K budget instead of a content-adaptive threshold. Context/sink tokens remain exact and are outside the video-block Top-K budget.

The ablation replaces only the baseline's fixed block partitioning with Spark-Reblock; routing, the sparse schedule, and all other settings remain unchanged.

Method density ↓ Attention-mass recall ↑ PSNR ↑ SSIM ↑ LPIPS ↓
BSA 10% 68.66% 18.61 0.68 0.26
BSA + Reblock 10% 81.54% 22.41 0.80 0.14

At a Top-K attention budget of 10%, Spark-Reblock increases attention-mass recall by 12.88 percentage points and improves mean PSNR by 3.80 dB. These results demonstrate the benefit of adaptive block partitioning over fixed block partitioning at the same attention budget.

Discussion

BSA, Sol-Attn, Veda, and Spark-Attn can be understood within a common block-sparse framework, with each method refining a different part of the pipeline.

  • BSA uses a fixed block partition, scores blocks from block-level means, and computes exact softmax attention for the Top-K selected blocks.
  • Sol-Attn additionally approximates each unselected block with its mean rather than discarding the remaining context. It assumes that the block-level attention logits approximately follow a Gaussian distribution and uses μ+τσ\mu+\tau\sigma as the cutoff for exact attention. This dynamic threshold routing adapts across layers, attention heads, and query blocks rather than using a fixed Top-K budget.
  • Veda learns a scoring model that is more accurate than block-mean scoring at predicting which blocks should receive exact attention, improving Top-K selection at the same budget while still discarding unselected blocks.
  • Spark-Attn improves both block construction and the treatment of unselected blocks: Spark-Reblock groups tokens with similar attention preferences rather than relying on fixed spatio-temporal neighborhoods, increasing the attention fidelity of the selected Top-K blocks at the same budget, while Spark-Reweight uses weighted summaries and a log-mass bias to estimate the remaining blocks more accurately.

Concurrent work VC-Attention uses single-level online kk-means guided by value-token similarity to reorder keys and values for low-bit quantization. Because the cluster sizes are unconstrained, the sorted sequence is subsequently partitioned at fixed hardware-block boundaries to retain a regular, fixed-size block layout. Spark-Reblock instead uses an iterative divide-and-conquer hierarchy guided by query-key-induced attention-preference similarity to construct fixed-capacity blocks for sparse attention.

LLSA also achieves O(Nlog⁡N)O(N\log N) sparse attention through a hierarchy, recursively pooling fixed neighboring blocks and applying coarse-to-fine Top-K selection. LLSA uses the hierarchy to search interactions over a fixed layout, while Spark-Reblock uses it to adapt the token layout itself. From a log-linear-attention perspective, Spark-Reblock can be understood as an O(Nlog⁡N)O(N\log N) attention-like probe of the underlying O(N2)O(N^2) dense attention distribution: instead of materializing the full attention matrix, it infers its structure and rearranges tokens so that fixed-size blocks align better with the resulting attention geometry.

02 / METHOD

Spark-Reweight

~2 min read

Mean pooling underestimates attention mass because of the Jensen gap. Attention exponentiates logits before normalization, but pooling first replaces E[exp⁡(x)]\mathbb{E}[\exp(x)] with exp⁡(E[x])\exp(\mathbb{E}[x]), which therefore underestimates the expected exponential. When pooled blocks are combined with exact blocks, this underestimate suppresses their contribution to the output.

Spark-Reweight mitigates the resulting bias by scoring each block once with a representative query. These scores produce weighted key and value summaries, along with a log-mass bias used to correct the block's total attention mass. The weighted summaries and bias are then reused to approximate the block contribution for each query. To avoid per-query weighting—which would incur the full N×NN \times N cost—we use the mean of the video queries as the representative. This only adds a linear O(N)O(N) scoring pass.

Mean pooling creates a Jensen gap, while a log-mass bias restores the expected exponential.
02In this two-logit illustration, mean pooling produces exp⁡(E[x])\exp(\mathbb{E}[x]), shown in blue, below the expected exponential E[exp⁡(x)]\mathbb{E}[\exp(x)], shown in orange. The green point shows the effective pooled logit after applying the log-mass bias; its exponential matches the expected exponential.

The animation illustrates how weighted summaries and a log-mass bias produce a corrected estimate of the compressed block's contribution.

Using the same BSA baseline and evaluation protocol as the Spark-Reblock ablation, this experiment retains the fixed block partitioning and disables reblocking. It changes only the compressed-branch estimator, replacing mean pooling with Spark-Reweight's weighted pooling and log-mass bias; routing remains unchanged. Absolute log-mass error measures the difference between the exact total log-mass of the compressed branch and its approximation.

Method density ↓ Abs log-mass error (nats) ↓ PSNR ↑ SSIM ↑ LPIPS ↓
BSA 10% 3.86 18.61 0.68 0.26
BSA + Reweight 10% 2.78 19.17 0.70 0.24

At the same 10% density, Spark-Reweight reduces mean absolute log-mass error by 28% and improves mean PSNR by 0.56 dB.

03 / EVALUATION

Benchmark Results

Evaluation setup. We compare Dense, Sol-Attn, and Spark-H3 variants on VBench prompts using a 20-point MiniMax-H3 schedule (19 denoiser steps), 1344×768 resolution, and 240 frames at 24 fps. Sol-Attn uses its official setting. We further evaluate a new efficient Spark-H3 variant, Spark-H3-Lite, which halves reblocking latency while delivering performance comparable to Spark-H3. Quality is paired against dense references with identical prompts and seeds.

Method density ↓ DiT
latency (s) ↓
DiT
speedup ↑
PSNR ↑ SSIM ↑ LPIPS ↓ ATTN
speedup ↑
Dense 100% 583.3 1.00× ∞ 1.00 0.00 1.00×
Sol-H3 — 355.4 1.64× 20.36 0.71 0.20 2.19×
Spark-H3, 10% Speed 10% 337.4 1.73× 23.30 0.80 0.14 2.33×
Spark-H3, 15% Balanced 15% 355.4 1.64× 24.45 0.83 0.11 —
Spark-H3, 20% Balanced 20% 371.1 1.57× 25.38 0.85 0.09 1.96×
Spark-H3, 30% Fidelity 30% 404.1 1.44× 27.07 0.88 0.07 1.76×
Spark-H3-Lite, 10% Speed 10% 330.2 1.77× 23.34 0.80 0.14 —
Spark-H3-Lite, 15% Balanced 15% 346.1 1.69× 24.47 0.83 0.11 —
Spark-H3-Lite, 20% Balanced 20% 362.6 1.61× 25.12 0.85 0.10 —
Spark-H3-Lite, 30% Fidelity 30% 395.4 1.48× 26.26 0.87 0.08 —
Additional densitiesShow 15% & 20%Collapse 15% & 20%

On the scored videos, VBench measures subject and background consistency, motion smoothness, imaging quality, and aesthetic quality; higher is better for every dimension. The standard Spark-H3 15% run has no VBench scores in this result set and is therefore omitted from this table.

Method Subject
consistency ↑
Background
consistency ↑
Motion
smoothness ↑
Imaging
quality ↑
Aesthetic
quality ↑
Dense 90.75 93.88 99.02 72.14 67.59
Sol-H3 91.14 94.10 98.96 72.12 67.49
Spark-H3, 10% Speed 90.77 94.24 99.01 71.81 68.22
Spark-H3, 20% Balanced 90.92 94.25 99.01 72.08 67.73
Spark-H3, 30% Fidelity 90.90 93.99 99.02 72.01 67.72
Spark-H3-Lite, 10% Speed 90.91 93.85 99.00 71.74 67.86
Spark-H3-Lite, 15% Balanced 90.92 94.18 99.02 71.84 68.01
Spark-H3-Lite, 20% Balanced 90.94 94.15 99.02 71.88 67.62
Spark-H3-Lite, 30% Fidelity 91.03 94.08 99.03 72.06 67.60
Additional densitiesShow 15% & 20%Collapse 15% & 20%

At 10% density, Spark-H3-Lite has lower denoising latency than Sol-H3 (330.2 versus 355.4 seconds) while preserving the dense output more faithfully across PSNR, SSIM, and LPIPS. Increasing density improves standard Spark-H3 fidelity monotonically; Spark-H3-Lite trades a small amount of fidelity at 20% and 30% for an additional 8–9 second reduction relative to standard Spark-H3 at the same density. All reported Spark-H3 variants remain close to Dense across the VBench dimensions.

04 / COMMUNITY

Spark-Integration

~4 min read

Few-step distillation reduces the number of denoising steps, while attention acceleration reduces the cost of each step. FastH3 and OpenVDN already combine these two complementary approaches. The experiments below follow the same strategy by integrating Spark-H3 into FastH3's Dense pipeline and the few-step pipelines from LightX2V, Larryvrh, Alibaba-PAI, and ByteDance.

Unless noted otherwise, the integration examples use 1344×768 resolution, 240 frames at 24 fps, and Spark-H3 at 10% attention density, denoted Spark-H3, 10% below. FastH3 and few-step runs use per-block torch.compile on a single NVIDIA RTX PRO 6000 Blackwell Server Edition GPU with 96 GB of memory. In the no-warmup setting, Spark-H3 runs from the first step. The warmup setting keeps the first transformer layer dense throughout and uses dense attention for the first step of each 4-step schedule and the first two steps of each 8-step schedule.

FastH3: an alternative for fixed block partitioning

The comparisons below show native FastH3 checkpoint variants alongside its V1 Dense checkpoint with Spark-H3, 10% as the attention backend. V1 variants use four steps; V2 VSA uses eight. The results illustrate how Spark-H3 transfers to a distilled dense checkpoint and suggest Spark-Reblock as an alternative to VSA's fixed spatio-temporal block partitioning.

Both VSA variants can load with all weights resident, but inference then exceeds the available 96 GB of GPU memory. We therefore use layerwise offloading. Their reported latencies include CPU–GPU transfer overhead and could be lower with more GPU memory.

Video comparisonsShow resultsHide results

LightX2V 8-step LoRA

LightX2V MiniMax-H3 few-step LoRA is evaluated with dense attention and two Spark-H3, 10% schedules using matching prompts and seeds under the shared protocol above.

Video comparisonsShow resultsHide results

Larryvrh 8-step LoRA

Larryvrh MiniMax-H3 few-step LoRA is evaluated with dense attention and two Spark-H3, 10% schedules using matching prompts and seeds under the shared protocol above.

Video comparisonsShow resultsHide results

Alibaba-PAI Acc 8-step LoRA

Alibaba-PAI MiniMax-H3-Acc is evaluated with dense attention and two Spark-H3, 10% schedules using matching prompts and seeds under the shared protocol above. The runs follow its official eight-step PDD schedule.

Video comparisonsShow resultsHide results

ByteDance DMAD 4-step LoRA

ByteDance DMAD is evaluated with dense attention and two Spark-H3, 10% schedules using matching prompts and seeds under the shared protocol above. The runs follow its official four-step re-noising schedule.

Video comparisonsShow resultsHide results

SelfLift two-stage sampling based on LBH upsampler

The LBH learned latent upsampler lifts the latent from 608×352 to 1344×768 after four low-resolution denoiser steps, followed by three high-resolution steps. Spark-H3 is applied only during the high-resolution steps. Both variants use matching prompts and seeds.

Video comparisonsShow resultsHide results

Video DeltaNet

Video DeltaNet (VDN) can be viewed as a special case of BSA in which each block corresponds to one latent frame and a fixed rule selects the exact branch. For each latent frame, this branch covers a local neighborhood of 15 latent frames together with the first and last latent frames, for 17 exact latent frames in total. The remaining context is handled by a DeltaNet branch adapted through training rather than by a pooled compressed branch.

Because there is no corresponding dense checkpoint for the Video DeltaNet weights, we cannot directly integrate Spark-H3 into the same model. We therefore compare their speedups under VDN's longer 345-frame, 14.4-second setting. The table reports average latency per denoising step and measures each speedup against Dense at the same precision.

Method Dtype Per-step DiT
latency (s) ↓
Per-step ATTN
latency (s) ↓
DiT
speedup ↑
ATTN
speedup ↑
Dense BF16 55.96 49.07 1.00× 1.00×
Spark-H3, 10% (w/ warmup) BF16 31.45 24.66 1.78× 1.99×
Spark-H3, 10% (w/o warmup) BF16 22.48 15.74 2.49× 3.12×
VDN (8 steps) BF16 24.51 17.93 2.28× 2.74×
Dense FP8 51.47 47.24 1.00× 1.00×
Spark-H3, 10% (w/ warmup) FP8 26.98 22.82 1.91× 2.07×
Spark-H3, 10% (w/o warmup) FP8 17.98 13.86 2.86× 3.41×
VDN (8 steps) FP8 19.76 15.81 2.60× 2.99×

In the fully sparse, no-warmup setting, Spark-H3 reaches per-step DiT speedups in a similar range to VDN: 2.49× versus 2.28× in BF16 and 2.86× versus 2.60× in FP8. However, without training adaptation, this fully sparse schedule may deviate more from the dense model. Its outputs may still be visually plausible, but fidelity to the dense model is not guaranteed. We therefore treat these results only as an optimistic speed upper bound, rather than evidence of comparable quality-preserving acceleration. Under the warmup schedule used to preserve dense fidelity, VDN remains faster. The upper-bound results instead suggest that training adaptation could potentially bring Spark-H3 closer to VDN-level speedups while maintaining fidelity.

05 / PREVIEW

Spark-Ref2VA Preview

~1 min read

This preview applies Spark-H3 to the official 5-second Ref2VA case at 1344×768 and 24 fps. All three outputs use the same prompt, seed, and 19-step schedule. The comparison shows Dense, Spark-H3, and Spark-H3 with Compact Cond under the same generation setup.

Video comparisonsShow resultsHide results
06 / SHOWCASE

Visual Comparisons

The comparisons below show 10-second videos at 1344×768 resolution. Dense and Spark-H3, 10% use matching prompts and seeds, and each video shows its measured DiT latency.

08 / CHANGELOG

Update history

  • October 8, 2026: Expanded the evaluation and integration results:
    1. Added results for the Spark-H3-Lite variant and expanded its density sweep to 10%, 15%, 20%, and 30%.
    2. Optimized the kernel implementation, remeasured DiT latency and speedup, and corrected the speed of the Spark-H3 w/ warmup variant in the VDN table.
    3. Added Acc-LoRA and DMAD-LoRA results.
    4. Added SelfLift-LBH results.
    5. Added evaluation results for the Ref2VA variant.
    6. Added an introduction to Veda.
  • September 28, 2026: Added Spark-H3-30pct benchmark results; updated the reported benchmark step count to 19 denoiser steps, corresponding to MiniMax-H3's 20-point sigma schedule.
  • September 23, 2026: Official publication (release).