youtube sent me here!! lol
could you break down which work flow you did some of your videos?
i love the music videos, they look awesome!!
Anways i am a noob so any explanation one which of your workflows to work with that would be awesome!
do you have videos on each?
Hi!
Here’s a quick breakdown of the main workflows I used:
Music videos / Image-to-Video with audio:
https://huggingface.co/SOLRICKS/LTX-2-5-ComfyUI-Workflows/blob/main/LTX-2.5_T2V_I2V_Audio.json
This is the best place to start if you want to animate an image and generate a video with synchronized audio or music.
Green-screen replacement and compositing:
https://huggingface.co/SOLRICKS/LTX-2-5-ComfyUI-Workflows/blob/main/LTX-2.5-ReTake-Green-Screen-Composite.json
I used this for the videos where the original green-screen background is replaced with a new AI-generated environment.
Animated characters and reference-based scenes:
https://huggingface.co/SOLRICKS/LTX-2-5-ComfyUI-Workflows/blob/main/LTX-2.5-ICLoRA-Ingredients-Two-Stage.json
This is the workflow I use for animation videos where character identity, clothing, proportions, and environment consistency need to be preserved.
If you’re just getting started, I recommend trying the T2V/I2V Audio workflow first. The workflows may look complicated at first, but don’t worry — we were all beginners once!
I don’t currently have a separate step-by-step tutorial for every workflow, but I’m planning to create more walkthrough videos.
In the meantime, if you run into any problems, feel free to post a screenshot or error here and I’ll try to help.
I am going to work on some tomorrow! do you have any examples of the The Audio-IC-LoRA workflow on your youtube page ? is it similar to scail2 but with lipsyncing?
For lip-syncing, be sure to use this: https://huggingface.co/SOLRICKS/LTX-2-5-ComfyUI-Workflows/blob/main/LTX-2.5_T2V_I2V_Audio.json
i got it working on my shadow pc! you need to post somewhere about the exact time on the audio i used ffmpeg
ffmpeg -i song.mp3 -ss 0 -t 8 -ar 48000 -c:a libmp3lame -q:a 2 vocals8.mp3
To verify:
ffprobe -v error -show_entries format=duration -of csv=p=0 vocals8.mp3
If it comes back 8.016 or similar, mp3 frame padding is the culprit — export to wav instead
ffmpeg -i song.mp3 -ss 0 -t 8 -ar 48000 vocals8.wav
I also saw questions today in your comments about the two stage.... is that for character sheets and if so you have a prompt for making one ? i thought they asked you but i may have misread
thanks that work flow is sweet !
so i used that two stage and got the craziest looking video lol i did a sheet like this and got this
and i kept your prompt, you mentioned in your YouTube comments you used this to lip sync?
how can i control the video size? by the reference sheet? but why is references bleeding into the video the way it is?
No, this is not the workflow. Use this image and this workflow. https://huggingface.co/SOLRICKS/LTX-2-5-ComfyUI-Workflows/resolve/main/LTX-2.5-ICLoRA-Ingredients-Two-Stage.json?download=true
Download the LoRa model and place it in https://huggingface.co/SOLRICKS/LTX-2-5-ComfyUI-Workflows/resolve/main/ltx-2.3-22b-ic-lora-ingredients-0.9.safetensors?download=true
your LoRa folder.
took me a ton of time to get it working on my end but i think i got it... are you working on a way to have this work with different audio too? i normally use a Google TTS in my workflows but no idea how that could even work here. the voice and lipsync are awesome here but consistency would become a issue
and yes i noticed the hay in the mouth issue after wards... i have been testing this on a NVIDIA RTX A4500, 19190 MiB and each test takes 40 mins, so i did not feel like testing or fixing any more today think i did 13 or 14 today and im exhausted i tried so many tweaks lol
I got it! well i just stripped the audio from last video and with some help from claude i was able to get this video
I couldn't figure out how to get this workflow like you did on huggingface so I added to pastebin : https://pastebin.com/0WbP27wK
Love to hear what you think and what you would add... and would love to see some prompting you did on that music video "Sneezers" industrial style metal song.. or heck the sax one is awesome too lol. did you do some of the effects in post on those or all AI?
Exactly. I'm not sure why you're facing so many problems, we receive it in a shorter time 10second maybe 14minute..
Hi, Plase use this v2 version with melband reformer for cleane lip-sync
https://huggingface.co/SOLRICKS/LTX-2-5-ComfyUI-Workflows/resolve/main/LTX-2.5-ICLoRA-Ingredients-Two-Stage-V2.json?download=trueLTX-2.5 IC-LoRA Ingredients + Audio Lip-Sync v2
Reference-driven I2V workflow for LTX-2.5 combining IC-LoRA Ingredients with MelBand RoFormer lip-sync conditioning.
• Multi-reference character & environment conditioning
• MelBand RoFormer vocal separation for cleaner lip-sync
not gonna lie, i cant recreate your results... the 1st time i tried it worked BUT the wrong character spoke... nothing i tried fixed that.
What prompt did you use? is there a set structure? i got like 10 of the same character speaking and this was my prompt, i downloaded your video and ripped the audio to see if something i was doing was wrong
here are a few examples i got i cant even get the original one working again lol
====================== Prompt 1 ======================
Reference sheet: The reference sheet contains the same blue donkey character, the same green crocodile wearing a brown cowboy hat and brown leather jacket, and the same sunlit garden with stone paving, flower beds, a stone bench, and a stone urn.
Generated video: Create one single full-screen cinematic scene entirely inside the exact same sunlit garden shown in the reference sheet. Use the reference sheet only to preserve the identity of the two characters and the garden environment. Do not reproduce the layout of the reference sheet. Do not show multiple panels. Do not show turnaround views. Do not show a collage. The entire frame must show only one continuous view of the garden.
Use one static medium shot with both characters large in frame, filling the lower two-thirds of the image, the garden path, flowers, and stone bench visible behind them. The blue donkey stands on the left. The green crocodile stands on the right, holding a small garden trowel. Keep the exact same warm golden sunlight, soft haze, and garden continuity throughout.
The audio contains four spoken lines separated by clear pauses. The characters alternate speakers, one line each, crocodile first.
First line: the crocodile tilts his head with a curious, slightly mischievous expression and speaks, asking his question. Only the crocodile's mouth moves. The donkey's mouth stays completely closed as he listens.
During the pause, both mouths are completely closed.
Second line: the donkey straightens up proudly and answers. Only the donkey's mouth moves. The crocodile's mouth stays completely closed, one eyebrow rising.
During the pause, both mouths are completely closed.
Third line: the crocodile glances around at the overgrown flower beds and speaks again, skeptical. Only the crocodile's mouth moves. The donkey's mouth stays completely closed as his smile falters slightly.
During the pause, both mouths are completely closed.
Fourth line: the donkey puffs his chest and replies with confident reassurance. Only the donkey's mouth moves. The crocodile's mouth stays completely closed, giving him a long flat look.
After the final line, both characters are silent with mouths closed. The donkey turns his head to look at a drooping flower beside them. The crocodile keeps his confident pose. Hold this as an exact continuity point for the next video.
Maintain strict character identity. The donkey must exactly match the blue donkey reference. The crocodile must exactly match the green crocodile reference, including his brown cowboy hat, brown leather jacket, and belt. Preserve the exact same garden, character scale, and warm golden lighting.
Clean, high-quality stylized 3D animated feature-film look, warm soft cinematic lighting, subtle facial animation, restrained body movement, natural breathing and blinking, and accurate lip synchronization matching the audio timing exactly.
One scene only. One camera only. No camera cuts. No camera movement. No walking. No exaggerated gestures. No extra dialogue. No narration.
Total duration: 9 seconds.
====================== Prompt 2 ======================
Reference sheet: The reference sheet contains the same blue donkey character, the same green crocodile wearing a brown cowboy hat and brown leather jacket, and the same sunlit garden with stone paving, flower beds, a stone bench, and a stone urn.
Generated video: Create one single full-screen cinematic scene entirely inside the exact same sunlit garden shown in the reference sheet. Use the reference sheet only to preserve the identity of the two characters and the garden environment. Do not reproduce the layout of the reference sheet. Do not show multiple panels. Do not show turnaround views. Do not show a collage. The entire frame must show only one continuous view of the garden.
Use one static medium shot with both characters large in frame, filling the lower two-thirds of the image, the garden path, flowers, and stone bench visible behind them. The blue donkey stands on the left. The green crocodile stands on the right, holding a small garden trowel. Keep the exact same warm golden sunlight, soft haze, and garden continuity throughout.
The scene follows the audio exactly, in four timed beats.
From 0.5 to 2.0 seconds: the donkey tilts his head with a curious, slightly mischievous expression and speaks, asking his question. Only the donkey's mouth moves. The crocodile's mouth stays completely closed as he pauses mid-motion with the trowel.
From 2.4 to 3.5 seconds: the crocodile straightens up proudly and answers. Only the crocodile's mouth moves. The donkey's mouth stays completely closed, one eyebrow rising.
From 3.8 to 4.8 seconds: the donkey glances around at the overgrown flower beds and speaks again, skeptical. Only the donkey's mouth moves. The crocodile's mouth stays completely closed as his smile falters slightly.
From 5.3 to 7.0 seconds: the crocodile puffs his chest and replies with confident reassurance. Only the crocodile's mouth moves. The donkey's mouth stays completely closed, giving him a long flat look.
From 7.0 to 8.0 seconds: both characters are silent with mouths closed. The donkey turns his head to look at a drooping flower beside them. The crocodile keeps his confident pose. Hold this as an exact continuity point for the next video.
Maintain strict character identity. The donkey must exactly match the blue donkey reference. The crocodile must exactly match the green crocodile reference, including his brown cowboy hat, brown leather jacket, and belt. Preserve the exact same garden, character scale, and warm golden lighting.
Clean, high-quality stylized 3D animated feature-film look, warm soft cinematic lighting, subtle facial animation, restrained body movement, natural breathing and blinking, and accurate lip synchronization matching the audio timing exactly.
One scene only. One camera only. No camera cuts. No camera movement. No walking. No exaggerated gestures. No extra dialogue. No narration.
Total duration: 8 seconds.
======================
prompt 1
and then
Prompt 2
what am i doing wrong? is it the prompting? and heck i just wanted the gator walking up and asking what donkey was doing lol
-- Yes, the problem here lies largely in the prompt structure. In particular, there is a critical contradiction between the two prompts.
Reference sheet: The reference sheet contains one blue donkey character and one green crocodile character wearing a brown cowboy hat, brown leather jacket, and belt, inside the same warm sunlit garden.
Generated video: Create one single continuous full-screen cinematic shot inside the garden from the reference sheet. Use the reference sheet only to preserve the exact character identities and garden environment. Do not reproduce the reference sheet layout, panels, turnaround views, or multiple character copies.
Use one static medium-wide camera shot.
The blue donkey is standing on the LEFT side of the frame, already working quietly in the garden.
The green crocodile enters from the RIGHT side of the frame, walks a few steps toward the donkey, stops beside him, looks at what the donkey is doing, and asks him a question.
IMPORTANT SPEAKER ASSIGNMENT:
The FIRST audible voice in the supplied audio belongs ONLY to the GREEN CROCODILE.
The SECOND audible voice belongs ONLY to the BLUE DONKEY.
Never swap the speakers.
Never assign the crocodile's voice to the donkey.
Never assign the donkey's voice to the crocodile.
While the crocodile speaks, only the crocodile performs visible speech and lip movement. The donkey listens silently with a closed relaxed mouth.
When the crocodile finishes speaking, both characters briefly pause.
Then the donkey answers. While the donkey speaks, only the donkey performs visible speech and lip movement. The crocodile listens silently with a closed relaxed mouth.
Follow the supplied audio naturally. Use the actual pauses in the audio rather than inventing additional dialogue or changing the speaker order.
Character identity must remain stable throughout the entire video.
There must be exactly ONE blue donkey and exactly ONE green crocodile.
The donkey must remain on the left.
The crocodile must remain on the right after approaching him.
Do not duplicate either character.
Do not transform one character into the other.
Do not create additional versions of either character.
Preserve the crocodile's brown cowboy hat, brown leather jacket, and belt throughout the entire shot.
Use subtle natural facial animation, accurate lip synchronization, natural blinking, small restrained body gestures, and gentle breathing.
The crocodile's walking motion should be short, slow, and natural. After reaching the donkey, both characters remain mostly stationary during the conversation.
Warm golden garden lighting, colorful flowers, stone garden path, soft cinematic depth of field, high-quality stylized 3D animated feature-film appearance.
One scene only.
One camera only.
No cuts.
No zoom.
No camera movement.
No narration.
No additional dialogue.
No subtitles.
No text.
Total duration: approximately 8–10 seconds.
Negative:
reference sheet layout, split screen, multiple panels, duplicated donkey, duplicated crocodile, extra characters, speaker swap, wrong character speaking, simultaneous talking, both mouths moving, character morphing, identity changes, clothing changes, camera cuts, camera movement, subtitles, text, watermark
i am having issue with the walking an talking... can you show me your prompt with the walking raccoon and the hedge hog
So mine seems to be "the one who stops first" is a fair way to say it — the first-planted character gets the first line. so like who ever is stationary first.
this is only thing i have gotten to actually work:
Reference sheet: The reference sheet contains one blue donkey character and one green crocodile character wearing a brown cowboy hat, brown leather jacket, and belt, inside the same warm sunlit garden.
Generated video: One single continuous full-screen cinematic shot inside the garden from the reference sheet. Use the reference sheet only to preserve the exact character identities and garden environment. Do not reproduce the reference sheet layout, panels, or multiple character copies. One static medium-wide camera shot.
The green crocodile is standing on the LEFT side of the frame, already in the garden beside the flower beds. The blue donkey stands on the RIGHT side of the frame a few steps from the crocodile, standing stiffly at attention like a little patrol guard, facing the crocodile.
The crocodile looks at the donkey and asks him a question. Only the crocodile's mouth moves. The donkey listens with his mouth closed.
The donkey lifts his head and answers. Only the donkey's mouth moves. The crocodile listens with his mouth closed.
The crocodile asks again, skeptical. Only the crocodile's mouth moves. The donkey listens with his mouth closed.
The donkey answers proudly. Only the donkey's mouth moves. The crocodile listens with his mouth closed.
After the last line, both stand silent with mouths closed and hold their positions.
Exactly one donkey and one crocodile. The crocodile keeps his brown cowboy hat, brown leather jacket, and belt throughout. Subtle natural facial animation, accurate lip synchronization, natural blinking, small restrained gestures. Warm golden garden lighting, colorful flowers, stone garden path, high-quality stylized 3D animated feature-film look. One camera, no cuts, no camera movement, no narration, no extra dialogue.
Total duration: 9 seconds.
thanks for helping out so much!
Reference sheet: The reference sheet contains one blue donkey character and one green crocodile character wearing a brown cowboy hat, brown leather jacket, and belt, inside the same warm sunlit garden.
Generated video: One single continuous full-screen cinematic shot inside the garden from the reference sheet. Use the reference sheet only to preserve the exact character identities and garden environment. Do not reproduce the reference sheet layout, panels, or multiple character copies. One static medium-wide camera shot.
The green crocodile is standing on the LEFT side of the frame, already in the garden beside the flower beds. The blue donkey stands on the RIGHT side of the frame a few steps from the crocodile, standing stiffly at attention like a little patrol guard, facing the crocodile.
The crocodile looks at the donkey and asks him a question. Only the crocodile's mouth moves. The donkey listens with his mouth closed.
The donkey lifts his head and answers. Only the donkey's mouth moves. The crocodile listens with his mouth closed.
The crocodile asks again, skeptical. Only the crocodile's mouth moves. The donkey listens with his mouth closed.
The donkey answers proudly. Only the donkey's mouth moves. The crocodile listens with his mouth closed.
After the last line, both stand silent with mouths closed and hold their positions.
Exactly one donkey and one crocodile. The crocodile keeps his brown cowboy hat, brown leather jacket, and belt throughout. Subtle natural facial animation, accurate lip synchronization, natural blinking, small restrained gestures. Warm golden garden lighting, colorful flowers, stone garden path, high-quality stylized 3D animated feature-film look. One camera, no cuts, no camera movement, no narration, no extra dialogue.
Total duration: 9 seconds.
example and it still gets the order wrong! i cant crack why it decides who talks first