Put two faces into the orange-booth duet. The moves, framing and sound stay exactly as in the template.
Left performer
Right performer
Best results: a sharp, front-facing photo with the whole face visible and no sunglasses. Only upload people who agreed to it.
15s with the original sound. Wan 3.0 bills the 15s template plus the 15s it renders.
Hotel Lobby AI
Drop two faces into the viral orange-booth duet. Upload one photo per person and Wan 3.0 swaps them into the template: every gesture, camera angle and beat of the original sound stays the same.
A real run of the effect. Nothing was edited by hand.


One clear, front-facing photo for each performer. Image 1 takes the left side of the microphone, Image 2 the right. Use the swap button to flip them.
The template, prompt and settings are already set. Wan 3.0 keeps the booth, the moves and the audio, and replaces only the two faces.
The 15-second clip with its original sound lands in the preview and in My videos, ready for TikTok, Reels or Shorts.
It edits the real performance instead of inventing a new one.
Wan 3.0 reads the template clip as a video reference, so the hand gestures, head bobs, lip movement and timing come straight from the source.
Each photo is tied to one performer. Faces don't blend or trade places halfway through the clip.
The duet keeps its original audio, so the result is ready to share without adding music in an editor.
The face-swap instructions are tuned for this template. You only choose who is in the booth.
Front-facing or slightly turned, eyes visible. Sunglasses, masks and hands over the face weaken the likeness.
Crop out friends in the background so the model knows exactly whose face to use.
Daylight or even indoor light works best. Blurry or heavily filtered selfies come out softer.
Best friends, a couple, you and your dad, the same person twice in different outfits: unexpected pairings get the most shares.
Wan 3.0 bills the 15-second template it reads plus the 15 seconds it renders, at 18 credits per second in 720P. Standard 480P costs 270 credits. Failed runs are refunded automatically.
Want your own scene? Try Wan 3.0Have a different question and can't find the answer you're looking for? Reach out to our support team by sending us an email and we'll get back to you as soon as we can.
It's a viral format built on an orange-booth rap duet with a hanging microphone. Creators swap the two performers for friends, partners, family or themselves and post the result on TikTok, Reels and Shorts.
The effect sends the template clip to Wan 3.0 as a video reference together with your two photos. The model keeps the motion, framing and sound of the clip and redraws only the two faces.
540 credits for a 15-second 720P video with sound, or 270 credits in 480P. As with every Wan 3.0 video reference, the 15-second template is billed along with the 15 seconds rendered. Failed runs are refunded automatically.
Image 1 replaces the performer on the left, Image 2 the performer on the right. Use the swap button between the photo slots to flip them before you generate.
Yes. Upload two different photos of the same person and they will appear on both sides of the microphone.
Usually 15 to 20 minutes, because Wan 3.0 redraws all 15 seconds of the template. You can leave the page; the finished video is saved to My videos.
Only photos of yourself or of people who agreed to it. Don't upload photos of minors, and don't use the effect to impersonate or mislead anyone.