Make any photo sing with AI in minutes

Turn a static image and a favourite song into a lifelike AI-generated singing video with realistic lip movements and natural facial expressions. Pick any portrait, drop in audio, and share to TikTok and Reels.

A smiling person sings inside a rounded white card with a floating Lip sync waveform panel marked Synced, blue Orby mascot peeking at the top-left.
152,608,831Videos generated
128,114,826Avatars generated
21,275,323Videos translated
company logo 1
company logo 2
company logo 3
company logo 4
company logo 5
company logo 6
company logo 7
company logo 8
company logo 9
company logo 10
company logo 11
company logo 12
company logo 13
company logo 14
company logo 15
company logo 16
company logo 17
company logo 18
company logo 19
company logo 20
company logo 21
company logo 22
company logo 23
company logo 24
company logo 25
company logo 26
company logo 27
company logo 28
company logo 29
company logo 30
company logo 31
company logo 32
company logo 33
company logo 34
company logo 35
company logo 36
Trusted by millions worldwide to bring their stories to life.
Stylised white car icon on a blue background.Key features

Features of the Make Photo Sing tool

Pixel-accurate AI lip-sync engine

Every syllable, breath, and beat in your favourite song lines up with the face. The AI lip sync enginesynchronises mouth shapes at the phoneme level, so each performance looks real instead of robotic. Even held notes and runs land clean.

A close-up portrait of a person mid-song inside a white card beside a Lip sync waveform timeline with phoneme markers and a green Synced pill.

Smart face detection on any portrait

Drop in a selfie, a pet picture, a cartoon character, or a vintage portrait. The image to video pipeline reads facial features, locks framing, and reconstructs subtle head movement so each animation holds together across long clips and close-up shots.

Three stacked thumbnails of a selfie, a dog, and a cartoon character, each with a green face-detection box and a Detect face pill inside a white card.

Bring your song or AI-generated voice

Drop in an MP3 or WAV at your full track length. Prefer a custom narration? Generate one in the AI voice generator and route it straight into the singing photo. Both options use the same advanced engine for matching quality.

A man plays an acoustic guitar with an MP3 icon and Japanese text overlays.

Expressive facial animation, not robotic

Pick a mood and the AI face follows. Subtle blinks, smiles, brow lifts, and head tilts appear automatically, so a ballad feels tender and a hype track feels fired up. No frame-by-frame work and no animation timeline to wrestle with later.

A photoreal face mid-expression inside a white card with a mood selector row of pill chips Happy, Tender, and Hyped with Tender selected in green.

Built-in captions and vertical exports

Burn in word-by-word captions with the subtitle generator and export 9:16 for Reels, 1:1 for feed, or 16:9 for YouTube. Safe text placement keeps your hook inside the frame on every platform from the first view, every time you post.

A man performing into a microphone, shown in three different video aspect ratios: 9:16, 16:9, and 1:1, with captions.
Green gift box icon.Use cases

Use cases

A blue cursor dragging a smiling Black woman’s photo into an image upload box on a digital interface.

Viral TikTok, Reels and Shorts clips

Static selfies don’t stop the scroll. Turn any photo into a singing image of the latest trend, post a few variations in a row, and let the format and timing do the work for you across feeds.

A warm portrait of a grandmother singing inside a white card with a “Happy Birthday” pill and a confetti accent on a lavender background.

Birthday and anniversary messages

A flat ecard feels like a chore. Make Grandma sing happy birthday in her own voice, send the in-joke version to your best friend, or hand a custom song over for a wedding toast that lands as fun and engaging.

A small dog wearing a yellow beanie and headphones in front of a microphone.

Pet and animal singing video clips

Pets are pure social fuel. Upload a clear photo of your dog, cat, or hamster, drop in a hilarious song to sing your favourite tunes, and watch I Love Happy Cats–style cuteness spread across every feed you post to.

A stylised cartoon VTuber avatar on a small stage inside a white card with an Animate panel and a blue VTuber pill on a lavender background.

Cartoon and VTuber singing performances

Animate faces from illustrations, channel mascots, AI-generated cartoon characters, or your virtual identity without rigging or motion capture. The animator gives every avatar a real on-stage performance, post after post, with no filming time in your calendar.

Happy cartoon boy with messy brown hair, glasses, and a red jumper, arms raised.

Music videos for independent artists

Got a track and no budget? Drop your single onto an artist portrait, clone your vocal with AI voice cloning for backing layers, and publish a lively singing video the same day you finish mixing the song.

A woman in sunglasses holds a white fluffy dog in a lift, with colourful abstract shapes floating around.

Language learning and classroom moments

Make historical figures sing their dates, animate book characters reading lyrics, or send a song through the AI video translator into 175+ languages so the same lesson is ready to share across multiple languages and classrooms.

Blurred white document icon with a play button on a light blue background.How it works

How Make Photo Sing works

Make your photo sing in four steps. Animate photos with no filming, no editing experience, and no plugins. Add the image, pick a mood, generate, and post in minutes.

step icon

Step 1: Upload a photo

Pick a clear front-facing image of a person, pet, or character. The AI detects facial features.

step icon

Step 2: Add your song

Drop in an MP3 or WAV file, pick a song from your library, or paste in lyrics for the AI to sing.

step icon

Step 3: Choose mood and aspect

Choose an emotion to guide the performance and the vertical or square format where you plan to post.

step icon

Step 4: Generate and download

Render the AI photo output, preview the singing result, and download an MP4 ready for any platform.

Frequently asked questions

What does make photo sing mean, and how does it work?

It means turning a still image into a short video where the face performs an audio track. The AI will automatically detect facial features, map phonemes to mouth shapes, then add blinks and head tilts so the photo brings images to life singing your chosen song.

How do I make a photo sing online for free with HeyGen?

Sign up for a free HeyGen account online, upload a clear front-facing portrait image, drop in an audio file, and hit generate. The free plan covers short clips with a watermark, while a paid plan unlocks longer renders and HD output.

Which photos and image types work best for singing videos?

Clear, well-lit, front-facing headshot images work best. Avoid heavy occlusions, sunglasses, extreme side angles, and low resolution. Pet shots, cartoons, and illustrations all work as long as the face is visible and centred in the frame.

Can I upload my own song or my own voice recording?

Yes. Drop in any MP3 or WAV, including original tracks, covers, voice memos, or instrumentals. The AI reads vocals, beat, and tone, then maps the performance to match the rhythm. Confirm you hold the rights to any commercial music before you post.

Can my photo sing in languages other than English?

Yes. HeyGen handles multiple languages with natural pronunciation that matches the rhythm of each language. Run the vocal through the AI video translator into Spanish, Hindi, Mandarin, or any language, and the lip-sync follows.

Can I edit the AI singing video after I generate the first version?

Yes. Open the clip in the AI video editor to trim, adjust expression intensity, swap audio, regenerate alternate takes, or add captions, music beds, and brand colours before exporting your final cut.

Can animals, cartoons, or anime characters sing as well?

Yes. If the image has a recognisable face, the AI animates it. Dogs, cats, illustrated cartoon characters, anime portraits, and AI-generated avatars all work. Pet and cartoon singing is one of the most popular formats on the platform.

What if a photo has more than one person or face in it?

The tool animates one face at a time for the cleanest result. Crop the image to the person you want as the singer, or generate separate clips for each face and stitch them together in post for a group scene.

Can I use the singing photo videos commercially for clients?

Yes, you own the outputs. Confirm you hold rights to any third-party music, photos, or voices you upload. HeyGen serves 85,000+ businesses, from solo creators like Anton Voroniuk reaching 1M+ students to global brands.

How fast is generation, and what video format do I receive?

Most short clips render in a few minutes. The AI-powered output is MP4, available in 9:16 vertical, 1:1 for feed, and 16:9 for YouTube. Captions can be burned in for autoplay or kept separate for repurposing across channels.

Explore more AI powered tools

Bring any photo to life with hyper‑realistic voice and movement using Avatar IV.

Start creating with HeyGen

Make any photo sing with AI in minutes. Upload a picture, drop in a song, and share on TikTok and Reels.

CTA background