ByteDance, the minds behind TikTok, just unveiled OmniHuman-1, a revolutionary AI model that’s redefining what’s possible in video generation. OmniHuman-1 blends images with audio or motion signals to create stunningly realistic human animations. The model leverages diverse data types (like audio and body poses) to produce natural movements and expressions without heavy data filtering. Check it out!
This is mostly likely the best image to video to date.
Link to OmniHuman Project Page - https://omnihuman-lab.github.io/ OmniHuman Arxiv
Link to Paper - https://arxiv.org/pdf/2502.01061
I dropped a video that went over the paper a bit in detail while showcasing some examples. You can watch it below.

